Cluster pursuit and evasion control method and system based on hierarchical game deep reinforcement learning
Through the method of deep reinforcement learning of hierarchical game, a hierarchical training model and a regularized auxiliary reward mechanism are built, which solves the problems of information incompleteness and computational complexity in multi-agent pursuit tasks, and improves the collaboration efficiency and adaptability of the agent system.
Patent Information
- Application Number
- CN202510450379.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing multi-agent pursuit method faces the problems of information acquisition limitations, low collaboration efficiency and high computing complexity in some observable dynamic environments, making it difficult to achieve efficient collaborative decision-making.
Using a method based on deep reinforcement learning of hierarchical games, a partial observable Markov decision-making process is constructed, and the collaboration strategy of the agent is optimized through hierarchical training of high-level strategy modules and low-level strategy modules, combined with regularized auxiliary reward mechanisms.
The collaboration efficiency and adaptability of multi-agent systems in complex environments are improved, and the problems of low collaboration efficiency and high computational complexity in some observable environments are solved, thus achieving efficient pursuit between agents.
Smart Images

Figure CN119960489B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of reinforcement learning, and in particular, to a cluster pursuit and evasion control method and system based on hierarchical game deep reinforcement learning. Background Art
[0002] In modern intelligent applications such as autonomous driving, UAV formation, intelligent security, and military reconnaissance, the multi-agent pursuit and evasion problem has gradually become a hot topic of concern in the academic and industrial fields. The core goal of the multi-agent pursuit and evasion task is to enable a group of "pursuers" to chase and capture an "evader" through cooperation. Such tasks are characterized by high real-time performance, complex environments, and frequent interactions, and usually require efficient cooperation and decision-making among multiple agents in a dynamic environment. Traditional pursuit and evasion methods mostly rely on global observation information and centralized control strategies, which have many limitations in actual application scenarios. Especially in partially observable environments, the lack of information acquisition and uncertainty become the main challenges in the multi-agent system pursuit and evasion task.
[0003] Existing multi-agent pursuit and evasion methods can be summarized into the following categories: traditional analytical methods, bio-inspired algorithms based on artificial intelligence, and multi-agent strategies based on deep reinforcement learning. These methods have certain practicality in some specific environments, but their limitations gradually become apparent when facing partially observable environments and large-scale agent scenarios.
[0004] Traditional analytical methods are based on differential game theory or geometric theory, and solve the pursuit and evasion problem by constructing an accurate mathematical model. Patent CN115891002A analyzes the pursuit and evasion game between multiple pursuers and evaders based on a mathematical model, and derives the optimal strategy of the pursuers through the model. Patent CN111947323A establishes a multi-agent confrontation task model based on differential equations, and obtains the pursuit and evasion strategy through the analysis and solution of the motion model of the agents. Although traditional analytical methods can play a certain role in fixed and simple environments, with the complexity of the scenario and the increase in the number of agents, traditional analytical methods become infeasible due to high computational complexity. Especially in partially observable dynamic environments, traditional methods cannot calculate the optimal decision in real time and are difficult to adapt to the uncertainty of the environment.
[0005] Bio-inspired algorithms achieve cooperation and adaptive behavior among multiple agents by simulating the hunting or group behavior of animals. Patent CN114422077A realizes the pursuit and evasion task allocation of agents based on the hunting behavior of wolf packs, and designs pursuit and evasion strategies by simulating the cooperation mechanism in wolf packs. Patent CN112233339A proposes an ant colony optimization algorithm for target search and path planning of unmanned aerial vehicles in uncertain environments. These bio-inspired algorithms have certain advantages in swarm intelligence and cooperation, but their computational complexity is relatively high, they are sensitive to parameters, have a large computational burden in dynamic environments, and are prone to falling into local optima, resulting in poor applicability in dynamic pursuit and evasion tasks.
[0006] In recent years, multi-agent deep reinforcement learning has shown great potential in solving multi-agent cooperation tasks. Patent CN114253649A uses a multi-agent deep reinforcement learning method to optimize the multi-agent cooperative pursuit and evasion strategy through a deep neural network, which is applicable to multi-agent cooperation in high-dimensional environments. This patent improves the policy learning effect of agents through centralized training and distributed execution. Patent CN113474985A uses reinforcement learning technology to achieve target tracking of robots in dynamic environments and improves policy performance through a self-play mechanism. Although these deep learning-based methods have achieved good results in multi-agent systems, they generally assume that the environment is globally observable and ignore the uncertainty brought by the limited perspective of agents in practical applications. In a partially observable environment, agents can only obtain limited local observation information, and this incomplete information significantly increases the difficulty of policy generation and optimization, resulting in poor performance of traditional multi-agent deep reinforcement learning methods in partially observable dynamic environments.
[0007] In a partially observable environment, the multi-agent pursuit and evasion task faces the following technical challenges:
[0008] Limitations in information acquisition: In a partially observable environment, agents can only obtain limited local information through sensors, resulting in the lack of global observation information. This incomplete information will affect the decision-making quality of agents, making it difficult for agents to make effective judgments when performing pursuit and evasion tasks.
[0009] Low cooperation efficiency: The effectiveness of multi-agent cooperation depends on real-time communication and information sharing among agents. However, in a partially observable environment, traditional centralized control systems are difficult to meet the cooperation requirements. When there is no effective cooperation strategy among multi-agents, it may lead to waste of resources and low capture efficiency.
[0010] High computational complexity: In a dynamic environment, multi-agents need to continuously update and adjust their strategies, and traditional centralized control methods have significant disadvantages in computational complexity. In addition, when the number of agents is large, centralized control will increase communication latency, resulting in a decrease in the real-time performance of the system.
[0011] In summary, the existing multi-agent pursuit-evasion methods have significant limitations when facing partially observable dynamic environments. How to design a distributed, flexible and efficient multi-agent cooperation strategy so that each agent can achieve the optimal pursuit-evasion strategy under incomplete information is an urgent problem to be solved currently. Summary of the Invention
[0012] The purpose of the embodiments of the present invention is to provide a cluster pursuit-evasion control method and system based on hierarchical game deep reinforcement learning, so as to improve the cooperation efficiency and response ability of multi-agent systems in complex environments.
[0013] In a first aspect, the present invention provides a cluster pursuit-evasion control method based on hierarchical game deep reinforcement learning, and the method includes:
[0014] Construct a cluster pursuit-evasion game scenario, where the cluster pursuit-evasion game scenario includes a game environment and multiple agents and obstacles in the game environment, and each agent is equipped with a sensor, and the sensor is used to detect local state information within the detection range of the sensor;
[0015] Construct a cooperative task model for multiple agents based on the partially observable Markov decision process;
[0016] Construct a deep reinforcement learning model for hierarchical games, where the deep reinforcement learning model includes a high-level policy module and a low-level policy module;
[0017] Collect the game information of the agents performing pursuit-evasion simulation according to the cooperative task model in the pursuit-evasion game scenario, where the game information includes local state information;
[0018] Use the game information to train the low-level policy module and the high-level policy module;
[0019] Obtain the guidance information of the agents performing the game by using the trained high-level policy module, and obtain the action information of the agents based on the guidance information through the low-level policy module, and control the agents to perform the pursuit-evasion game based on the action information.
[0020] In an optional implementation manner, the step of using the game information to train the low-level policy module and the high-level policy module includes:
[0021] Use the game information to train the low-level policy module;
[0022] When the number of training iterations of the low-level policy module reaches the update frequency of the high-level policy module, train the high-level policy module based on the game information.
[0023] In an alternative embodiment, the low-level policy module includes a shared value network and a shared policy network;
[0024] The step of training the low-level policy module using the game information includes:
[0025] For each agent, calculate the target value based on the game information of the agent, and update the shared value network by minimizing the loss function constructed by the current value and the target value;
[0026] Use the observation information in the game information and the guidance information provided by the high-level policy module as the input of the shared policy network, output the corresponding action information using the shared policy network, and update the shared policy network in a gradient ascent manner based on the cumulative reward under the action information.
[0027] In an alternative embodiment, the cumulative reward is obtained from a reward function with an auxiliary reward regularization term obtained by the agent within the cumulative time steps;
[0028] The auxiliary reward regularization term is constructed based on detecting whether the pursuer agent captures the evader agent, whether the pursuer agent detects the evader agent, whether the pursuer agent collides with an obstacle, and the number of evader agents captured by the pursuer agent.
[0029] In an alternative embodiment, the game information includes the observation information, action information, reward function information of the agent at the current time step, and the observation information at the next time step;
[0030] The step of training the high-level policy module based on the game information includes:
[0031] Input the observation information, action information, reward function information of the agent at the current time step, and the observation information at the next time step included in the game information into the high-level policy module to obtain the guidance information at the current time step and the guidance information at the next time step;
[0032] Calculate the current expected value based on the observation information and guidance information at the current time step;
[0033] Calculate the target expected value based on the observation information and guidance information at the next time step;
[0034] Conduct guidance training on the high-level policy module based on the loss function constructed by the current expected value and the target expected value.
[0035] In an alternative embodiment, the agents include a pursuer agent and an evader agent, and the tasks of the pursuer agent include a pursuit task and an avoidance task;
[0036] The guidance information includes target ID assignment, cooperation coefficient, and priority weight distribution;
[0037] The target ID assignment is used to determine the ID of the evader agent that each pursuer agent pursues;
[0038] The cooperation coefficient is used to determine the degree of cooperation between each pursuer agent and other pursuer agents;
[0039] The priority weight distribution includes the weight of the pursuit task and the weight of the avoidance task of each pursuer agent.
[0040] In an alternative embodiment, each agent determines whether it detects an observed object based on the carried sensor, where the observed object is any other agent or an obstacle, in the following manner:
[0041] Obtain the projected distance between the agent and the observed object in the sensor direction, and obtain the actual distance between the agent and the observed object;
[0042] Detect the actual relative velocity and the projected relative velocity between the agent and the observed object;
[0043] Based on the projected distance, actual distance, actual relative velocity, projected relative velocity, the radius of the observed object, and the sensor vector, determine whether the agent detects the observed object according to a preset judgment formula.
[0044] In an alternative embodiment, the cooperation task model includes an observation space model, and the steps of constructing the observation space model include:
[0045] Obtain the internal state information and external state information of each agent at each time step, where the external state information is the external local state information detected by the sensors of the agent;
[0046] Construct an observation space model based on the internal state information and the external state information.
[0047] In an alternative embodiment, the cooperation task model includes an action space model, and the steps of constructing the action space model include:
[0048] Obtain the acceleration information of each agent in the x-axis and the acceleration information in the y-axis at each time step;
[0049] Construct an action space model based on the acceleration information of each agent on the x-axis and the acceleration information on the y-axis.
[0050] In a second aspect, the present invention provides a swarm pursuit-evasion control system based on hierarchical game deep reinforcement learning. The system includes:
[0051] A construction module for constructing a swarm pursuit-evasion game scenario. The swarm pursuit-evasion game scenario includes a game environment and multiple agents and obstacles in the game environment. Each agent is equipped with a sensor, and the sensor is used to detect local state information within the detection range of the sensor.
[0052] The construction module is further configured to construct a cooperative task model for multiple agents based on the partially observable Markov decision process.
[0053] The construction module is further configured to construct a deep reinforcement learning model for hierarchical games. The deep reinforcement learning model includes a high-level policy module and a low-level policy module.
[0054] An acquisition module for acquiring the game information of the agents performing pursuit-evasion simulation according to the cooperative task model in the pursuit-evasion game scenario. The game information includes local state information.
[0055] A training module for training the low-level policy module and the high-level policy module using the game information.
[0056] A guidance module for obtaining guidance information for the agents performing the game using the trained high-level policy module, and obtaining action information of the agents based on the guidance information through the low-level policy module, and controlling the agents to perform the pursuit-evasion game based on the action information. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0058] Figure 1 It is a flowchart of the swarm pursuit-evasion control method based on hierarchical game deep reinforcement learning provided by the embodiment of the present invention.
[0059] Figure 2 It is a detection schematic diagram of the sensor on the agent in the embodiment of the present invention.
[0060] Figure 3 It is a schematic diagram of the architecture of the deep reinforcement learning model in the embodiment of the present invention.
[0061] Figure 4 It is a schematic diagram of the architecture of the high-level policy module in the embodiment of the present invention;
[0062] Figure 5 It is a functional module block diagram of the cluster pursuit and evasion control system based on hierarchical game deep reinforcement learning provided by the embodiment of the present invention;
[0063] Figure 6 It is a structural block diagram of the electronic device provided by the embodiment of the present invention. Detailed implementation manners
[0064] Next, the technical solutions in the embodiments of the present invention will be described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0065] Please refer to Figure 1 , which is a flowchart of the cluster pursuit and evasion control method based on hierarchical game deep reinforcement learning provided by the embodiment of the present invention. The cluster pursuit and evasion control method based on hierarchical game deep reinforcement learning can be executed by the cluster pursuit and evasion control system based on hierarchical game deep reinforcement learning. The cluster pursuit and evasion control system based on hierarchical game deep reinforcement learning can be implemented by software and / or hardware, and can be configured in an electronic device, and the electronic device can be a computer device. The detailed steps of the cluster pursuit and evasion control method based on hierarchical game deep reinforcement learning are introduced as follows.
[0066] S11. Construct a cluster pursuit and evasion game scenario, where the cluster pursuit and evasion game scenario includes a game environment and multiple agents and obstacles in the game environment. Each agent is equipped with a sensor, and the sensor is used to detect local state information within the detection range of the sensor.
[0067] S12. Construct a cooperative task model for multiple agents based on the partially observable Markov decision process.
[0068] S13. Construct a deep reinforcement learning model for hierarchical games, where the deep reinforcement learning model includes a high-level policy module and a low-level policy module.
[0069] S14. Collect the game information of the agents performing pursuit and evasion simulation according to the cooperative task model in the pursuit and evasion game scenario, and the game information includes local state information.
[0070] S15. Use the game information to train the low-level policy module and the high-level policy module.
[0071] S16. Use the trained high-level policy module to obtain the guidance information of the agents performing the game, and obtain the action information of the agents based on the guidance information through the low-level policy module. Control the agents to perform the pursuit and evasion game based on the action information.
[0072] In this embodiment, a cluster pursuit-evasion game scenario can be constructed based on the game simulation requirements. The game environment includes a certain area for game simulation, and the agents include agents of different camps, such as pursuer agents and evader agents. Among them, the agents can be unmanned aerial vehicle agents, robot agents, autonomous vehicle agents, etc. The obstacles include static obstacles and dynamic obstacles.
[0073] Among them, the numbers of the pursuer agents and the evader agents can be adjusted based on requirements to meet the cluster pursuit-evasion quantity requirements in scenarios such as small-scale, medium-scale, and large-scale scenarios.
[0074] The numbers of the static obstacles and the dynamic obstacles can also be adjusted arbitrarily, and the setting of the system environment obstacles can meet the situations of no obstacles, only static obstacles, only dynamic obstacles, and coexistence of static obstacles and dynamic obstacles.
[0075] Sensors are configured on each agent, and the sensors can be used to detect local state information in its surrounding environment, such as information about obstacles or other agents within the detection range of the sensors.
[0076] In the game simulation, the actions and decisions of each agent can be obtained based on algorithms, so as to be able to perform pursuit-evasion simulation in the game environment. Specifically, the goal of the pursuer agent is to capture as many evader agents as possible, while the evader agent tries to avoid being captured by the pursuer agent. All agents sense the surrounding environment through sensors and take actions according to the sensed information. The existence of static obstacles and dynamic obstacles increases the difficulty of the pursuit-evasion task. In order to increase the exploration and strategy diversity of agent training, the initial positions of the agents and the obstacles can be randomly generated during each subsequent training.
[0077] By optimizing the actions and decisions of each agent through algorithms, it is possible to enable the agents of different camps to have better pursuit or evasion actions.
[0078] When each agent performs pursuit-evasion simulation, its actions, observations, states, etc. need to meet relevant models, which are called cooperative task models in this embodiment. This cooperative task model is constructed based on the partially observable Markov decision process (POMDP) and reformulated as a collective decentralized partially observable Markov decision process (C-Dec POMDP).
[0079] Partial observability here means that each agent can only observe the local state information within the detection range of its sensors, and guides its actions based on this local state information without being able to understand the global information. Therefore, for a particular agent, only the agents or obstacles within a certain range around it can be detected by the sensors of that agent. Specifically, based on the sensors it carries, an agent can determine whether it has detected an observation object in the following ways, where the observation object is any other agent or obstacle:
[0080] Obtain the projection distance between the agent and the observation object in the sensor direction, and obtain the actual distance between the agent and the observation object; detect the actual relative velocity and the projection relative velocity between the agent and the observation object; based on the projection distance, actual distance, actual relative velocity, projection relative velocity, the radius of the observation object, and the sensor vector, judge whether the agent has detected the observation object according to a preset judgment formula.
[0081] In this embodiment, assume that the radius of the pursuer agent P1 is , and the radius of the evader agent is . The radii of the static obstacle and the dynamic obstacle P2 are and respectively. Due to the limitations of the sensor perception ability, an agent can only observe the local state information within the detection range of its sensors. To describe the relationships and observation situations in the environment, take the static obstacle as an example (similarly applicable to other evader agents and dynamic obstacles). As shown in Figure 2 , where is the projection distance vector, representing the projection distance between the agent and the observation object in the sensor direction. is the sensor vector, representing the observation direction and range of the sensor. is the distance vector, representing the actual distance between the agent and the observation object. is the actual relative velocity, representing the velocity difference between the agent and the observation object. is the projection relative velocity, representing the projection relative velocity between the agent and the observation object.
[0082] Based on the above, if the agent and the observation object satisfy the following preset judgment formula, it can be determined that the agent can observe the observation object:
[0083]
[0084]
[0085]
[0086] In this embodiment, a cooperative task model for multiple agents is constructed based on the partially observable Markov decision process. The cooperative task model includes an observation space model, an action space model, a state space model, a reward function, a transition function, etc.
[0087] Specifically, the multi-agent cooperative task is re-modeled as a C-Dec POMDP. The C-Dec POMDP is defined by an 8-tuple as follows:
[0088] , representing the set of
[0089] pursuer agents. Each pursuer agent can be modeled as a 5-tuple . Specifically, is the set of internal states of the pursuer agent at time step , including the positions and velocities that the pursuer agent may be in during the task. is the set of action states of the pursuer agent at time step . is the set of observations of the pursuer agent at time step , representing all the information that the pursuer agent can observe in a given state. is the set of auxiliary information of the pursuer agent at time step , representing the additional information transmitted from the outside. is the policy of the pursuer agent at time step
[0090] , mapping observations to actions during the task.
[0091] is the set of states, including the states of the chaser agents, the evader agent, the dynamic obstacles, and the static obstacles.
[0092] defines the probability of transitioning from state to state under the joint action .
[0093] is the reward function, providing a global reward based on the current state and the joint action.
[0094] is the observation set of the agent.
[0095] is the discount factor applied to future rewards.
[0096] At time step , given the current state each pursuer agent contains a combination of its own acquired local information and auxiliary observation information . The agent, under its local observation and auxiliary observation, selects an action according to the strategy of its own component , and combines with other agents to form a joint action , interacts with the environment and obtains the reward given by the environment , and the environment transfers to the next state according to the state transition function .
[0097] Specifically, the detailed definitions and construction methods of the various elements in the 8-tuple of C-Dec POMDP in the cooperative task model are as follows.
[0098] At time step , the state information in the global state space model can be established based on the following formula:
[0099]
[0100] where represents the state of the pursuer agent at time step , is the number of pursuer agents; represents the state space of the evader agent at time, represents the state of the evader agent at time, is the number of evader agents; represents the initial state space of static obstacles, represents the initial state of static obstacle ; represents the state space of dynamic obstacles at time, represents the state of dynamic obstacle at time, is the number of dynamic obstacles.
[0101] The action space model in the cooperative task model can be constructed in the following way:
[0102] Obtain the acceleration information of each agent on the axis and the acceleration information on the axis at each time step for each agent; construct the action space model based on the acceleration information of each agent on the axis and the acceleration information on the axis.
[0103] Specifically, at time step , the joint action space of the pursuer agent can be expressed as follows:
[0104]
[0105] where represents the action of the pursuer agent at time . represents the acceleration action of the pursuer agent at time along the axis, and represents the acceleration action of the pursuer agent at time along the axis. These actions determine the speed and position changes of the pursuer agent at the next time step.
[0106] Therefore, at time , the joint action space of the multi-agent pursuers is expressed as:
[0107]
[0108] The above state transition function describes the probability of transitioning from the current state to the next state under the joint action , and this joint action includes the actions of all pursuers at time . This function includes the state transitions of pursuers, evaders, static obstacles, and dynamic obstacles.
[0109] For each pursuer, its state transition probability is determined by the above state update equations, and these equations update its state according to the position and speed at time . Similarly, for each evader, its state transition probability is . The state of static obstacles remains unchanged because they do not move. For dynamic obstacles, their state transition probability It is determined by the same motion equation and is used to update its position and velocity.
[0110] At time The state transition function of the multi-agent system can be summarized as:
[0111]
[0112]
[0113] The present invention studies the collective pursuit-evasion problem under partially observable information, where the agents cannot obtain global state information. Therefore, the agents can only obtain observation information through their own sensors. Based on this, the observation space model in the cooperative task model can be constructed in the following way:
[0114] Obtain the internal state information and external state information of each agent at each time step. The external state information is the external local state information detected by the agent's sensor; construct the observation space model based on the internal state information and the external state information.
[0115] Specifically, at time , the pursuer 's observation space is , where represents the internal state of the pursuer at time , represents the external state information observed by the pursuer at time through its sensor, including the distance and velocity information of other objects, and whether there is a collision with the evader or dynamic obstacle.
[0116] External state information is defined as follows:
[0117]
[0118] Where:
[0119] represents the state of other pursuers observed by the pursuer at time .
[0120] represents the distance and velocity of the evader at time .
[0121] represents the distance of the static obstacle at time .
[0122] Indicates at time the distance and speed of dynamic obstacles.
[0123] Indicates at time the distance to the boundary.
[0124] Indicates at time whether the evader is captured.
[0125] Indicates at time whether a collision occurs with an obstacle (static or dynamic).
[0126] Therefore, the global observation space of the pursuer is:
[0127]
[0128] The design of the reward function aims to guide the multi-agent pursuer system to successfully capture the evader. Therefore, the target reward function of the pursuer is defined by the following formula:
[0129]
[0130] Obviously, only the target reward makes the pursuit-evasion game a dynamic and complex problem with sparse rewards. When the pursuer does not detect the evader for a period of time, or detects the evader but does not capture them, the reward remains zero, thus making the policy gradient also zero and unable to achieve policy improvement. Therefore, to address the learning difficulties brought about by sparse rewards, an auxiliary reward function is introduced as a regularization term for the target reward function.
[0131] First, to encourage the pursuer to quickly capture the evader, at time , if the tracker uses sensors to detect the evader , then the detection reward is distributed among all trackers that detect the same evader. In this case, the agent will receive a small reward as follows:
[0132]
[0133] where, is a positive constant representing the reward intensity, is the total positive reward for detecting the evader, is the number of trackers that detect the same evader .
[0134] Secondly, to avoid collisions with static and dynamic obstacles, at time if the pursuer collides with any obstacle, a negative reward will be given as follows:
[0135]
[0136] where is a positive constant representing the punishment intensity, is a positive number greater than zero.
[0137] Finally, to encourage the quick capture of the evader, a time penalty reward function is designed. At time if no evader is captured, a negative reward will be given according to the time used, defined as follows:
[0138]
[0139] where and are positive constants representing the time penalty intensity.
[0140] In summary, the designed reward function and its auxiliary regularization term are defined as:
[0141]
[0142] Based on the above, this embodiment constructs a deep reinforcement learning model for hierarchical game, and this learning model is a regularized auxiliary term deep deterministic policy gradient (HRG-MADDPG) algorithm model based on hierarchical game. Hierarchical game can be represented as a multi-level decision-making process, where each decision level has its own policy set and utility function . The levels influence and optimize each other through hierarchical relationships.
[0143] Please refer to Figure 3 for combination. Hierarchical game is an important framework of the HRG-MADDPG algorithm in this embodiment. It includes a high-level policy module (HLS) and a low-level policy module (LLS). The high-level policy module is responsible for making global decisions, that is, formulating the intelligent agent grouping target allocation and individual target allocation strategies, and providing guidance for the low-level policy module. The low-level policy module then executes specific tasks according to the guidance of the high-level policy module. By combining the high-level policy module and the low-level policy module, the decision-making ability of the multi-agent system in a complex environment is enhanced.
[0144] The high-level strategy module plays a crucial role in coordinating multi-agent systems. HLS uses a high-level policy network (HLNN) to process inputs and generate policy outputs.
[0145] In this embodiment, the game information of the agents that perform pursuit-evasion simulation according to the cooperative task model in the pursuit-evasion game scenario can be stored in the replay buffer pool That is to say, the past experiences of the agents are stored in the replay buffer pool, including the observation information, action information, reward information, next observation information, etc. of the agents. Among them, the observation information includes the local state information of the agents.
[0146] By collecting the game information of the agents from the replay buffer pool, the high-level and low-level policy modules in the deep reinforcement learning model are trained based on the game information.
[0147] Since the high-level policy module is mainly responsible for setting global goals and task allocation, its update frequency is lower than that of the low-level policy module. Therefore, when training the high-level and low-level policy modules based on the game information, it can be achieved in the following way:
[0148] The low-level policy module is trained using the game information; when the number of training iterations of the low-level policy module reaches the update frequency of the high-level policy module, the high-level policy module is trained based on the game information.
[0149] Among them, the training and update process of the low-level policy module includes processes such as calculating action information, storing experiences (the game information of the agents), sampling small sample data, calculating target values, updating the network of the low-level policy module, and performing soft updates. First, the agents generate actions based on local observations and the guidance information provided by the high-level policy module, and then the experience data of these actions are stored in the replay buffer. Then, the LLS randomly samples a small sample of data from the replay buffer for training. Finally, a soft update is performed to ensure that the target network parameters gradually approach the actual network parameters. This process is executed at each time step, enabling the LLS to quickly adapt to environmental changes and optimize the agent behavior.
[0150] The HLS update frequency is expressed as , during the training of the LLS, when the number of training iterations reaches the HLS update frequency , the HLS samples data from the replay buffer and performs training and update of the HLS based on the sampled data and the constructed loss function.
[0151] Through this hierarchical training method, the HLS and LLS work together, which can enhance the overall effectiveness and stability of the algorithm.
[0152] Specifically, the training and update of the high-level policy module are realized based on the game information of the sampling-based agent. The game information includes the observation information, action information, reward function information of the agent at the current time step, and the observation information of the next time step, which can be expressed as . The steps of training the high-level policy module based on the game information can be implemented in the following ways:
[0153] Calculate the current expected value according to the observation information and guidance information at the current time step; calculate the target expected value according to the observation information and guidance information of the next time step; and conduct guidance training on the high-level policy module based on the loss function constructed based on the current expected value and the target expected value.
[0154] Please refer to Figure 4 specifically. In the high-level policy module, the agent accepts the observation state of the current cluster agent sensor , action information , the reward of the low-level policy module , and the observation of the next moment as inputs at each time step. The main output of the HLS is the guidance information of each agent . The guidance information includes target ID assignment, cooperation coefficient, and priority weight distribution.
[0155] The target ID assignment is used to determine the ID of the evader agent that each pursuer agent corresponds to pursue. When the agent observes the evader sequence , where and , the high-level policy will assign a corresponding pursuer to each evader .
[0156] The cooperation coefficient is used to determine the degree of cooperation between each pursuer agent and other pursuer agents. Each pursuer will obtain a cooperation coefficient , which represents the degree of closeness of the pursuer cooperating with other pursuers when performing tasks. The cooperation coefficient ranges from 0 to 1, and the higher the value, the higher the degree of cooperation. For example, as Figure 4 shows, when two agents observe the same evader at the same time, their cooperation coefficient will be higher; conversely, when the agent does not observe any evaders, or only one pursuer observes the evader, the cooperation coefficient will be lower.
[0157] The tasks of the pursuer agents include pursuit tasks and evasion tasks. The priority weight distribution includes the weights of the pursuit tasks and the evasion tasks for each pursuer agent. That is, it is used to determine the priority when each pursuer executes various tasks. Each pursuer is assigned a priority weight , which represents the priority of the pursuer in the current task. As Figure 4 shown, the priority weight can determine whether the primary task of the pursuer is to pursue the assigned evader or to avoid obstacles.
[0158] Specifically, when training HLS, at time step , HLS randomly samples a mini-batch of data from the replay buffer . The replay buffer stores the past experiences of the agents, including observation information, action information, reward function information, and next observation information. Sampling a mini-batch of data for training enables the algorithm to learn from diverse past interactions, thereby improving the robustness and generalization ability of the policy. At this time, the calculation formula for the guidance information of the agent is:
[0159]
[0160] Next, the algorithm calculates the current expected value , denoted as , which represents the expected value that the agent can achieve starting from the current state and following the calculated guidance information . This value is a key component for evaluating the effectiveness of the policy, and it uses the current observation and the guidance information :
[0161]
[0162] To optimize the policy, HLS uses the Temporal Difference (TD) error algorithm. It selects a new sample from the mini-batch of data, which includes the observation information at the next time step, and calculates the guidance information at the next time step. The guidance information at the next time step is calculated using the same HLNN but applied to the observation at the next moment:
[0163]
[0164] After calculating the current policy and the target policy, the algorithm calculates the target expected value , which represents the expected value obtained from the future state. This target The value is obtained by adding the immediate reward received at time from to the discounted value of the next state as follows: where
[0165]
[0166] is the discount factor. The loss function
[0167] is calculated as the mean squared error (MSE) between the current expected value and the target expected value . Additionally, a regularization term is included to penalize large weights in the policy parameters to prevent overfitting. The regularization term is controlled by the hyperparameter which determines the strength of the penalty for large weights. The weights represent the individual independent parameters in the policy parameters . The loss function is calculated as follows: Finally, the HLS parameters
[0168]
[0169] are updated by performing gradient descent on the loss function :
[0170]
[0171] This update adjusts the parameters to minimize the loss, thus effectively improving the HLS over time.
[0172] This iterative process continues until the policy parameters converge to the optimal solution or a predetermined number of iterations is reached to complete the training of the high-level policy module. The final output of the HLS is the high-level guidance information which is passed to the LLS to guide the agent's actions at a finer-grained level.
[0173] The core of the low-level policy module is to use a classic multi-agent deep reinforcement learning algorithm MADDPG, which consists of two phases: the offline training phase and the online execution phase. During the offline training phase, the LLS includes the training of the shared value network and the shared policy network.
[0174] During the online execution phase, based on the shared policy network, a policy improvement method is adopted to enhance the decision-making of each agent. At time , the pursuer Execute specific tasks according to the guidance information provided by HLS and the current observation information. Considering the large scale and agent isomorphism in the cluster pursuit-evasion problem, in this embodiment, LLS adopts a method of policy sharing to simplify the network, that is, all pursuers share a shared policy network with parameters of and a shared value network with parameters of . In addition, the shared policy network also has its corresponding target policy network with parameters , and the shared value network also has its corresponding target value network with parameters . The use of the shared network greatly reduces the number of networks and also has lower computational complexity.
[0175] Based on this, when training the low-level policy module using game information, it can be specifically implemented in the following ways:
[0176] For each agent, calculate the target value based on the agent's game information, and update the shared value network by minimizing the loss function constructed by the current value and the target value;
[0177] Take the observation information in the game information and the guidance information provided by the high-level policy module as the input of the shared policy network, use the shared policy network to output the corresponding action information, and update the shared policy network in a gradient ascent manner based on the cumulative reward under the action information.
[0178] In this embodiment, for the shared policy network, at time , the environmental state observed by the sensor of the pursuer is represented as . For each pursuer , the shared policy network takes the observation at the current time step as the input and generates the corresponding action at the next time step. The policy generates the corresponding action for the pursuer based on the observation state and the guidance information :
[0179]
[0180] In the shared value network, the shared value network adopts the joint observation value and the joint action to calculate the pursuer's value. Denote at time the value of the pursuer calculated through as is , then:
[0181]
[0182] Similarly, the value output based on the target network can be defined as:
[0183]
[0184] The goal of the pursuer is to maximize its expected cumulative discounted return , where represents the cumulative return over time, defined as:
[0185]
[0186] where is the discount factor, is the pursuer at time step obtained with the auxiliary term.
[0187] The cumulative return is obtained from the reward function with the auxiliary reward regularization term obtained by the agent within the cumulative time steps. The auxiliary reward regularization term is constructed based on whether the pursuer agent captures the evader agent, whether the pursuer agent detects the evader agent, whether the pursuer agent collides with the obstacle, and the number of evader agents captured by the pursuer agent.
[0188] Specifically, is calculated as follows:
[0189]
[0190] The definition and calculation method of each auxiliary reward regularization term can be referred to the relevant elaboration above in this embodiment.
[0191] Then, the shared policy network of the pursuer is updated through gradient ascent, and its gradient expression is:
[0192]
[0193] Here, the experience replay buffer contains the tuple , recording the experiences of all agents. The centralized action-value function is used to train the shared value network . The loss function consists of the current value and the target value , and is constructed according to the following formula:
[0194]
[0195] where and are defined as above.
[0196] Based on the above method, the training of the high-level policy module and the low-level policy module in the deep reinforcement learning model is completed.
[0197] The trained deep reinforcement learning model is applied to different test environments for pursuit-evasion simulation test verification. Specifically, the trained high-level policy module can obtain the guidance information of the agents for the game. Through the low-level policy module, the action information of the agents is obtained based on the guidance information, and the agents are controlled to conduct the pursuit-evasion game based on the action information.
[0198] In this embodiment, a partially observable swarm pursuit-evasion control scheme based on hierarchical game deep reinforcement learning is proposed. By introducing a hierarchical game control structure and a regularization-assisted reward mechanism, each agent can efficiently cooperate in a partially observable environment, and the cooperation efficiency and adaptability of the swarm agents in the partially observable environment are improved.
[0199] Through the hierarchical game structure, the system can generate cooperative strategies at the global and local levels to maximize the swarm cooperation efficiency. Specifically, the high-level policy module is responsible for the global goal assignment of tasks, generates task assignment strategies based on the current observations of the agents, including target assignment, cooperation coefficients, and priority weights. This strategy is generated by a high-level neural network (HLNN) and is used to guide the cooperative behavior of each agent in a partially observable environment. The low-level policy module (LLS) adopts the multi-agent deep deterministic policy gradient (MADDPG) algorithm, and combines the regularization-assisted reward term to optimize the policy convergence and learning efficiency of each agent during the centralized training phase; during the distributed execution phase, each agent independently executes the strategy of the high-level task assignment based on the local observation information to conduct specific pursuit behaviors.
[0200] In addition, a regularization-assisted reward mechanism is introduced, and auxiliary reward terms such as detection rewards, collision penalties, time penalties, and capture rewards are introduced to ensure the cooperation efficiency among agents, reduce ineffective behaviors, and improve the overall policy convergence speed. It solves the problem of difficult policy optimization caused by sparse rewards in a partially observable environment and ensures the adaptability and robustness of the system in complex environments.
[0201] In summary, this embodiment provides a reliable solution for swarm pursuit-evasion, which can be trained in a single environment and verified in multiple different test environments, demonstrating strong robustness and generalization ability. This solution not only realizes agent cooperation in task allocation, but also adds a regularization-assisted reward mechanism in the pursuit-execution stage to optimize the convergence and stability of the strategy. The hierarchical game control architecture improves the pursuit-evasion effect of multi-agent swarms in partially observable environments, overcomes the information bottleneck problem of traditional centralized control in complex environments, and makes the system more robust and real-time. The proposed framework has adaptability, scalability, and fast responsiveness to complex dynamic environments in various practical applications.
[0202] Based on the same inventive concept, please refer to Figure 5 , this embodiment of the present invention also provides a schematic diagram of the functional modules of a swarm pursuit-evasion control system based on hierarchical game deep reinforcement learning. This embodiment can divide the functional modules of the swarm pursuit-evasion control system based on hierarchical game deep reinforcement learning according to the above method embodiment. For example, each functional module can be corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in this embodiment of the present invention is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0203] For example, in the case of dividing each functional module corresponding to each function, Figure 5 the shown swarm pursuit-evasion control system based on hierarchical game deep reinforcement learning is only a schematic diagram of a device. The swarm pursuit-evasion control system based on hierarchical game deep reinforcement learning may include a construction module, a collection module, a training module, and a guidance module. The functions of each functional module of the swarm pursuit-evasion control system based on hierarchical game deep reinforcement learning will be elaborated in detail below.
[0204] The construction module is used to construct a swarm pursuit-evasion game scenario, where the swarm pursuit-evasion game scenario includes a game environment and multiple agents and obstacles in the game environment. Each agent is equipped with a sensor, and the sensor is used to detect local state information within the sensor range;
[0205] The construction module is also used to construct a cooperative task model for multi-agents based on the partially observable Markov decision process;
[0206] The construction module is also used to construct a deep reinforcement learning model for hierarchical games, and the deep reinforcement learning model includes a high-level policy module and a low-level policy module;
[0207] A collection module, configured to collect the game information of an agent that performs pursuit-evasion simulation according to the cooperation task model in the pursuit-evasion game scenario, where the game information includes local state information;
[0208] A training module, configured to train the low-level policy module and the high-level policy module by using the game information;
[0209] A guidance module, configured to obtain guidance information for an agent that performs a game by using the trained high-level policy module, and obtain action information of the agent based on the guidance information through the low-level policy module, and control the agent to perform a pursuit-evasion game based on the action information.
[0210] It can be understood that the above construction module, collection module, training module, and guidance module can be used to execute S11 to S16 above. For the detailed implementation manners of the construction module, collection module, training module, and guidance module, reference can be made to the content related to S11 to S16 above.
[0211] The cluster pursuit-evasion control system based on hierarchical game deep reinforcement learning provided in this embodiment has the same, similar, or corresponding technical features as the cluster pursuit-evasion control method based on hierarchical game deep reinforcement learning in the above embodiment, and has the same technical effects. For the relevant content of this control system, reference can be made to the description related to the above control method, and this embodiment will not be elaborated here.
[0212] Please refer to Figure 6 , which is a structural block diagram of an electronic device provided in an embodiment of the present invention. The electronic device may be a computer device or the like. The electronic device includes a memory, a processor, and a communication module. Each element of the memory, the processor, and the communication module is directly or indirectly electrically connected to each other to realize data transmission or interaction. For example, these elements may be electrically connected to each other through one or more communication buses or signal lines.
[0213] Among them, the memory is used to store computer programs or data. The memory may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM), etc.
[0214] The processor is used to read / write the data or programs stored in the memory and execute the cluster pursuit-evasion control method based on hierarchical game deep reinforcement learning provided by any embodiment of the present invention.
[0215] The communication module is used to establish a communication connection between the electronic device and other communication terminals through the network and is used to send and receive data through the network.
[0216] It should be understood that Figure 6 The structure shown is only a schematic diagram of the structure of the electronic device, and the electronic device may also include more or fewer components than those shown in Figure 6 and may have a different configuration from that shown in Figure 6 .
[0217] Furthermore, an embodiment of the present invention also provides a computer-readable storage medium, which stores machine-executable instructions that, when executed, implement the cluster pursuit-evasion control method based on hierarchical game deep reinforcement learning provided by the above embodiment.
[0218] Specifically, the computer-readable storage medium can be a general storage medium, such as a removable disk, a hard disk, etc. When the computer program on the computer-readable storage medium runs, it can execute the above-mentioned cluster pursuit-evasion control method based on hierarchical game deep reinforcement learning. Regarding the process involved when the machine-executable instructions in the computer-readable storage medium run, reference can be made to the relevant descriptions in the above method embodiment and will not be elaborated here.
[0219] In summary, the cluster pursuit-evasion control method and system provided by the embodiment of the present invention model the multi-agent cooperation task as a collective decentralized partially observable Markov decision process through a hierarchical game structure and propose a hierarchical game multi-agent deep deterministic policy gradient algorithm. The algorithm includes a high-level policy and a low-level policy. The high-level policy is responsible for target allocation and task coordination, and the low-level policy optimizes specific action decisions through centralized training and distributed execution, and combines a regularization auxiliary function to improve the convergence speed and execution efficiency of the policy. This solution can effectively improve the cooperation efficiency and response ability of the multi-agent system in a complex environment.
[0220] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0221] In addition, the units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0222] Furthermore, in each embodiment of the present invention, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0223] It should be noted that if the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0224] The above are only the embodiments of the present invention and are not used to limit the protection scope of the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A cluster pursuit-evasion control method based on hierarchical game deep reinforcement learning, characterized in that The method includes: Constructing a cluster pursuit-evasion game scenario, where the cluster pursuit-evasion game scenario includes a game environment and multiple agents and obstacles in the game environment. Each agent is equipped with a sensor, and the sensor is used to detect local state information within the detection range of the sensor; Constructing a cooperative task model for multiple agents based on the partially observable Markov decision process; Constructing a deep reinforcement learning model for hierarchical games, where the deep reinforcement learning model includes a high-level policy module and a low-level policy module; Collecting the game information of the agents performing pursuit-evasion simulation according to the cooperative task model in the pursuit-evasion game scenario, where the game information includes local state information; Training the low-level policy module and the high-level policy module using the game information; Obtaining the guidance information of the agents for the game using the trained high-level policy module, and obtaining the action information of the agents based on the guidance information through the low-level policy module. Controlling the agents to perform the pursuit-evasion game based on the action information; The cooperative task model includes an observation space model and an action space model. The steps of constructing the observation space model include: Obtaining the internal state information and external state information of each agent at each time step, where the external state information is the external local state information detected by the sensor of the agent; constructing the observation space model based on the internal state information and the external state information; The steps of constructing the action space model include: Obtaining the acceleration information of each agent on the x-axis and the acceleration information on the y-axis at each time step; constructing the action space model based on the acceleration information of each agent on the x-axis and the acceleration information on the y-axis; The steps of training the low-level policy module and the high-level policy module using the game information include: Training the low-level policy module using the game information; When the number of training iterations of the low-level policy module reaches the update frequency of the high-level policy module, training the high-level policy module based on the game information.
2. The cluster pursuit and escape control method based on hierarchical game deep reinforcement learning according to claim 1, wherein The low-level policy module includes a shared value network and a shared policy network; The steps of training the low-level policy module using the game information include: For each agent, calculating the target value based on the game information of the agent, and updating the shared value network by minimizing the loss function constructed by the current value and the target value; Taking the observation information in the game information and the guidance information provided by the high-level policy module as the input of the shared policy network, using the shared policy network to output the corresponding action information, and updating the shared policy network in a gradient ascent manner based on the cumulative reward under the action information; 3. The cluster pursuit and evasion control method based on hierarchical game deep reinforcement learning according to claim 2, characterized in that The cumulative reward is obtained from the reward function with an auxiliary reward regularization term obtained by the agent within the cumulative time steps; The auxiliary reward regularization term is constructed based on whether the pursuer agent captures the evader agent, whether the pursuer agent detects the evader agent, whether the pursuer agent collides with an obstacle, and the number of evader agents captured by the pursuer agent.
4. The cluster pursuit and evasion control method based on hierarchical game deep reinforcement learning according to claim 1, wherein The game information includes the observation information, action information, reward function information of the agent at the current time step, and the observation information at the next time step; The step of training the high-level policy module based on the game information includes: Inputting the observation information, action information, reward function information of the agent at the current time step, and the observation information at the next time step included in the game information into the high-level policy module to obtain the guidance information at the current time step and the guidance information at the next time step; Calculating the current expected value according to the observation information and guidance information at the current time step; Calculating the target expected value according to the observation information and guidance information at the next time step; Guiding and training the high-level policy module based on the loss function constructed based on the current expected value and the target expected value.
5. The cluster pursuit and evasion control method based on hierarchical game deep reinforcement learning according to claim 1, characterized in that The agent includes a pursuer agent and an evader agent, and the tasks of the pursuer agent include a pursuit task and an avoidance task; The guidance information includes target ID assignment, cooperation coefficient, and priority weight distribution; The target ID assignment is used to determine the ID of the evader agent that each pursuer agent pursues; The cooperation coefficient is used to determine the degree of cooperation between each pursuer agent and other pursuer agents; The priority weight distribution includes the weights of the pursuit task and the avoidance task of each pursuer agent.
6. The cluster pursuit and evasion control method based on hierarchical game deep reinforcement learning according to claim 1, characterized in that Each agent determines whether it detects an observed object through the following method based on the carried sensor, and the observed object is any other agent or an obstacle: Obtaining the projected distance between the agent and the observed object in the sensor direction, and obtaining the actual distance between the agent and the observed object; Detecting the actual relative speed and the projected relative speed between the agent and the observed object; Judging whether the agent detects the observed object based on the projected distance, actual distance, actual relative speed, projected relative speed, the radius of the observed object, and the sensor vector according to a preset judgment formula.
7. A cluster pursuit and evasion control system based on hierarchical game deep reinforcement learning, characterized in that A system for implementing the multi-agent cluster pursuit and evasion control method based on hierarchical game deep reinforcement learning according to any one of claims 1-6, the system includes: A construction module for constructing a multi-agent cluster pursuit and evasion game scenario, the multi-agent cluster pursuit and evasion game scenario includes a game environment and multiple agents and obstacles in the game environment, and each agent carries a sensor, and the sensor is used to detect local state information within the detection range of the sensor; The construction module is further configured to construct a cooperative task model of multi-agents based on a partially observable Markov decision process; The construction module is further configured to construct a deep reinforcement learning model of a hierarchical game, and the deep reinforcement learning model includes a high-level policy module and a low-level policy module; A collection module, configured to collect the game information of an agent that performs pursuit-evasion simulation according to the cooperation task model in the pursuit-evasion game scenario, where the game information includes local state information; A training module, configured to train the low-level policy module and the high-level policy module by using the game information; A guidance module, configured to obtain the guidance information of the agent performing the game by using the trained high-level policy module, and obtain the action information of the agent based on the guidance information through the low-level policy module, and control the agent to perform the pursuit-evasion game based on the action information.
Citation Information
Patent Citations
Vacuum flat plate heat collector and manufacturing method and equipment thereof
CN111947323A
Protection mechanism for intelligent self-service ticket taking machine
CN112233339A
Electric motor driving device
CN113474985A
Image rendering method, device and equipment and readable storage medium
CN114253649A
Optical fiber radio frequency replication jammer
CN114422077A