Cluster pursuit control method and system based on hierarchical game deep reinforcement learning

By adopting a deep reinforcement learning method based on hierarchical game in the multi-agent system, a cooperative task model is built and the strategy is optimized, the limitations of multi-agent pursuit and escape in some observable environments are solved, and the collaboration efficiency and response capabilities are improved.

CN119960489AActive Publication Date: 2025-05-09TONGJI UNIV

Patent Information

Application Number
CN202510450379.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-09
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing multi-agent pursuit method has significant limitations when facing some observable dynamic environments, including limitations in information acquisition, inefficient collaboration and high computational complexity.

Method used

The deep reinforcement learning method based on hierarchical game is adopted to build a cooperative task model for multiple agents, and the pursuit strategy of agents is optimized through the combination of high-level policy modules and low-level policy modules.

Benefits of technology

It improves the collaboration efficiency and coping capabilities of multi-agent systems in complex environments, and can realize the optimal pursuit strategy in the case of incomplete information, reducing computing complexity and communication delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960489A_ABST
    Figure CN119960489A_ABST
Patent Text Reader

Abstract

The invention provides a cluster pursuit control method and system based on hierarchical game deep reinforcement learning, a multi-agent cooperation task is modeled into a collective decentration partial observable Markov decision process through a hierarchical game structure, and a hierarchical game multi-agent depth deterministic strategy gradient algorithm model is provided. The algorithm model comprises a high-level strategy module and a low-level strategy module, the high-level strategy module is responsible for target allocation and task coordination, and the low-level strategy module optimizes specific action decisions through centralized training and distributed execution. According to the scheme, the cooperation efficiency and the coping capacity of the multi-agent system in a complex environment can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of reinforcement learning, and in particular to a cluster pursuit control method and system based on hierarchical game deep reinforcement learning. Background Art

[0002] In modern intelligent applications such as autonomous driving, drone formations, intelligent security, and military reconnaissance, the problem of multi-agent pursuit has gradually become a hot topic in academia and industry. The core goal of the multi-agent pursuit task is to enable a group of "pursuers" to chase and capture the "fugitive" through collaboration. This type of task has the characteristics of high real-time, complex environment, and frequent interactions, and usually requires efficient collaboration and decision-making among multiple agents in a dynamic environment. Traditional pursuit methods mostly rely on global observation information and centralized control strategies, which have many limitations in actual application scenarios. Especially in partially observable environments, the lack of information acquisition and uncertainty have become the main challenges in the pursuit task of multi-agent systems.

[0003] Existing multi-agent pursuit methods can be summarized into the following categories: traditional analytical methods, artificial intelligence-based biologically inspired algorithms, and multi-agent strategies based on deep reinforcement learning. These methods have certain practicality in certain specific environments, but their limitations gradually emerge when facing some observable environments and large-scale intelligent agent scenarios.

[0004] Traditional analytical methods are based on differential game theory or geometric theory, and solve the pursuit problem by constructing precise mathematical models. Patent CN115891002A analyzes the pursuit game between multiple pursuers and fugitives based on a mathematical model, and derives the optimal strategy of the pursuers through the model. Patent CN111947323A establishes a multi-agent confrontation task model based on differential equations, and obtains the pursuit strategy by analyzing and solving the motion model of the agent. Although traditional analytical methods can play a certain role in fixed and simple environments, with the complexity of the scene and the increase in the number of agents, traditional analytical methods become infeasible due to high computational complexity. Especially in partially observable dynamic environments, traditional methods cannot calculate the optimal decision in real time and are difficult to adapt to environmental uncertainty.

[0005] Bio-inspired algorithms achieve collaboration and adaptive behavior among multiple agents by simulating the hunting or group behavior of animals. Patent CN114422077A implements the task allocation of pursuit and escape tasks of agents based on the hunting behavior of wolf packs, and designs pursuit and escape strategies by simulating the collaboration mechanism in wolf packs. Patent CN112233339A proposes an ant colony optimization algorithm for target search and path planning of drones in uncertain environments. These bio-inspired algorithms have certain advantages in swarm intelligence and collaboration, but their computational complexity is also high, and they are sensitive to parameters. There is a large computational burden in dynamic environments, and they are prone to falling into local optimality, resulting in poor applicability in dynamic pursuit and escape tasks.

[0006] In recent years, multi-agent deep reinforcement learning has shown great potential in solving multi-agent collaboration tasks. Patent CN114253649A adopts a multi-agent deep reinforcement learning method to optimize the multi-agent collaborative pursuit strategy through a deep neural network, which is suitable for multi-agent collaboration in a high-dimensional environment. This patent improves the strategy learning effect of the agent through centralized training and distributed execution. Patent CN113474985A uses reinforcement learning technology to achieve robot target tracking in a dynamic environment and improves strategy performance through a self-game mechanism. Although these deep learning-based methods have achieved good results in multi-agent systems, they generally assume that the environment is globally observable and ignore the uncertainty caused by the limitations of the agent's perspective in practical applications. In a partially observable environment, the agent can only obtain local observation information. This incomplete information situation significantly increases the difficulty of strategy generation and optimization, resulting in poor performance of traditional multi-agent deep reinforcement learning methods in partially observable dynamic environments.

[0007] In partially observable environments, multi-agent pursuit tasks face the following technical challenges: Limitations of information acquisition: In some observable environments, the agent can only obtain limited local information through sensors, resulting in the lack of global observation information. This incomplete information will affect the decision-making quality of the agent, making it difficult for the agent to make effective judgments when performing pursuit tasks.

[0008] Inefficient collaboration: The effectiveness of multi-agent collaboration depends on real-time communication and information sharing between agents. However, in partially observable environments, traditional centralized control systems are difficult to meet collaboration requirements. When there is a lack of effective collaboration strategies between multiple agents, it may lead to waste of resources and low capture efficiency.

[0009] High computational complexity: In a dynamic environment, multiple agents need to constantly update and adjust strategies, and traditional centralized control methods have significant disadvantages in computational complexity. In addition, when the number of agents is large, centralized control will increase communication delays, resulting in a decrease in the real-time performance of the system.

[0010] In summary, existing multi-agent pursuit methods have significant limitations when facing partially observable dynamic environments. How to design a distributed, flexible and efficient multi-agent collaboration strategy so that each agent can achieve the optimal pursuit strategy under incomplete information is an urgent problem to be solved. Summary of the invention

[0011] The purpose of the embodiments of the present invention is to provide a cluster pursuit control method and system based on hierarchical game deep reinforcement learning, so as to improve the collaboration efficiency and response capability of the multi-agent system in a complex environment.

[0012] In a first aspect, the present invention provides a cluster pursuit control method based on hierarchical game deep reinforcement learning, the method comprising: Constructing a cluster pursuit game scenario, wherein the cluster pursuit game scenario includes a game environment and a plurality of intelligent agents and obstacles in the game environment, each of the intelligent agents carries a sensor, and the sensor is used to detect local state information within a detection range of the sensor; Construct a multi-agent cooperative task model based on partially observable Markov decision process; Constructing a deep reinforcement learning model of hierarchical game, wherein the deep reinforcement learning model includes a high-level strategy module and a low-level strategy module; Collecting game information of an intelligent agent that performs pursuit and escape simulation according to the cooperative task model in the pursuit and escape game scenario, wherein the game information includes local state information; Using the game information to train the low-level strategy module and the high-level strategy module; The trained high-level strategy module is used to obtain guidance information for the intelligent agent playing the game, and the action information of the intelligent agent is obtained based on the guidance information through the low-level strategy module, and the intelligent agent is controlled to play the pursuit-escape game based on the action information.

[0013] In an optional implementation, the step of training the low-level strategy module and the high-level strategy module using the game information includes: Using the game information to train the low-level strategy module; When the number of training iterations of the low-level policy module reaches the update frequency of the high-level policy module, the high-level policy module is trained based on the game information.

[0014] In an optional embodiment, the low-level policy module includes a shared value network and a shared policy network; The step of training the low-level strategy module using the game information includes: For each agent, a target value is calculated based on the game information of the agent, and the shared value network is updated by minimizing a loss function constructed by a current value and a target value; The observation information in the game information and the guidance information provided by the high-level strategy module are used as inputs of the shared strategy network, the shared strategy network is used to output corresponding action information, and the shared strategy network is updated using a gradient ascent method based on the cumulative rewards under the action information.

[0015] In an optional embodiment, the cumulative reward is obtained by a reward function with an auxiliary reward regularization term obtained by the agent in the cumulative time step; The auxiliary reward regularization term is constructed based on detecting whether the pursuer agent captures the evader agent, whether the pursuer agent detects the evader agent, whether the pursuer agent collides with an obstacle, and the number of evader agents captured by the pursuer agent.

[0016] In an optional embodiment, the game information includes observation information, action information, reward function information of the agent at the current time step and observation information of the next time step; The step of training the high-level strategy module based on the game information includes: Inputting the observation information, action information, reward function information of the agent at the current time step and the observation information of the next time step included in the game information into the high-level strategy module to obtain the guidance information of the current time step and the guidance information of the next time step; The current expected value is calculated based on the observation information and guidance information of the current time step; The target expected value is calculated based on the observation information and guidance information of the next time step; The high-level strategy module is trained based on a loss function constructed based on the current expected value and the target expected value.

[0017] In an optional embodiment, the agent includes a pursuer agent and an evader agent, and the tasks of the pursuer agent include a pursuit task and an evasion task; The guidance information includes target ID allocation, cooperation coefficient and priority weight distribution; The target ID allocation is used to determine the ID of the evader agent that each pursuer agent is pursuing; The cooperation coefficient is used to determine the degree of cooperation between each pursuer agent and other pursuer agents; The priority weight distribution includes the weight of the pursuit task and the weight of the avoidance task of each pursuer agent.

[0018] In an optional embodiment, each of the intelligent agents determines whether an observation object is detected based on the sensors carried by the intelligent agents in the following manner, where the observation object is any other intelligent agent or an obstacle: Obtaining a projection distance between the agent and the observed object in the sensor direction, and obtaining an actual distance between the agent and the observed object; Detecting an actual relative speed and a projected relative speed between the intelligent agent and the observed object; Based on the projection distance, actual distance, actual relative speed, projection relative speed, radius of the observed object and sensor vector, it is determined whether the intelligent agent detects the observed object according to a preset judgment formula.

[0019] In an optional implementation, the cooperative task model includes an observation space model, and the step of constructing the observation space model includes: Obtaining internal state information and external state information of each of the agents in each time step, wherein the external state information is external local state information detected by the sensor of the agent; An observation space model is constructed based on the internal state information and the external state information.

[0020] In an optional implementation, the cooperative task model includes an action space model, and the step of constructing the action space model includes: Obtaining acceleration information of each of the agents on the x-axis and on the y-axis at each time step; An action space model is constructed based on the acceleration information of each intelligent agent on the x-axis and the acceleration information on the y-axis.

[0021] In a second aspect, the present invention provides a cluster pursuit control system based on hierarchical game deep reinforcement learning, the system comprising: A construction module is used to construct a cluster pursuit game scenario, wherein the cluster pursuit game scenario includes a game environment and multiple intelligent agents and obstacles in the game environment, each of the intelligent agents carries a sensor, and the sensor is used to detect local state information within the sensor detection range; The building block is also used to construct multi-agent cooperative task models based on partially observable Markov decision processes; The construction module is also used to construct a deep reinforcement learning model of hierarchical game, wherein the deep reinforcement learning model includes a high-level strategy module and a low-level strategy module; A collection module, used to collect game information of an intelligent agent that performs pursuit and escape simulation according to the cooperative task model in the pursuit and escape game scenario, wherein the game information includes local state information; A training module, used for training the low-level strategy module and the high-level strategy module using the game information; The guidance module is used to use the trained high-level strategy module to obtain guidance information for the intelligent agent playing the game, and obtain the action information of the intelligent agent based on the guidance information through the low-level strategy module, and control the intelligent agent to play the pursuit and escape game based on the action information. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments of the present invention are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 A flow chart of a cluster pursuit control method based on hierarchical game deep reinforcement learning provided by an embodiment of the present invention; Figure 2 Schematic diagram of detection by sensors on an intelligent body in an embodiment of the present invention; Figure 3 Schematic diagram of the architecture of a deep reinforcement learning model in an embodiment of the present invention; Figure 4 Schematic diagram of the architecture of a high-level policy module in an embodiment of the present invention; Figure 5 A functional module block diagram of a cluster pursuit control system based on hierarchical game deep reinforcement learning provided by an embodiment of the present invention; Figure 6 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The technical solutions in the embodiments of the present invention will be described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0025] See also Figure 1 , which is a flow chart of a cluster pursuit control method based on hierarchical game deep reinforcement learning provided by an embodiment of the present invention. The cluster pursuit control method based on hierarchical game deep reinforcement learning can be executed by a cluster pursuit control system based on hierarchical game deep reinforcement learning. The cluster pursuit control system based on hierarchical game deep reinforcement learning can be implemented by software and / or hardware and can be configured in an electronic device, which can be a computer device. The detailed steps of the cluster pursuit control method based on hierarchical game deep reinforcement learning are introduced as follows.

[0026] S11, constructing a cluster pursuit game scenario, wherein the cluster pursuit game scenario includes a game environment and multiple intelligent agents and obstacles in the game environment, each of the intelligent agents carries a sensor, and the sensor is used to detect local state information within the sensor detection range.

[0027] S12, constructing a multi-agent cooperative task model based on partially observable Markov decision process.

[0028] S13, constructing a deep reinforcement learning model of hierarchical game, wherein the deep reinforcement learning model includes a high-level strategy module and a low-level strategy module.

[0029] S14, collecting game information of an intelligent agent that performs pursuit and escape simulation according to the cooperative task model in the pursuit and escape game scenario, wherein the game information includes local state information.

[0030] S15, using the game information to train the low-level strategy module and the high-level strategy module.

[0031] S16, using the trained high-level strategy module to obtain guidance information for the intelligent agent playing the game, and using the low-level strategy module to obtain action information of the intelligent agent based on the guidance information, and controlling the intelligent agent to play the pursuit and escape game based on the action information.

[0032] In this embodiment, a cluster pursuit game scenario can be constructed based on the game simulation requirements. The game environment includes a certain area constructed for game simulation. The intelligent agents include agents of different camps, such as pursuer agents and evader agents. Among them, the intelligent agents can be drone agents, robot agents, autonomous vehicle agents, etc. Obstacles include static obstacles and dynamic obstacles.

[0033] Among them, the number of pursuer agents and evader agents can be adjusted based on demand to meet the requirements of cluster pursuit in small-scale, medium-scale, and large-scale scenarios.

[0034] The number of static obstacles and dynamic obstacles can also be adjusted arbitrarily. The setting of system environment obstacles can meet the situations of no obstacles, only static obstacles, only dynamic obstacles, coexistence of static obstacles and dynamic obstacles, etc.

[0035] Each intelligent agent is equipped with sensors, which can be used to detect local state information in its surrounding environment, such as obstacles or information about other intelligent agents within the sensor's detection range.

[0036] In the game simulation, the actions and decisions of each agent can be obtained based on the algorithm, so that pursuit simulation can be carried out in the game environment. Specifically, the goal of the pursuer agent is to capture as many evader agents as possible, while the evader agent tries to avoid being captured by the pursuer agent. All agents perceive the surrounding environment through sensors and take actions based on the perceived information. The presence of static and dynamic obstacles increases the difficulty of the pursuit task. In order to increase the exploratory nature and strategy diversity of agent training, the initial positions of the agents and obstacles can be randomly generated in each subsequent training.

[0037] By optimizing the actions and decisions of each agent through algorithms, agents in different camps can all have better pursuit or escape actions.

[0038] The actions, observations, and states of each agent in the pursuit simulation must satisfy the relevant model, which is called the cooperative task model in this embodiment. The cooperative task model is built based on the partially observable Markov decision process (POMDP) ​​and is restated as the collective decentralized partially observable Markov decision process (C-Dec POMDP).

[0039] Partial observability here means that each agent can only observe local state information within the sensor detection range, guide actions based on local state information, and cannot understand global information. Therefore, for a certain agent, only agents or obstacles within a certain range around it can be detected by the agent's sensor. Specifically, the agent can determine whether the observed object is detected based on the sensors it carries in the following ways, and the observed object is any other agent or obstacle: Obtain the projected distance between the intelligent agent and the observed object in the sensor direction, and obtain the actual distance between the intelligent agent and the observed object; detect the actual relative speed and the projected relative speed between the intelligent agent and the observed object; and determine whether the intelligent agent detects the observed object according to a preset judgment formula based on the projected distance, the actual distance, the actual relative speed, the projected relative speed, the radius of the observed object and the sensor vector.

[0040] In this embodiment, it is assumed that the radius of the pursuer agent P1 is , the radius of the evader agent is , the radii of the static obstacle and the dynamic obstacle P2 are and . Under the limitation of sensor perception, the agent can only observe the local state information within the detection range of its sensor. To describe the relationship and observation in the environment, static obstacles are taken as an example (the same applies to other evader agents and dynamic obstacles). Figure 2 As shown in, is the projected distance vector, which represents the projected distance between the agent and the observed object in the sensor direction. is the sensor vector, which indicates the observation direction and range of the sensor. is a distance vector, which represents the actual distance between the agent and the observed object. is the actual relative speed, which represents the speed difference between the agent and the observed object. is the projected relative velocity, which represents the projected relative velocity between the agent and the observed object.

[0041] On the basis of the above, if the following preset judgment formula is satisfied between the agent and the observed object, it can be determined that the agent can observe the observed object:

[0042]

[0043]

[0044] In this embodiment, a multi-agent cooperative task model is constructed based on a partially observable Markov decision process. The cooperative task model includes, for example, an observation space model, an action space model, a state space model, a reward function, a transfer function, etc.

[0045] Specifically, the multi-agent cooperation task is remodeled as a C-Dec POMDP. C-Dec POMDP consists of an 8-tuple The definition is as follows: ,express A collection of pursuer agents.

[0046] is the set of pursuer agents. Each pursuer agent Can be modeled as a 5-tuple . Specifically, Is the pursuer agent At time step The set of internal states of , including the possible positions and velocities of the pursuer agent during the task. Is the pursuer agent At time step The set of action states. Is the pursuer agent At time step The observation set represents all the information that the pursuer agent can observe in a given state. Is the pursuer agent At time step A collection of auxiliary information, indicating additional information transmitted from the outside. Is the pursuer agent At time step A strategy that maps observations to actions during a task.

[0047] is a set of states, including the states of the pursuer agent, the evader agent, the dynamic obstacles, and the static obstacles.

[0048] is the action set, including the actions of all pursuer agents.

[0049] Defined in joint action Next, from the state Transfer to state probability.

[0050] is a reward function that provides a global reward based on the current state and joint action.

[0051] is the set of observations of the agent.

[0052] is the discount factor applied to future rewards.

[0053] At time step , given the current state , each pursuer agent Both contain the local information obtained by themselves and the combination of auxiliary observation information The agent observes in its local With auxiliary observation, according to the strategy of your own components Select Action , and combine with other agents to form joint actions , interact with the environment and get rewards from the environment , the environment is based on the state transition function Move to the next state.

[0054] Specifically, the detailed definition and construction method of each element in the 8-tuple of C-Dec POMDP in the cooperative task model are as follows.

[0055] At time step , the state information in the global state space model can be established based on the following formula:

[0056] in, Indicates that at time step Time Chaser Agent status, is the number of pursuer agents; Indicates that the evader agent is The state space at time, represents the evader agent exist The state of the moment, is the number of evader agents; represents the initial state space of the static obstacle, Indicates static obstacles The initial state of Indicates that dynamic obstacles are The state space at time, Indicates dynamic obstacles exist The state of the moment, is the number of dynamic obstacles.

[0057] The action space model in the cooperative task model can be constructed in the following ways: Get the agent's Axis acceleration information and Axis acceleration information; based on the acceleration information of each agent Axis acceleration information and The motion space model is constructed based on the acceleration information of the axis.

[0058] Specifically, at the time step , the joint action space of the pursuer agent can be expressed as follows:

[0059] in, represents the pursuer agent In time action. Represents the pursuer agent In time along The acceleration of the axis, Represents the pursuer agent In time along The acceleration actions of the axes. These actions determine the velocity and position change of the pursuer agent at the next time step.

[0060] Therefore, at time , the joint action space of multi-agent pursuers It is expressed as:

[0061] The above state transition function Describes the joint action Next, from the current state Transition to next state The probability that the joint action contains all pursuers at time This function includes the state transition of the pursuer, evader, static obstacle and dynamic obstacle.

[0062] For each pursuer, its state transition probability is Determined by the above state update equations, these equations are based on time The position and speed of the escaper are updated. Similarly, for each escaper, the state transition probability is The states of static obstacles remain unchanged because they do not move. For dynamic obstacles, their state transition probabilities are Determined by the same equations of motion used to update its position and velocity.

[0063] In time The state transition function of the multi-agent system can be summarized as:

[0064]

[0065] This paper studies the collective pursuit-escape problem under partial observable information. The agent cannot obtain global state information. Therefore, the agent can only obtain observation information through its own sensors. Based on this, the observation space model in the cooperative task model can be constructed in the following way: The internal state information and external state information of each agent in each time step are obtained, and the external state information is the external local state information detected by the sensor of the agent; the observation space model is constructed based on the internal state information and the external state information.

[0066] Specifically, at time , the pursuer The observation space is ,in Indicates the pursuer In time The internal state of Indicates the pursuer In time The external state information observed by its sensors includes the distance and speed information of other objects, and whether there is a collision with the evader or dynamic obstacle.

[0067] External status information The definition is as follows:

[0068] in: Indicates at time By the pursuer The observed status of other pursuers.

[0069] Indicates that the evader is in time distance and speed.

[0070] Indicates at time The distance of static obstacles.

[0071] Indicates at time Distance and speed of dynamic obstacles.

[0072] Indicates at time The distance to the border.

[0073] Indicates at time Whether to capture evaders.

[0074] Indicates at time Whether there is a collision with an obstacle (static or dynamic).

[0075] Therefore, the pursuer's global observation space for:

[0076] The reward function is designed to guide the multi-agent pursuer system to successfully capture the evader. The objective reward function is defined by the following formula:

[0077] Obviously, only the target reward makes the pursuit-escape game a dynamic and complex problem with sparse rewards. When the pursuer does not detect the evader for a period of time, or detects the evader but does not capture them, the reward remains zero, which makes the policy gradient also zero and cannot achieve policy improvement. Therefore, in order to cope with the learning difficulties caused by sparse rewards, an auxiliary reward function is introduced as a regularization term of the target reward function.

[0078] First, in order to motivate the pursuer to quickly capture the evader, , if the tracker Detecting evaders using sensors , then the detection reward is distributed among all pursuers that detected the same evader. In this case, the agent You will get a small reward as follows:

[0079] in, is a positive constant indicating the strength of the reward, is the total positive reward for detecting evaders, The same evader is detected The number of followers.

[0080] Secondly, in order to avoid collisions with static and dynamic obstacles, If the pursuer Colliding with any obstacle will give a negative reward ,as follows:

[0081] in, is a positive constant indicating the intensity of the penalty. is a positive number greater than zero.

[0082] Finally, to encourage To quickly capture evaders, a time penalty reward function is designed. ,if If no evaders are caught, a negative reward is given based on the elapsed time, defined as follows:

[0083] in, and is a positive constant indicating the intensity of the time penalty.

[0084] In summary, the designed reward function and its auxiliary regularization term is defined as:

[0085] Based on the above, this embodiment constructs a deep reinforcement learning model for hierarchical games. The learning model is a regular auxiliary term deep deterministic policy gradient (HRG-MADDPG) algorithm model based on hierarchical games. Hierarchical games can be represented as a multi-level decision-making process, in which each decision level Has its own set of policies and utility function . Each level influences and optimizes each other through hierarchical relationships.

[0086] Please refer to Figure 3,Hierarchical game is an important framework of the HRG-MADDPG algorithm in this embodiment, which includes a high-level strategy module (HLS) and a low-level strategy module (LLS). The high-level strategy module is responsible for making global decisions, that is, formulating strategies for grouping and assigning targets to agents and individual targets, and providing guidance for the low-level strategy module. The low-level strategy module performs specific tasks according to the guidance of the high-level strategy module. By combining the high-level strategy module and the low-level strategy module, the decision-making ability of the multi-agent system in a complex environment is enhanced.

[0087] The high-level policy module plays a vital role in coordinating multi-agent systems. HLS uses a high-level policy network (HLNN) to process inputs and generate policy outputs.

[0088] In this embodiment, the game information of the intelligent agent that performs the pursuit and escape simulation according to the cooperative task model in the pursuit and escape game scenario can be stored in the playback buffer pool. The replay buffer pool stores the agent's past experience, including the agent's observation information, action information, reward information, next observation information, etc. The observation information includes the agent's local state information.

[0089] By collecting the agent's game information from the replay buffer pool, the high-level policy module and the low-level policy module in the deep reinforcement learning model are trained based on the game information.

[0090] Since the high-level policy module is mainly responsible for setting global goals and task allocation, its update frequency is lower than that of the low-level policy module. Therefore, when training the high-level policy module and the low-level policy module based on game information, it can be achieved in the following ways: The low-level strategy module is trained using the game information; when the number of training iterations of the low-level strategy module reaches the update frequency of the high-level strategy module, the high-level strategy module is trained based on the game information.

[0091] The training and updating process of the low-level policy module includes processes such as calculating action information, storing experience (game information of the agent), collecting small sample data, calculating target values, updating the network of the low-level policy module, and performing soft updates. First, the agent generates actions based on local observations and guidance information provided by the high-level policy module, and then stores the experience data of these actions in the replay buffer. Next, LLS randomly samples a small sample data from the replay buffer for training. Finally, soft updates are performed to ensure that the target network parameters gradually approach the actual network parameters. This process is performed at each time step, allowing LLS to quickly adapt to environmental changes and optimize agent behavior.

[0092] The HLS update frequency is expressed as , during LLS training, when the number of training iterations reaches the HLS update frequency When HLS is played, it samples data from the playback buffer and performs HLS training updates based on the sampled data and the constructed loss function.

[0093] Through this hierarchical training method, HLS and LLS work together to enhance the overall effect and stability of the algorithm.

[0094] Specifically, the training and updating of the high-level strategy module is realized based on the game information of the sampled agent. The game information includes the observation information, action information, reward function information of the current time step of the agent and the observation information of the next time step, which can be expressed as The steps of training the high-level strategy module based on game information can be achieved in the following ways: The current expected value is calculated based on the observation information and guidance information of the current time step; the target expected value is calculated based on the observation information and guidance information of the next time step; and the loss function constructed based on the current expected value and the target expected value guides the training of the high-level strategy module.

[0095] Please refer to Figure 4 Specifically, in the high-level policy module, the agent performs Accept the observation status of the current cluster agent sensor , action information , rewards for low-level strategy modules , the observation at the next moment As input, the main output of HLS is the guidance information for each agent The guidance information includes target ID assignment, cooperation coefficient and priority weight distribution.

[0096] Target ID assignment is used to determine the ID of the evader agent that each pursuer agent is pursuing. ,in and When Assign corresponding pursuers . The cooperation coefficient is used to determine the degree of cooperation between each pursuer agent and other pursuer agents. Each will receive a cooperation coefficient , which represents the pursuer The degree of cooperation with other pursuers when performing tasks. The cooperation coefficient ranges from 0 to 1, with higher values ​​indicating a higher degree of cooperation. For example, Figure 4As shown in Figure 2, when two agents observe the same evader at the same time, their cooperation coefficient will be higher; conversely, when the agents do not observe any evaders, or only one pursuer observes the evader, the cooperation coefficient will be lower.

[0097] The tasks of the pursuer agent include pursuit tasks and avoidance tasks. The priority weight distribution includes the weight of the pursuit task and the weight of the avoidance task of each pursuer agent. That is, it is used to determine the priority of each pursuer when performing various tasks. A priority weight is assigned , the weight indicates the priority of the pursuer in the current task. Figure 4 As shown, the priority weights can determine whether the pursuer's primary task is to pursue the assigned evader or to avoid obstacles.

[0098] Specifically, when training HLS, at the time step , HLS from the playback buffer pool Randomly sample a mini-batch of data from The replay buffer stores the agent's past experience, including observation information, action information, reward function information, and next observation information. Sampling a small batch of data for training enables the algorithm to learn from diverse past interactions, thereby improving the robustness and generalization ability of the strategy. The calculation formula of the guidance information is:

[0099] Next, the algorithm calculates the current expected value , denoted as , indicating that the agent starts from the current state and follows the calculated guidance information The expected value that can be achieved. The value is a key component in evaluating the effectiveness of a strategy, which uses the current observation and guidance information :

[0100] To optimize the strategy, HLS uses the Temporal Difference (TD) error algorithm. It selects a new sample from the mini-batch data that includes the observation information of the next time step. , and calculate the guidance information for the next time step The guidance information for the next time step is calculated using the same HLNN, but applied to the observation at the next moment:

[0101] After calculating the current strategy and the target strategy, the algorithm calculates the target expected value , which represents the expected value obtained from the future state. The value is passed through the time from Instant rewards received Discount with next status Adding the values ​​gives:

[0102] in is the discount factor.

[0103] Loss Function Calculated as the current expectation value Target expectations value In addition, a regularization term is included to penalize the policy parameters The regularization term is determined by the hyperparameter Control, determines the intensity of the penalty for large weights. Weight Indicates policy parameters Each independent parameter in. Loss function The calculation is as follows:

[0104] Finally, the HLS parameters are updated by performing gradient descent on the loss function :

[0105] This update adjusts the parameters to minimize the loss, effectively improving HLS over time.

[0106] This iterative process continues until the policy parameters converge to the optimal solution, or until the predetermined number of iterations is reached to complete the training of the high-level policy module. The final output of HLS is the high-level guidance information ,This guidance information is passed to LLS to guide the agent’s actions at a more fine-grained level.

[0107] The core of the low-level policy module is to use a classic multi-agent deep reinforcement learning algorithm MADDPG, which includes two stages: offline training and online execution. In the offline training stage, LLS includes the training of shared value network and shared policy network.

[0108] In the online execution phase, based on the shared policy network, a policy improvement method is used to enhance the decision-making of each agent. , the pursuer According to the guidance information provided by HLS Considering the large scale and isomorphism of intelligent agents in the cluster pursuit problem, in this embodiment, LLS adopts a strategy sharing method to simplify the network, that is, all pursuers share a single policy with parameters Shared Strategy Network and one with parameters Shared Value Network In addition, the shared strategy network There is also its corresponding target strategy network , the parameters are , a shared value network There is also a corresponding target value network , the parameters are The use of shared networks greatly reduces the number of networks and reduces the computational complexity.

[0109] Based on this, when using game information to train the low-level strategy module, it can be achieved in the following ways: For each agent, the target value is calculated based on the agent's game information, and the shared value network is updated by minimizing the loss function constructed by the current value and the target value; The observation information in the game information and the guidance information provided by the high-level strategy module are used as the input of the shared strategy network, the shared strategy network is used to output the corresponding action information, and the shared strategy network is updated using a gradient ascent method based on the cumulative rewards under the action information.

[0110] In this embodiment, for the shared policy network, at time , the pursuer The environmental state observed by the sensor is expressed as For each pursuer , shared policy network The observation at the current time step As input, generate the corresponding action for the next time step .Strategy Based on the observed state and guidance information For the pursuer Generate corresponding actions :

[0111] In a shared value network, a shared value network Using joint observations and joint actions To calculate the pursuer's Value. Recorded at time moment, through Calculating the pursuer The value of ,but:

[0112] Similarly, the value output based on the target network can be defined as:

[0113] The pursuer's goal is to maximize its expected cumulative discounted return ,in represents the cumulative return over time and is defined as:

[0114] in is the discount factor, It's the pursuer At time step Rewards obtained with assists.

[0115] The cumulative reward is obtained by the reward function with an auxiliary reward regularization term obtained by the agent in the cumulative time step. The auxiliary reward regularization term is constructed based on detecting whether the pursuer agent captures the evader agent, whether the pursuer agent detects the evader agent, whether the pursuer agent collides with an obstacle, and the number of evader agents captured by the pursuer agent.

[0116] Specifically, The calculation is as follows:

[0117] The definition and calculation method of each auxiliary reward regularization term can refer to the relevant description above in this embodiment.

[0118] Then, the pursuer The shared policy network is updated by gradient ascent, and its gradient expression is:

[0119] Here, the experience replay buffer Contains tuples , records the experience of all agents. Centralized action value function Used to train the shared value network by minimizing the squared TD error The loss function is composed of the current value and target value , and is constructed according to the following formula:

[0120] in and See above for the definition of .

[0121] Based on the above method, the training of high-level strategy modules and low-level strategy modules in the deep reinforcement learning model is completed.

[0122] The trained deep reinforcement learning model is applied to different test environments for pursuit and escape simulation test verification. Specifically, the trained high-level strategy module can obtain guidance information for the intelligent agent playing the game, and the low-level strategy module obtains the action information of the intelligent agent based on the guidance information, and controls the intelligent agent to play the pursuit and escape game based on the action information.

[0123] In this embodiment, a partially observable cluster pursuit control scheme based on hierarchical game deep reinforcement learning is proposed. By introducing a hierarchical game control structure and a regularized auxiliary reward mechanism, each intelligent agent can collaborate efficiently in a partially observable environment, thereby improving the collaboration efficiency and adaptability of cluster agents in a partially observable environment.

[0124] Through the hierarchical game structure system, collaborative strategies can be generated at the global and local levels to maximize the efficiency of cluster collaboration. Specifically, the high-level strategy module is responsible for the global target allocation of tasks and generates task allocation strategies based on the current observations of the agents, including target allocation, coordination coefficients, and priority weights. This strategy is generated by a high-level neural network (HLNN) to guide the collaborative behavior of each agent in a partially observable environment. The low-level strategy module (LLS) uses the multi-agent deep deterministic policy gradient (MADDPG) algorithm to optimize the strategy convergence and learning efficiency of each agent in the centralized training phase in combination with regularized auxiliary rewards; in the distributed execution phase, each agent independently executes the high-level task allocation strategy based on local observation information and performs specific pursuit behaviors.

[0125] In addition, a regularized auxiliary reward mechanism is introduced, and auxiliary reward items such as detection reward, collision penalty, time penalty and capture reward are introduced to ensure the efficiency of cooperation between agents, reduce invalid behaviors, and improve the overall strategy convergence speed. The difficulty of strategy optimization caused by sparse rewards in some observable environments is solved, ensuring the adaptability and robustness of the system in complex environments.

[0126] In summary, this embodiment provides a reliable solution for cluster pursuit, which can be trained in a single environment and verified in multiple different test environments, showing strong robustness and generalization capabilities. This solution not only realizes the collaboration of intelligent agents in task allocation, but also adds a regularized auxiliary reward mechanism in the pursuit execution phase, which optimizes the convergence and stability of the strategy. The hierarchical game control architecture improves the pursuit effect of multi-agent clusters in partially observable environments, overcomes the information bottleneck problem of traditional centralized control in complex environments, and makes the system more robust and real-time. The proposed framework is adaptable, scalable, and responsive to complex dynamic environments in various practical applications.

[0127] Based on the same inventive concept, please refer to Figure 5 , an embodiment of the present invention also provides a functional module diagram of a cluster pursuit control system based on deep reinforcement learning of hierarchical games. This embodiment can divide the functional modules of the cluster pursuit control system based on deep reinforcement learning of hierarchical games according to the above method embodiment. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present invention is schematic, which is only a logical function division. There may be other division methods in actual implementation.

[0128] For example, when each functional module is divided into corresponding functional modules, Figure 5 The cluster pursuit control system based on hierarchical game deep reinforcement learning shown is only a schematic diagram of a device. The cluster pursuit control system based on hierarchical game deep reinforcement learning can include a construction module, a collection module, a training module and a guidance module. The functions of each functional module of the cluster pursuit control system based on hierarchical game deep reinforcement learning are described in detail below.

[0129] A construction module is used to construct a cluster pursuit game scenario, wherein the cluster pursuit game scenario includes a game environment and multiple intelligent agents and obstacles in the game environment, each of the intelligent agents carries a sensor, and the sensor is used to detect local state information within the sensor range; The building block is also used to construct multi-agent cooperative task models based on partially observable Markov decision processes; The construction module is also used to construct a deep reinforcement learning model of hierarchical game, wherein the deep reinforcement learning model includes a high-level strategy module and a low-level strategy module; A collection module, used to collect game information of an intelligent agent that performs pursuit and escape simulation according to the cooperative task model in the pursuit and escape game scenario, wherein the game information includes local state information; A training module, used for training the low-level strategy module and the high-level strategy module using the game information; The guidance module is used to use the trained high-level strategy module to obtain guidance information for the intelligent agent playing the game, and obtain the action information of the intelligent agent based on the guidance information through the low-level strategy module, and control the intelligent agent to play the pursuit and escape game based on the action information.

[0130] It can be understood that the above-mentioned construction module, acquisition module, training module and guidance module can be used to execute the above-mentioned S11 to S16. The detailed implementation method of the construction module, acquisition module, training module and guidance module can refer to the relevant contents of the above-mentioned S11 to S16.

[0131] The cluster pursuit control system based on hierarchical game deep reinforcement learning provided in this embodiment has the same, similar or corresponding technical features as the cluster pursuit control method based on hierarchical game deep reinforcement learning in the above embodiment, and has the same technical effects. For relevant contents of the control system, please refer to the description related to the above control method, and this embodiment will not be repeated here.

[0132] See also Figure 6 , is a block diagram of an electronic device provided in an embodiment of the present invention, the electronic device may be a computer device, etc., and the electronic device includes a memory, a processor, and a communication module. The memory, the processor, and the communication module are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines.

[0133] The memory is used to store computer programs or data. The memory can be, but is not limited to, random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0134] The processor is used to read / write data or programs stored in the memory, and execute the cluster pursuit control method based on hierarchical game deep reinforcement learning provided by any embodiment of the present invention.

[0135] The communication module is used to establish a communication connection between the electronic device and other communication terminals through the network, and is used to send and receive data through the network.

[0136] It should be understood that Figure 6 The structure shown is only a schematic diagram of the structure of the electronic device. The electronic device may also include Figure 6 More or fewer components as shown, or with Figure 6 Different configurations are shown.

[0137] Furthermore, an embodiment of the present invention also provides a computer-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are executed, the cluster pursuit and escape control method based on hierarchical game deep reinforcement learning provided in the above embodiment is implemented.

[0138] Specifically, the computer-readable storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the computer-readable storage medium is executed, the above-mentioned cluster pursuit control method based on hierarchical game deep reinforcement learning can be executed. Regarding the process involved when the computer-readable storage medium and its executable instructions are executed, reference can be made to the relevant description in the above-mentioned method embodiment, which will not be described in detail here.

[0139] In summary, the cluster pursuit control method and system based on hierarchical game deep reinforcement learning provided by the embodiment of the present invention models the multi-agent collaborative task as a collective decentralized partially observable Markov decision process through a hierarchical game structure, and proposes a hierarchical game multi-agent deep deterministic policy gradient algorithm. The algorithm includes high-level strategies and low-level strategies. The high-level strategy is responsible for target allocation and task coordination. The low-level strategy optimizes specific action decisions through centralized training and distributed execution, and combines regularized auxiliary functions to improve the convergence speed and execution efficiency of the strategy. This solution can effectively improve the collaborative efficiency and response capabilities of multi-agent systems in complex environments.

[0140] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0141] In addition, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0142] Furthermore, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0143] It should be noted that if the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0144] The above description is only an embodiment of the present invention and is not intended to limit the protection scope of the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A cluster pursuit control method based on hierarchical game deep reinforcement learning, characterized in that: The method comprises: Constructing a cluster pursuit game scenario, wherein the cluster pursuit game scenario includes a game environment and a plurality of intelligent agents and obstacles in the game environment, each of the intelligent agents carries a sensor, and the sensor is used to detect local state information within a detection range of the sensor; Construct a multi-agent cooperative task model based on partially observable Markov decision process; Constructing a deep reinforcement learning model of hierarchical game, wherein the deep reinforcement learning model includes a high-level strategy module and a low-level strategy module; Collecting game information of an intelligent agent that performs pursuit and escape simulation according to the cooperative task model in the pursuit and escape game scenario, wherein the game information includes local state information; Using the game information to train the low-level strategy module and the high-level strategy module; The trained high-level strategy module is used to obtain guidance information for the intelligent agent playing the game, and the action information of the intelligent agent is obtained based on the guidance information through the low-level strategy module, and the intelligent agent is controlled to play the pursuit-escape game based on the action information.

2. The cluster pursuit control method based on hierarchical game deep reinforcement learning according to claim 1 is characterized in that: The step of training the low-level strategy module and the high-level strategy module using the game information comprises: Using the game information to train the low-level strategy module; When the number of training iterations of the low-level policy module reaches the update frequency of the high-level policy module, the high-level policy module is trained based on the game information.

3. The cluster pursuit control method based on hierarchical game deep reinforcement learning according to claim 2 is characterized in that: The low-level strategy module includes a shared value network and a shared strategy network; The step of training the low-level strategy module using the game information includes: For each agent, a target value is calculated based on the game information of the agent, and the shared value network is updated by minimizing a loss function constructed by a current value and a target value; The observation information in the game information and the guidance information provided by the high-level strategy module are used as inputs of the shared strategy network, the shared strategy network is used to output corresponding action information, and the shared strategy network is updated using a gradient ascent method based on the cumulative rewards under the action information.

4. The cluster pursuit control method based on hierarchical game deep reinforcement learning according to claim 3 is characterized in that: The cumulative reward is obtained by the reward function with auxiliary reward regularization term obtained by the agent in the cumulative time step; The auxiliary reward regularization term is constructed based on detecting whether the pursuer agent captures the evader agent, whether the pursuer agent detects the evader agent, whether the pursuer agent collides with an obstacle, and the number of evader agents captured by the pursuer agent.

5. The cluster pursuit control method based on hierarchical game deep reinforcement learning according to claim 2 is characterized in that: The game information includes observation information, action information, reward function information of the current time step of the agent and observation information of the next time step; The step of training the high-level strategy module based on the game information includes: Inputting the observation information, action information, reward function information of the agent at the current time step and the observation information of the next time step included in the game information into the high-level strategy module to obtain the guidance information of the current time step and the guidance information of the next time step; The current expected value is calculated based on the observation information and guidance information of the current time step; The target expected value is calculated based on the observation information and guidance information of the next time step; The high-level strategy module is trained based on a loss function constructed based on the current expected value and the target expected value.

6. The cluster pursuit control method based on hierarchical game deep reinforcement learning according to claim 2 is characterized in that: The intelligent agents include a pursuer intelligent agent and an evader intelligent agent, and the tasks of the pursuer intelligent agent include a pursuit task and an evasion task; The guidance information includes target ID allocation, cooperation coefficient and priority weight distribution; The target ID allocation is used to determine the ID of the evader agent that each pursuer agent is pursuing; The cooperation coefficient is used to determine the degree of cooperation between each pursuer agent and other pursuer agents; The priority weight distribution includes the weight of the pursuit task and the weight of the avoidance task of each pursuer agent.

7. The cluster pursuit control method based on hierarchical game deep reinforcement learning according to claim 1 is characterized in that: Each of the intelligent agents determines whether an observation object is detected based on the sensors carried by it in the following manner, where the observation object is any other intelligent agent or obstacle: Obtaining a projection distance between the agent and the observed object in the sensor direction, and obtaining an actual distance between the agent and the observed object; Detecting an actual relative speed and a projected relative speed between the intelligent agent and the observed object; Based on the projection distance, actual distance, actual relative speed, projection relative speed, radius of the observed object and sensor vector, it is determined whether the intelligent agent detects the observed object according to a preset judgment formula.

8. The cluster pursuit control method based on hierarchical game deep reinforcement learning according to claim 1 is characterized in that: The cooperative task model includes an observation space model, and the step of constructing the observation space model includes: Obtaining internal state information and external state information of each of the agents in each time step, wherein the external state information is external local state information detected by the sensor of the agent; An observation space model is constructed based on the internal state information and the external state information.

9. The cluster pursuit control method based on hierarchical game deep reinforcement learning according to claim 1 is characterized in that: The cooperative task model includes an action space model, and the steps of constructing the action space model include: Obtaining acceleration information of each of the agents on the x-axis and on the y-axis at each time step; An action space model is constructed based on the acceleration information of each intelligent agent on the x-axis and the acceleration information on the y-axis.

10. A cluster pursuit control system based on hierarchical game deep reinforcement learning, characterized in that: The system comprises: A construction module is used to construct a cluster pursuit game scenario, wherein the cluster pursuit game scenario includes a game environment and multiple intelligent agents and obstacles in the game environment, each of the intelligent agents carries a sensor, and the sensor is used to detect local state information within the sensor detection range; The building block is also used to construct multi-agent cooperative task models based on partially observable Markov decision processes; The construction module is also used to construct a deep reinforcement learning model of hierarchical game, wherein the deep reinforcement learning model includes a high-level strategy module and a low-level strategy module; A collection module, used to collect game information of an intelligent agent that performs pursuit and escape simulation according to the cooperative task model in the pursuit and escape game scenario, wherein the game information includes local state information; A training module, used for training the low-level strategy module and the high-level strategy module using the game information; The guidance module is used to use the trained high-level strategy module to obtain guidance information for the intelligent agent playing the game, and obtain the action information of the intelligent agent based on the guidance information through the low-level strategy module, and control the intelligent agent to play the pursuit and escape game based on the action information.

Citation Information

Patent Citations

  • Vacuum flat plate heat collector and manufacturing method and equipment thereof

    CN111947323A

  • Protection mechanism for intelligent self-service ticket taking machine

    CN112233339A

  • Electric motor driving device

    CN113474985A

  • Image rendering method, device and equipment and readable storage medium

    CN114253649A

  • Optical fiber radio frequency replication jammer

    CN114422077A

Cited By

  • Environmental disturbance-oriented pursuit game strategy solving method

    CN120851183A