Multi-agent hunting method based on obstacle terrain information optimization

Through the multi-agent roundup method based on obstacle terrain information optimization, combined with reinforcement learning and artificial potential field method, the agent can efficiently use obstacles for roundup in complex environments, solving the problem of difficulty in effectively utilizing obstacles in the prior art, and improving the task execution efficiency and success rate.

CN120046642APending Publication Date: 2025-05-27天津(滨海)人工智能创新中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411889863.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize obstacles in the environment as terrain advantages to assist multiple agents in completing round-up tasks, especially in complex environments, where agents need to avoid obstacles and use obstacles to block prey action routes.

Method used

A multi-agent roundup method based on obstacle terrain information optimization is proposed. By obtaining the observation information of each pursuer agent, combining a pre-trained reinforcement learning model for calculation, the action execution instructions of each pursuer agent are obtained, and using artificial potential field method and attention coding technology to optimize the movement of the agent to use obstacles to round up more efficiently.

Benefits of technology

It realizes the use of obstacles more efficiently in complex environments for roundup, improves the execution efficiency and success rate of tasks, and alleviates the low efficiency and dimensional disasters of large-scale high-dimensional data utilization in reinforcement learning of multi-agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046642A_ABST
    Figure CN120046642A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent hunting method based on obstacle terrain information optimization. The multi-agent hunting method comprises the following steps: acquiring observation information of each pursuer agent; calculating based on the observation information of each pursuer agent in combination with a reinforcement learning model to obtain an action execution instruction of each pursuer agent; performing a corresponding hunting action based on the action execution instruction of each pursuer agent; obstacles can be used as terrain advantages to block the action route of preys, and hunting tasks are completed; it is guaranteed that feature embedding of the agents is irrelevant to the number of the agents through attention coding, the problems of low utilization efficiency of large-scale high-dimensional data and dimension disasters in multi-agent reinforcement learning are effectively solved, and the learning efficiency is improved; irregular and non-convex obstacles are converted into a simplified convex hull form through a key point discretization method, so that a pursuer can more efficiently utilize the topographic advantage to carry out hunting; obstacle avoidance processing is carried out through an artificial potential field method, and local minimum can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of multi-agent systems, and particularly relates to a multi-agent pursuit method optimized based on obstacle terrain information. Background Art

[0002] In multi-agent pursuit tasks, the influence of obstacles is multi-faceted. On the one hand, when an agent encounters an obstacle during movement, it poses a serious threat to the safety of the agent. The agent must adopt appropriate obstacle avoidance strategies to ensure its own safety to ensure the smooth completion of the task and achieve the desired effect. On the other hand, the agent can skillfully utilize the terrain advantages of obstacles in the environment for avoidance. For example, the agent can use the shielding effect of obstacles to hide itself and enhance concealment. Or, by leveraging the terrain advantages of obstacles, it can skillfully conduct pursuit and interception to enhance the execution efficiency of the task.

[0003] In most research works, the focus is mainly on the study of obstacle avoidance. One type of research adopts a rule-based method, such as introducing Voronoi polygons, defining the obstacle perception boundary vector as a safe area, and each agent continuously calculates the safe unit and plans its motion control in a recursive manner. Some other research introduces the virtual potential field method, which obtains the resultant force acting on the pursuer by calculating the repulsive force between the pursuer and the obstacle and the attractive force between the pursuer and the prey, thereby guiding the movement of the pursuer to complete the task while avoiding obstacles. Although these methods have achieved a certain degree of success in agent obstacle avoidance, most of these methods require accurate global information or need to manually design rules, which are difficult to obtain in a complex multi-agent pursuit environment.

[0004] Deep reinforcement learning is a promising trend for solving multi-agent cooperation and has become the main method for solving the multi-agent pursuit problem. That is, agents interact with the environment, collect and learn experience information, and thus learn pursuit strategies. The paper "Multi-robot cooperative pursuit via potential field-enhanced reinforcement learning" published by Zhang et al. in the "International Conference on Robotics and Automation" in 2023 combines deep reinforcement learning and artificial potential field, and proposes a distributed cooperative tracking algorithm DACOOP. The artificial potential field method is used to calculate the resultant force to determine the reference heading of the pursuers. Further, in the paper "DACOOP-A: Decentralized Adaptive Cooperative Pursuit via Attention" published by the authors in the "IEEE Robotics and Automation Letters" in 2023, the artificial potential field and attention mechanism are used to enhance the ability of reinforcement learning, and an attention-based distributed adaptive tracking method DACOOP-A is proposed. The weights of the learned neighboring agents are integrated into the observation state and interaction rules through attention. The paper "Multi-target pursuit by a decentralized heterogeneous uav swarm using deep multi-agent reinforcement learning" published by Kouzeghar et al. in the "International Conference on Robotics and Automation" in 2023 proposes a role-based MADDPG method to handle the non-stationary multi-target pursuit problem in an unknown environment with partially observable and non-static random obstacles. By assigning two roles to the pursuers, chasing known targets and searching for potential targets. The Thiessen polygon is used to assign tasks and rewards to optimize the exploration of the environment by the agents.

[0005] However, the above methods mainly focus on avoiding obstacles in the environment to ensure the safety of the agents. However, the impact of obstacles in the environment is multi-faceted. Some studies have shown that pursuers can use obstacles to guide the prey to dead ends or enclosed areas, greatly restricting the available movement space of the prey (the paper "ALeapfrog Strategy for Pursuit-Evasion in aPolygonal Environment" published by Ames et al. in the "International Journal of Computational Geometry Applications" in 2014). When there are obstacles in the environment, the pursuers can adjust their positions according to the distribution of the obstacles to better form a uniform distribution around the prey. In this way, the pursuers can reduce the possibility of the prey breaking through the encirclement using the gaps. In a highly complex environment, the pursuers not only need to avoid collisions with obstacles, but can also use obstacles to block the movement routes of the prey to assist in completing the encirclement task. Summary of the Invention

[0006] To solve the problem of how to utilize the obstacles in the environment as a terrain advantage to assist in completing the multi-agent encirclement task in the prior art, this application proposes a multi-agent encirclement method optimized based on obstacle terrain information, including:

[0007] Obtain the observation information of each pursuer agent;

[0008] Based on the observation information of each pursuer agent and combined with a pre-trained reinforcement learning model, calculate to obtain the action execution instructions for each pursuer agent;

[0009] Perform corresponding encirclement actions based on the action execution instructions of each pursuer agent;

[0010] Among them, the reinforcement learning model is obtained by training the models of each pursuer agent using a multi-agent reinforcement learning algorithm based on the observation information of the pursuer agent.

[0011] Preferably, the training process of the reinforcement learning model includes:

[0012] Based on the pursuer agents equipped with the reinforcement learning model, the prey agent, and the obstacles, construct a simulation scenario for the multi-agent encirclement task;

[0013] Based on the simulation scenario, use the multi-agent reinforcement learning algorithm to train the reinforcement learning models corresponding to each pursuer agent.

[0014] Preferably, training the reinforcement learning model corresponding to each pursuer agent by using the multi-agent reinforcement learning algorithm based on the simulation scenario includes:

[0015] S1 Obtain the position and speed information of the pursuer agent itself, the position and speed information of other pursuer agents, the position information of obstacles, and the position and speed information of the prey agent based on the simulation scenario, and perform attention encoding based on the observation information of the pursuer agent to obtain the processed state features, and enter S2;

[0016] S2 Perform key-point discretization processing based on the position information of the obstacles to obtain the processed obstacle information, and enter S3;

[0017] S3 Perform artificial potential field calculation based on the position and speed information of the pursuer agent itself, the position and speed information of the prey agent, and the processed obstacle information to obtain the resultant force received by the pursuer agent, and enter S4;

[0018] S4 Calculate based on the processed state features, the action direction obtained by inputting the processed obstacle information into the policy network, and the physical guidance direction provided by the resultant force received by the pursuer agent to obtain the action execution instruction of the pursuer agent, and enter S5;

[0019] S5 Execute the corresponding hunting action based on the action execution instruction of the pursuer agent, return to S1 after obtaining the feedback reward value until the hunting is successful.

[0020] Preferably, S3 performs artificial potential field calculation based on the position and speed information of the pursuer agent itself, the position and speed information of the prey agent, and the processed obstacle information to obtain the resultant force received by the pursuer agent, including:

[0021] Calculate based on the position of the pursuer agent and the key-point position of the obstacle agent closest to the pursuer agent to obtain the repulsive force between the pursuer agent and the obstacle;

[0022] Calculate based on the position of the prey agent and the position of the pursuer agent to obtain the attraction of the prey agent to the pursuer agent;

[0023] Calculate the resultant force received by the pursuer agent by adding the repulsive force between the pursuer agent and the obstacle and the attraction of the prey agent to the pursuer agent.

[0024] Preferably, the calculation formula of the reward value is as follows:

[0025]

[0026] In the formula, R pos_obstacleThe reward value designed according to the dot product; dot product is the dot product;

[0027]

[0028] In the formula, R dir_obstacle is the reward value designed according to the included angle; cosθ is the included angle between the moving direction of the pursuer agent and the direction of the prey agent to the obstacle;

[0029]

[0030] In the formula, R capture is the reward value for the pursuer agent to complete the encirclement; d i is the distance between the pursuer agent and the prey agent; l is the encirclement radius; std(θ) is the standard deviation of the angle distribution of the pursuer agent; threshold is the set threshold.

[0031] Preferably, the observation information of each pursuer agent includes: the position and speed information of the pursuer agent itself, the position and speed information of other pursuer agents, the position information of the obstacle, and the position and speed information of the prey agent.

[0032] Based on the same application concept, the present application also proposes a multi-agent encirclement system optimized based on obstacle terrain information, including:

[0033] An information acquisition module for acquiring the observation information of each pursuer agent;

[0034] A model calculation module for calculating based on the observation information of each pursuer agent in combination with a pre-trained reinforcement learning model to obtain the action execution instructions of each pursuer agent;

[0035] An instruction execution module for performing corresponding encirclement actions based on the action execution instructions of each pursuer agent;

[0036] Among them, the reinforcement learning model is obtained by training the models of each pursuer agent using a multi-agent reinforcement learning algorithm based on the observation information of the pursuer agent.

[0037] Preferably, it further includes a model training module; the model training module includes:

[0038] A simulation scenario construction sub-module for constructing a simulation scenario of a multi-agent encirclement task based on the pursuer agent, the prey agent, and the obstacle provided with a reinforcement learning model;

[0039] The reinforcement learning model training sub-module is used to train the reinforcement learning models corresponding to each pursuer agent based on the simulation scenario using the multi-agent reinforcement learning algorithm.

[0040] Preferably, the reinforcement learning model training sub-module is specifically used for:

[0041] S1: Obtain the position and speed information of the pursuer agent itself, the position and speed information of other pursuer agents, the position information of obstacles, and the position and speed information of the prey agent based on the simulation scenario, and perform attention encoding based on the observation information of the pursuer agent itself to obtain the processed state features, and enter S2;

[0042] S2: Perform key point discretization processing based on the position information of obstacles to obtain the processed obstacle information, and enter S3;

[0043] S3: Perform artificial potential field calculation based on the position and speed information of the pursuer agent itself, the position and speed information of the prey agent, and the processed obstacle information to obtain the resultant force received by the pursuer agent, and enter S4;

[0044] S4: Calculate based on the action direction obtained by inputting the processed state features and the processed obstacle information into the policy network, and combine the physical guidance direction provided by the resultant force received by the pursuer agent to obtain the action execution instruction of the pursuer agent, and enter S5;

[0045] S5: Execute the corresponding hunting action based on the action execution instruction of the pursuer agent, return to S1 after obtaining the feedback reward value until the hunting is successful.

[0046] Preferably, in S3 of the reinforcement learning model training sub-module, performing artificial potential field calculation based on the position and speed information of the pursuer agent itself, the position and speed information of the prey agent, and the processed obstacle information to obtain the resultant force received by the pursuer agent includes:

[0047] Calculate based on the position of the pursuer agent and the key point position of the obstacle agent closest to the pursuer agent to obtain the repulsive force between the pursuer agent and the obstacle;

[0048] Calculate based on the position of the prey agent and the position of the pursuer agent to obtain the attraction of the prey agent to the pursuer agent;

[0049] Calculate the resultant force received by the pursuer agent by adding the repulsive force between the pursuer agent and the obstacle and the attraction of the prey agent to the pursuer agent.

[0050] Preferably, the calculation formula of the reward value in the reinforcement learning model training sub-module is as follows:

[0051]

[0052] Wherein, R pos_obstacle is the reward value designed according to the dot product; dot product is the dot product;

[0053]

[0054] Wherein, R dir_obstacle is the reward value designed according to the included angle; cosθ is the included angle between the moving direction of the pursuer agent and the direction of the prey agent to the obstacle;

[0055]

[0056] Wherein, R capture is the reward value for the pursuer agent to complete the encirclement; d i is the distance between the pursuer agent and the prey agent; l is the encirclement radius; std(θ) is the standard deviation of the angle distribution of the pursuer agent; threshold is the set threshold.

[0057] Preferably, the observation information of each pursuer agent includes: the position and speed information of the pursuer agent itself, the position and speed information of other pursuer agents, the position information of the obstacle, and the position and speed information of the prey agent.

[0058] Compared with the prior art, the beneficial effects of the present application are:

[0059] A multi-agent hunting method optimized based on obstacle terrain information, including: obtaining the observation information of each pursuer agent; calculating based on the observation information of each pursuer agent in combination with a pre-trained reinforcement learning model to obtain the action execution instructions of each pursuer agent; performing corresponding hunting actions based on the action execution instructions of each pursuer agent; wherein, the reinforcement learning model is trained by using a multi-agent reinforcement learning algorithm based on the observation information of the pursuer agent for the models set for each pursuer agent; the reinforcement learning model of the present application can use obstacles as terrain advantages to block the action route of the prey and complete the hunting task;

[0060] The present application ensures that the feature embedding of the agent is independent of the number of agents through attention encoding, effectively alleviating the problems of low utilization efficiency of large-scale high-dimensional data and dimensionality disaster in multi-agent reinforcement learning, and improving the learning efficiency;

[0061] The present application converts irregular and non-convex obstacles into a simplified convex hull form through the method of key point discretization, enabling the pursuers to more efficiently utilize terrain advantages for hunting;

[0062] This application performs obstacle avoidance processing through the artificial potential field method, which can avoid local minima, thereby preventing the agent from falling into an undesirable state during movement. Description of the Drawings

[0063] Figure 1 It is a flowchart of a multi-agent pursuit method optimized based on obstacle terrain information of this application;

[0064] Figure 2 It is the average reward curve graph of the agent of this application;

[0065] Figure 3 It is a flowchart of a multi-agent pursuit method optimized based on obstacle terrain information of this application;

[0066] Figure 4 It is a schematic diagram of the main step process of the multi-agent pursuit method optimized based on obstacle terrain information of this application;

[0067] Figure 5 It is a system structure diagram of a multi-agent pursuit method optimized based on obstacle terrain information of this application. Detailed Embodiment

[0068] As disclosed in the background art, in the multi-agent pursuit task, the influence of obstacles is multi-faceted. On the one hand, when an agent encounters an obstacle during movement, it poses a serious threat to the safety of the agent. The agent must adopt appropriate obstacle avoidance strategies to ensure its own safety to ensure the smooth completion of the task and achieve the desired effect. On the other hand, the agent can skillfully utilize the terrain advantages of obstacles in the environment for evasion. For example, the agent can use the shielding effect of obstacles to hide itself and enhance concealment. Or by leveraging the terrain advantages of obstacles, it can skillfully conduct encirclement and interception to enhance the execution efficiency of the task. In order to better utilize environmental information and encourage the pursuers to use the obstacles in the environment as terrain advantages to assist in completing the multi-agent pursuit task, this application proposes a multi-agent pursuit method optimized based on obstacle terrain information. To better understand this application, the content of this application will be further described below in conjunction with the drawings in the specification and embodiments.

[0069] A multi-agent pursuit method optimized based on obstacle terrain information, the specific process is as Figure 1 shown, including:

[0070] Step 1, obtain the observation information of each pursuer agent;

[0071] Step 2, based on the observation information of each pursuer agent and combined with a pre-trained reinforcement learning model for calculation, obtain the action execution instructions of each pursuer agent;

[0072] Step 3: Perform corresponding encirclement actions based on the action execution instructions of each pursuer agent;

[0073] Among them, the reinforcement learning model is obtained by training the models of each pursuer agent using a multi-agent reinforcement learning algorithm based on the observation information of the pursuer agents.

[0074] Before step 1, there is also a training process for the reinforcement learning model, specifically including:

[0075] Construct a multi-agent encirclement task simulation scenario including pursuer agents equipped with a reinforcement learning model, prey agents, and obstacles, including:

[0076] Construct a multi-agent encirclement task simulation scenario in the multi-agent simulation environment of Open AI's open-source multiagent-particle-envs to prepare for the training of a multi-agent reinforcement learning model optimized based on obstacle terrain information.

[0077] Install the MPE simulation environment on any computer equipped with ubuntu and the pytorch deep learning framework, and construct an agent encirclement task simulation scenario.

[0078] In the constructed simulation scenario, set the size of the overall simulation map to 1×1, set the number of agents and obstacles in the environment, the speed, acceleration, size, and color of the agents, and the initial positions, sizes, and colors of the obstacles.

[0079] An agent refers to an unmanned node with capabilities such as perception, communication, movement, storage, and computing, such as drones, robots, etc., including but not limited to the agent particles constructed in the simulation environment. The described simulation encirclement task environment is an entity that interacts with the agent based on encirclement scenario parameters. The agent observes the state of the environment and acts in the environment based on this state according to the control instructions.

[0080] An agent consists of a perception module, an action module, a storage module, a multi-agent reinforcement learning model optimized based on obstacle terrain information, and a control module.

[0081] The perception module is connected to the feature extraction module and the storage module. The perception module obtains from the simulation task environment information including the state of the agent including the position information and speed information of the agent itself and the state of the environment including the position and speed information of other agents and the position information of the obstacles, and sends the information to the feature encoding module and the storage module.

[0082] The action module is the executor of the agent's control instructions, connected to the control module, receiving instructions from the control module, and moving in the simulation task environment according to the control instructions.

[0083] The storage module is a memory with more than 1GB of available space, connected to the perception module, the control module, and the data augmentation module. It receives observation information from the perception module, control instruction information from the control module, and reward information from the simulation capture task environment, and combines the observation information, control instruction information, and reward information into the trajectory data of the interaction between the agent and the simulation capture task environment, simply referred to as trajectory data. The trajectory data is stored in the form of a quadruple (s t ,a t ,r t ,s t+1 ), where: s t is the observation state information received from the perception module when the agent interacts with the simulation capture task environment for the t-th time, a t is the control instruction from the control module executed when the agent interacts with the simulation capture task environment for the t-th time, r t is the reward value feedback by the environment for the control instruction a t when the agent interacts with the simulation capture task environment for the t-th time, and s t+1 is the observation state information received from the perception module after the agent interacts with the simulation capture task environment for the t-th time and causes a change in the environmental state, and is also called the observation state information when the agent interacts with the simulation capture task environment for the (t + 1)-th time.

[0084] The MAPPO multi-agent reinforcement learning algorithm is used as the basic algorithm to train the reinforcement learning models corresponding to each pursuer agent in the multi-agent capture task simulation scenario, including:

[0085] S1: Based on the simulation scenario, obtain the position and speed information of the pursuer agent itself, the position and speed information of other pursuer agents, the position information of obstacles, and the position and speed information of the prey agent, and perform attention encoding based on the observation information of the pursuer agent itself to obtain the processed state features, and enter S2;

[0086] The attention observation encoding network consists of a fully connected layer with a ReLU activation function, takes the position information and speed information of the agent as input, and outputs a 128-dimensional embedding vector;

[0087] S2: Based on the position information of the obstacles, perform key point discretization processing to obtain the processed obstacle information, and enter S3;

[0088] The obstacle key point discretization process deals with irregular obstacles and non-convex obstacles in the environment. After processing, complex obstacles can be processed into simple geometric characteristics;

[0089] S3 calculates the artificial potential field based on the position and velocity information of the pursuer agent itself, the position and velocity information of the prey agent, and the processed obstacle information to obtain the resultant force acting on the pursuer agent, and proceeds to S4;

[0090] Artificial potential field calculation takes the observed information of the agent and the obstacle information after discretization of the obstacle key points as inputs, calculates the repulsive force between the pursuer and the obstacle, and the attractive force between the pursuer and the prey respectively, so as to calculate the resultant force acting on the pursuer;

[0091] The calculation process of the artificial potential field is as follows:

[0092] The position of the pursuer is p p , and the position of the prey target is p e . According to the discretization of key points to describe the irregular obstacle, the position of the key point of the obstacle closest to the pursuer is selected as p o . The repulsive force F r between the pursuer and the obstacle is mathematically expressed as:

[0093]

[0094] where ρ 0 is the influence range of the obstacle, and η is a proportionality coefficient.

[0095] The attractive force F a of the prey on the pursuer is a unit vector, that is, a vector with a direction of 1, towards the direction of the prey. The magnitude of this vector does not change with distance, is constant, and only provides a direction information towards the prey target.

[0096]

[0097] The resultant force F APF acting on the pursuer is: F APF = F r + F a .

[0098] S4 calculates based on the processed state features, the action direction obtained by inputting the processed obstacle information into the policy network, and combines the physical guidance direction provided by the resultant force acting on the pursuer agent to obtain the action execution instruction of the pursuer agent, and proceeds to S5;

[0099] The policy network takes the state features processed by the attention observation encoding network and the obstacle information after discretization of the key points as inputs, passes through 3 hidden layers of 128 dimensions, captures complex feature relationships, combines the resultant force output by the human potential field calculation module, and comprehensively outputs the agent's action;

[0100] The action a output by the policy network represents the action direction that the agent should take based on the learned policy, while the resultant force F acting on the pursuer APF provides a physical guiding direction a APF 。By synthesizing these two directions, the final agent action a is generated final 。

[0101] a final = α·a + β·a APF

[0102] where α is the first fusion coefficient and β is the second fusion coefficient, used to adjust the weights between the policy network and the artificial potential field

[0103] S5 executes the corresponding hunting action based on the action execution instruction of the pursuer agent. After obtaining the feedback reward value, it returns to S1 until the hunting is successful

[0104] The value network receives the state features processed by the attention observation encoding network and the obstacle information processed by the key point discretization processing module as inputs. After passing through 3 hidden layers with 128 dimensions each, it outputs a scalar, which is used to represent the expected return of the agent in the current state

[0105] The reward function of the pursuer agent equipped with the reinforcement learning model is as follows

[0106] For the pursuer, it can complete the encirclement and capture of the prey through cooperation, or use obstacles to drive the prey towards the corners or long sides of the obstacles, thereby restricting the escape route of the prey and reducing the number of required cooperative teammates. When the pursuer chases the prey near the obstacle, we can use the obstacle to complete the encirclement and capture

[0107] Assume the position of the pursuer is p, the position of the prey is e, and the position of the obstacle is o. Calculate the connection vector between the obstacle and the prey The vector between the pursuer and the prey Next, calculate and The dot product dot of product , if the dot product is positive, it means that the pursuer is on the connection line direction between the obstacle and the prey, or closer to the prey side. The reward R designed according to the dot product value pos_obstacle is

[0108]

[0109] On the other hand, the included angle between the moving direction of the pursuer and the direction from the prey to the obstacle. Calculate The included angle between is the direction of the pursuer’s movement. If θ is small enough, it means that the pursuer is forcing the prey to move towards the obstacle, and a reward R designed according to the angle is given. dir_obstacle :

[0110]

[0111] The cooperative encirclement reward is designed to encourage the pursuers to form an evenly distributed encirclement around the prey and ensure that the distance between the pursuer and the prey is within the encirclement radius l. The angle between adjacent pursuers is calculated and its standard deviation is calculated. The smaller the standard deviation, the higher the reward. For each pursuer p i Vector relative to the prey x i is the horizontal coordinate of the pursuer's position, x e is the horizontal coordinate of the prey position, y i is the ordinate of the pursuer’s position, y e is the ordinate of the prey’s position. i and p i+1 , calculate their vector and The angle θ between i :

[0112]

[0113] Calculate the angle θ between all adjacent pursuers i The smaller the standard deviation, the more evenly the pursuers are distributed, so the higher the reward. In the case of obstacles, we designed rewards that take advantage of obstacles. Due to the presence of obstacles, the movement range of the prey is limited, so we only need to calculate the uniform distribution of the pursuers within the remaining movement range. The angles between adjacent pursuers are still calculated, but only the standard deviation of the angles within the remaining movement range is considered. The smaller the standard deviation, the more evenly the pursuers are distributed, so the higher the reward.

[0114] When the pursuer surrounds the prey within the pursuit radius l and is evenly distributed around the prey, the capture is successful. When the capture is successful, that is, when the prey is surrounded by the pursuer and within the capture range, the pursuer is given a larger reward.

[0115]

[0116] In the formula, R capture The reward value for the pursuer agent to complete the encirclement; d i is the distance between the pursuer agent and the prey agent; l is the encirclement radius; std(θ) is the standard deviation of the angle distribution of the pursuer agent; threshold is the set threshold.

[0117] Finally, the trained multi-agent reinforcement learning model optimized based on obstacle terrain information is saved in the.pt format.

[0118] In this embodiment, two modes of capture scenarios are designed: normal mode and difficult mode. In the normal mode, there are square and long rectangular obstacles, while in the difficult mode, there are non-convex obstacles and a combined mode of other shaped obstacles. The number ratios of pursuers to prey are designed to be 3:1, 5:2, and 7:3 respectively, and training tests are carried out in obstacle environment modes with different difficulty levels.

[0119] As Figure 2 shown, in the obstacle environment of the normal mode, due to fewer obstacles and relatively simple tasks, all three methods completed the tasks well. However, it can be clearly seen that compared with the MAPPO method, the RL-MECA and RL-T-MECA methods have a faster rising reward function curve at the initial stage of the task. This is because the introduction of attention encoding enables the agents to perceive environmental information faster, thus exploring better strategies. In the obstacle environment of the difficult mode, as the complexity of the obstacles in the environment increases, the task difficulty significantly improves. The reward curve of the RL-T-MECA method converges to a higher level. This is because we introduced a reward strategy that utilizes obstacles to assist in capture, which can guide the agents to better utilize the terrain in complex environments, thus more efficiently completing the capture task and learning better strategies. Mean Episode Rewards reward; Pursuer:Prey = 7:3 in standard obstacle environment the ratio of pursuers to obstacles is 7:3 in the standard obstacle environment; Pursuer:Prey = 5:2 in difficult obstacle environment the ratio of pursuers to obstacles is 5:2 in the difficult obstacle environment; Updates updates.

[0120] As shown in Table 1, it shows the success rates and completion times of different numbers of agents completing the capture and obstacle avoidance tasks in the normal obstacle environment. When the number of agents is small, all three methods show high success rates and similar completion durations. However, as the number of agents increases, the performance advantage of the RL-T-MECA method gradually emerges. When the number of agents is 7 and 9 respectively, the success rate of the RL-T-MECA method is increased by 10.06% and 11.71% compared with MAPPO, and at the same time, the capture duration is shorter. P represents the proportion of pursuers; E represents the proportion of prey; S% represents the success rate; MEL represents the completion time.

[0121]

[0122] Table 1 Success rate and completion time of different methods in the ordinary obstacle environment

[0123] As shown in Table 2, in the difficult obstacle environment, even when the number of agents is only 4, the success rate of the RL-T-MECA method is increased by 6.53% compared with MAPPO, indicating that the agents learn to use obstacles to complete the encirclement task. As the number of agents increases, the complex environmental information further increases the difficulty for agents to explore effective strategies. When the number of agents is 7 and 9, the success rate of MAPPO drops significantly, while the success rate of the RL-MECA method is still 10% higher than that of MAPPO, which indicates that the use of attention-based observational information encoding improves the agents' understanding ability of the environment. In addition, the RL-T-MECA method has a 20% higher success rate than MAPPO with less completion time, further indicating that the agents in the RL-T-MECA method can analyze environmental information more effectively, use obstacles as terrain advantages, make more efficient execution decisions, and thus better complete the encirclement and obstacle avoidance task.

[0124]

[0125] Table 2 Success rate and completion time of different methods in the difficult obstacle environment

[0126] The following combines Figure 3 and Figure 4 to elaborate on this embodiment in detail.

[0127] In step 1, obtaining the observational information of each pursuer agent specifically includes:

[0128] Obtaining the observational information of each pursuer agent, including: the position and velocity information of the pursuer agent itself, the position and velocity information of other pursuer agents, the position information of obstacles, and the position and velocity information of the prey agent.

[0129] In step 2, based on the observational information of each pursuer agent and combined with a pre-trained reinforcement learning model for calculation, obtaining the action execution instructions of each pursuer agent specifically includes:

[0130] Inputting the observational information of each pursuer agent into the reinforcement learning model of each pursuer agent, and outputting the action execution instructions of each pursuer agent.

[0131] In step 3, performing corresponding encirclement actions based on the action execution instructions of each pursuer agent specifically includes:

[0132] Performing a hunting operation based on the action execution instructions of each pursuer agent.

[0133] In this embodiment, a reinforcement learning model is designed. The observation information of the agent is input into the attention encoding module to calculate the attention weights between the agent and other surrounding entities, enabling the agent to dynamically focus on the entities useful for the task and guiding the agent's action decision-making, thereby realizing a distributed cooperation strategy. The attention observation encoding ensures that the feature embedding of the agent is independent of the number of agents, effectively alleviating the problems of low utilization efficiency of large-scale high-dimensional data and the curse of dimensionality in multi-agent reinforcement learning, and improving the learning efficiency. By using the method based on key-point discretization, irregular and non-convex obstacles are converted into a simplified convex hull form, enabling the pursuers to more efficiently utilize the terrain advantages for encirclement. This method allows the pursuers to guide the prey into the corners or boundary areas of the map, or use the processed obstacles to trap the prey, thus maximizing the restriction of the prey's activity space and achieving an efficient encirclement task. On the other hand, for the obstacles simplified by the convex hull method, the artificial potential field method is used for obstacle avoidance, which can avoid local minima and prevent the agent from falling into an ideal state during movement.

[0134] Embodiment 2:

[0135] A multi-agent encirclement system optimized based on obstacle terrain information has a structure as Figure 5 shown, including:

[0136] An information acquisition module for acquiring the observation information of each pursuer agent;

[0137] A model calculation module for calculating, based on the observation information of each pursuer agent and in combination with a pre-trained reinforcement learning model, the action execution instructions of each pursuer agent;

[0138] An instruction execution module for performing corresponding encirclement actions based on the action execution instructions of each pursuer agent;

[0139] Among them, the reinforcement learning model is obtained by training the models of each pursuer agent using a multi-agent reinforcement learning algorithm based on the observation information of the pursuer agent.

[0140] It further includes a model training module; the model training module includes:

[0141] A simulation scenario construction sub-module for constructing a simulation scenario of a multi-agent encirclement task based on the pursuer agents, prey agents, and obstacles provided with a reinforcement learning model;

[0142] A reinforcement learning model training sub-module for training the reinforcement learning models corresponding to each pursuer agent based on the simulation scenario using a multi-agent reinforcement learning algorithm.

[0143] The reinforcement learning model training sub-module is specifically used for:

[0144] S1 obtains the position and velocity information of the pursuer agent itself, the position and velocity information of other pursuer agents, the position information of obstacles, and the position and velocity information of the prey agent based on the simulation scenario, and performs attention encoding based on the observation information of the pursuer agent itself to obtain the processed state features, and enters S2;

[0145] S2 performs key-point discretization processing based on the position information of obstacles to obtain the processed obstacle information, and enters S3;

[0146] S3 performs artificial potential field calculation based on the position and velocity information of the pursuer agent itself, the position and velocity information of the prey agent, and the processed obstacle information to obtain the resultant force received by the pursuer agent, and enters S4;

[0147] S4 calculates based on the action direction obtained by inputting the processed state features and the processed obstacle information into the policy network, and combines the physical guidance direction provided by the resultant force received by the pursuer agent to obtain the action execution instruction of the pursuer agent, and enters S5;

[0148] S5 executes the corresponding hunting actions based on the action execution instruction of the pursuer agent, returns to S1 after obtaining the feedback reward value until the hunting is successful.

[0149] In the S3 of the reinforcement learning model training sub-module, artificial potential field calculation is performed based on the position and velocity information of the pursuer agent itself, the position and velocity information of the prey agent, and the processed obstacle information to obtain the resultant force received by the pursuer agent, including:

[0150] Calculating based on the position of the pursuer agent and the key-point position of the obstacle agent closest to the pursuer agent to obtain the repulsive force between the pursuer agent and the obstacle;

[0151] Calculating based on the position of the prey agent and the position of the pursuer agent to obtain the attraction of the prey agent to the pursuer agent;

[0152] Calculating the resultant force received by the pursuer agent by adding the repulsive force between the pursuer agent and the obstacle and the attraction of the prey agent to the pursuer agent.

[0153] The calculation formula of the reward value in the reinforcement learning model training sub-module is as follows:

[0154]

[0155] Where R pos_obstacle is the reward value designed according to the dot product; dot product is the dot product;

[0156]

[0157] Wherein, R dir_obstacle is the reward value designed according to the included angle; cosθ is the included angle between the moving direction of the pursuer agent and the direction from the prey agent to the obstacle;

[0158]

[0159] Wherein, R capture is the reward value for the pursuer agent to complete the encirclement; d i is the distance between the pursuer agent and the prey agent; l is the encirclement radius; std(θ) is the standard deviation of the angular distribution of the pursuer agent; threshold is the set threshold.

[0160] The observed information of each pursuer agent includes: the position and speed information of the pursuer agent itself, the position and speed information of other pursuer agents, the position information of the obstacle, and the position and speed information of the prey agent.

[0161] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0162] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0163] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions in the process Figure 1one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.

[0164] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 the steps of the functions specified in one block or multiple blocks.

[0165] The above are only embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application are included within the scope of the claims of the present application pending approval.

Claims

1. A multi-agent capture method based on obstacle terrain information optimization, characterized in that: include: Obtain observation information of each pursuer agent; Based on the observation information of each pursuer agent and the pre-trained reinforcement learning model, calculation is performed to obtain the action execution instructions of each pursuer agent; Perform corresponding round-up actions based on the action execution instructions of each pursuer agent; The reinforcement learning model is obtained by training the model set for each pursuer agent using a multi-agent reinforcement learning algorithm based on the observation information of the pursuer agent.

2. The method according to claim 1, characterized in that The training process of the reinforcement learning model includes: Construct a simulation scenario of a multi-agent hunting task based on the pursuer agent, prey agent, and obstacles with a reinforcement learning model; Based on the simulation scenario, a multi-agent reinforcement learning algorithm is used to train the reinforcement learning model corresponding to each pursuer agent.

3. The method according to claim 2, characterized in that The method of training the reinforcement learning model corresponding to each pursuer agent by using a multi-agent reinforcement learning algorithm based on the simulation scenario includes: S1 obtains the position and speed information of the pursuer agent, the position and speed information of other pursuer agents, the position information of obstacles, and the position and speed information of the prey agent based on the simulation scene, and performs attention encoding based on the observation information of the pursuer agent to obtain the processed state features, and enters S2; S2 performs key point discretization processing based on the position information of the obstacle, obtains the processed obstacle information, and enters S3; S3 calculates the artificial potential field based on the position and speed information of the pursuer agent, the position and speed information of the prey agent, and the processed obstacle information to obtain the resultant force on the pursuer agent, and then enters S4; S4 calculates the action direction obtained by inputting the processed state features and the processed obstacle information into the strategy network, and combines the physical guidance direction provided by the combined force of the pursuer agent to obtain the action execution instruction of the pursuer agent, and enters S5; S5 executes the corresponding hunting action based on the action execution instruction of the pursuer intelligent agent, and returns to S1 after obtaining the feedback reward value until the hunting is successful.

4. The method according to claim 3, characterized in that The S3 performs artificial potential field calculation based on the position and speed information of the pursuer agent, the position and speed information of the prey agent, and the processed obstacle information to obtain the resultant force on the pursuer agent, including: Based on the position of the pursuer agent and the key point position of the obstacle agent closest to the pursuer agent, the repulsive force between the pursuer agent and the obstacle is calculated; Based on the position of the prey agent and the position of the pursuer agent, the attraction of the prey agent to the pursuer agent is calculated; Based on the addition of the repulsive force between the pursuer agent and the obstacle and the attractive force of the prey agent on the pursuer agent, the resultant force on the pursuer agent is calculated.

5. The method according to claim 3, characterized in that: The calculation formula of the reward value is as follows: In the formula, R pos_obstacle is the reward value designed according to the dot product; dot product is the dot product; In the formula, R dir_obstacle is the reward value designed according to the angle; cosθ is the angle between the moving direction of the pursuer agent and the direction from the prey agent to the obstacle; In the formula, R capture The reward value for the pursuer agent to complete the encirclement; d i is the distance between the pursuer agent and the prey agent; l is the encirclement radius; std(θ) is the standard deviation of the angle distribution of the pursuer agent; threshold is the set threshold.

6. The method according to claim 1, characterized in that The observation information of each pursuer agent includes: the position and speed information of the pursuer agent itself, the position and speed information of other pursuer agents, the position information of obstacles, and the position and speed information of prey agents.

7. A multi-agent capture system based on obstacle terrain information optimization, characterized in that: include: An information acquisition module, used to obtain observation information of each pursuer agent; A model calculation module, used to calculate based on the observation information of each pursuer intelligent agent in combination with a pre-trained reinforcement learning model to obtain action execution instructions for each pursuer intelligent agent; An instruction execution module, used for performing corresponding round-up actions based on the action execution instructions of each pursuer agent; The reinforcement learning model is obtained by training the model set for each pursuer agent using a multi-agent reinforcement learning algorithm based on the observation information of the pursuer agent.

8. The system according to claim 7, characterized in that It also includes a model training module; the model training module includes: A simulation scenario construction submodule is used to construct a simulation scenario of a multi-agent hunting task based on a pursuer agent, a prey agent, and obstacles with a reinforcement learning model; The reinforcement learning model training submodule is used to train the reinforcement learning model corresponding to each pursuer agent using a multi-agent reinforcement learning algorithm based on the simulation scenario.

9. The system according to claim 8, characterized in that The reinforcement learning model training submodule is specifically used for: S1 obtains the position and speed information of the pursuer agent, the position and speed information of other pursuer agents, the position information of obstacles, and the position and speed information of the prey agent based on the simulation scene, and performs attention encoding based on the observation information of the pursuer agent itself to obtain the processed state features, and enters S2; S2 performs key point discretization processing based on the position information of the obstacle, obtains the processed obstacle information, and enters S3; S3 calculates the artificial potential field based on the position and speed information of the pursuer agent, the position and speed information of the prey agent, and the processed obstacle information to obtain the resultant force on the pursuer agent, and then enters S4; S4 calculates the action direction obtained by inputting the processed state features and the processed obstacle information into the strategy network, and combines the physical guidance direction provided by the combined force of the pursuer agent to obtain the action execution instruction of the pursuer agent, and enters S5; S5 executes the corresponding hunting action based on the action execution instruction of the pursuer intelligent agent, and returns to S1 after obtaining the feedback reward value until the hunting is successful.

10. The system according to claim 8, characterized in that In the reinforcement learning model training submodule, S3 performs artificial potential field calculation based on the position and speed information of the pursuer agent, the position and speed information of the prey agent, and the processed obstacle information to obtain the resultant force on the pursuer agent, including: Based on the position of the pursuer agent and the key point position of the obstacle agent closest to the pursuer agent, the repulsive force between the pursuer agent and the obstacle is calculated; Based on the position of the prey agent and the position of the pursuer agent, the attraction of the prey agent to the pursuer agent is calculated; Based on the addition of the repulsive force between the pursuer agent and the obstacle and the attractive force of the prey agent on the pursuer agent, the resultant force on the pursuer agent is calculated.

11. The system according to claim 8, characterized in that The calculation formula of the reward value in the reinforcement learning model training submodule is as follows: In the formula, R pos_obstacle is the reward value designed according to the dot product; dot product is the dot product; In the formula, R dir_obstacle is the reward value designed according to the angle; cosθ is the angle between the moving direction of the pursuer agent and the direction from the prey agent to the obstacle; In the formula, R capture The reward value for the pursuer agent to complete the encirclement; d i is the distance between the pursuer agent and the prey agent; l is the encirclement radius; std(θ) is the standard deviation of the angle distribution of the pursuer agent; threshold is the set threshold.

12. The system according to claim 7, characterized in that The observation information of each pursuer agent includes: the position and speed information of the pursuer agent itself, the position and speed information of other pursuer agents, the position information of obstacles, and the position and speed information of prey agents.