Multi-robot multi-target large-scale pursuit escape method based on reinforcement learning
By employing reinforcement learning and self-attention mechanisms, a multi-robot pursuit and escape method is developed, addressing issues such as uneven target allocation, insufficient obstacle avoidance safety, and robot collisions in large-scale scenarios within multi-target tasks. This approach enables more efficient and safer multi-robot pursuit and escape missions.
Patent Information
- Application Number
- CN202511003165.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-17
AI Technical Summary
Existing multi-robot pursuit-escape research has shortcomings in multi-target tasks, obstacle avoidance safety and large-scale application. It is difficult to adapt to complex practical scenarios, lacks efficient target allocation and coordination strategies, is not safe enough, and collisions between robots are frequent in large-scale environments.
A large-scale pursuit and escape method based on reinforcement learning for multiple robots and multiple targets is adopted. The model is built using the self-attention mechanism of the Transformer architecture, the robot's perception radius and activity area are set, autonomous obstacle avoidance and teammate cooperation mechanisms are introduced, and intelligent target selection and allocation are carried out by using a partially observable Markov decision process and a multi-level reward mechanism. The training paradigm of centralized training and decentralized execution is adopted to gradually increase the difficulty of the task.
It improves the efficiency and safety of multi-robot systems in multi-objective tasks, reduces the risk of collisions between robots and between robots and obstacles, enhances adaptability and collaboration in large-scale scenarios, and ensures the efficient completion of tasks.
Smart Images

Figure CN120791759A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of multi-target pursuit and escape, and particularly relates to a multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning. BACKGROUND
[0002] The multi-robot pursuit-escape problem is a complex and challenging research field that combines robotics, control theory, game theory, optimization algorithms, graph theory, and other disciplines. It studies how a group of pursuers cooperates to capture one or more evaders in a shared environment, while the evaders try to avoid being captured.
[0003] Current similar technical solutions in the multi-robot pursuit-escape field have three defects, which limit their applicability in complex real-world scenarios and large-scale, multi-target, and safety performance.
[0004] First, the scene setting is simplified and difficult to adapt to multi-target tasks. Existing researches mostly focus on the scenario of multiple pursuers surrounding a single evader, assuming that the pursuers are slower and the evaders are faster. However, actual applications often involve multiple evading targets, and existing solutions lack efficient target allocation and coordination strategies, with low robustness, making it difficult to be directly used for multi-target tasks such as UAV group hunting.
[0005] Second, the safety consideration is insufficient, and there is a lack of effective obstacle avoidance mechanism. Many studies assume an ideal environment without considering obstacles and lack obstacle avoidance strategies. When pursuing, the relative position with teammates is ignored, which easily leads to collisions between robots. Due to the lack of safety constraints, it is difficult to deploy in scenarios with high safety requirements such as urban environments. Improving obstacle avoidance capability and reducing collisions are optimization priorities.
[0006] Third, the applicability in large-scale scenarios is insufficient. Existing researches mostly focus on small-scale scenarios, while in large-scale tasks, there are problems such as imperfect target selection mechanism (easy to repeat or miss pursuit), lack of understanding of environmental relationships (simple decision-making method without fully utilizing environmental information to plan a globally optimal strategy), and high-frequency collisions between robots (difficult to coordinate motion), and related exploration is still insufficient.
[0007] Therefore, the existing research on multi-robot pursuit-escape problem still has many limitations, especially in multi-target tasks, obstacle avoidance safety, and large-scale applications. Therefore, in view of these defects, a new solution that can balance task efficiency, safety, and large-scale adaptability is needed to improve the feasibility and robustness of multi-robot systems in practical applications. SUMMARY
[0008] The application aims at providing a multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning to solve the above problems.
[0009] The technical scheme adopted by the application is as follows: a multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning, the pursuit and escape method comprising the following steps: Initializing modeling, modeling key variables and tasks, and initializing variables; Pursuit and escape modeling, setting all robots in a fixed perception radius, and dividing the robots into pursuers and escapees; dynamically adjusting the key variable weight through the self-attention mechanism of the Transformer architecture, and performing pursuit and escape modeling; Calculating the distance between the pursuer and the escapee, and determining whether the pursuer captures the escapee.
[0010] Further, the key variables include robot size, obstacle position, robot activity area and its boundary; the tasks include the pursuit of the pursuer to the escapee; Modeling each robot as an individual with uniform size and subject to the same kinematic constraints; initializing the random distribution of obstacle positions, and the obstacle positions are subject to dynamic interference; setting the activity area boundary of the robot as a soft boundary, and the activity space as a two-dimensional bounded space; Setting the pursuit of the pursuer to the escapee to a set threshold, and the task is completed.
[0011] Further, the pursuer makes decisions through a reinforcement learning strategy, and the escapee uses a predefined artificial potential field (APF) strategy to escape; and the pursuer and the escapee do not communicate with each other and do not master each other's strategy information.
[0012] Further, the pursuer is set as N, the escapee is set as M, and the judgment method of the pursuer to the escapee is: the pursuer approaches the escapee to a set threshold distance , while avoiding collision with other robots or obstacles; When the Euclidean distance between an escapee and at least one pursuer is , the pursuer successfully pursues the escapee : ; wherein, is the position of the th pursuer, is the position of the th escapee, is the position of the pursuer , for the position of the evader, a set threshold distance for the success of the pursuit.
[0013] Further, since there is no communication between the pursuers and the evader, the pursuit evasion problem is modeled as a partially observable Markov decision process (POMDP), which is defined as follows: ; where, is the state space, which describes the state information of the pursuers and the evader; is the action space, which defines the available actions for the pursuers, including acceleration and angular velocity; is the state transition probability, which describes the dynamic changes in the state of the robots over time; is the reward function, which encourages the robots to complete the pursuit task quickly and avoid collisions; is the observation space, which defines the local environmental information that each pursuer can observe; is the observation probability distribution, which describes the observation probability of the pursuers in different states; is the discount factor, which controls the weight of future rewards.
[0014] Further, at each time step , each pursuer receives a local observation and selects an action according to the current policy, the environment updates the states of all pursuers and the evader according to the state transition probability , and the pursuers receive a reward signal to optimize the long-term cumulative reward, and the goal of the optimal policy is to maximize the discounted cumulative reward : ; where, is the mathematical expectation, which is used to calculate the expected value of the discounted cumulative reward, and is the core indicator of measuring the long-term strategy income of the pursuers, is the policy, which is the decision-making rule obtained by the pursuers based on reinforcement learning, and determines the logic of selecting actions under different local observations, is the state at time step t, including the position, velocity, and other dynamic information of the pursuers and the evader, as well as the distribution of obstacles and environmental boundaries, which is the comprehensive representation of the environment at that moment.
[0015] Furthermore, at each time step In each pursuer Local observation of It consists of the following information: ; in, The status information of the pursuer itself; The observation information of other pursuers within the sensing range, i.e. the pursuer's teammates; Observe information for escapees; Obstacle information; The action space of each pursuer is defined as: ; in, is the linear acceleration, used to control the speed change, is the angular velocity, which is used to control the direction change.
[0016] Furthermore, the pursuit-escape method adopts a training paradigm. During the training phase, the pursuer accesses global information, including the positions and velocities of all pursuers and escapers, to optimize the robot's behavioral strategy during training and learn more efficient collaborative pursuit methods. During the execution phase, global information is removed, and each pursuer makes decisions based only on its own local observations, ensuring that the pursuer can still complete efficient multi-target pursuit tasks under non-communication conditions.
[0017] Furthermore, the reward function Including safety reward function , target allocation reward function and the task completion reward function ; Safety Reward Function Specifically include: ; ; ; in, is the minimum Euclidean distance between the pursuer and all observed teammates, escapers, and obstacles; is the distance from the pursuer to the nearest escapee; is the threshold distance; It is a safety threshold used to reduce the risk of collision; Target allocation reward function is the global reward mechanism, for each escaper , calculate the number of pursuers around it , the number of pursuers within the relative distance is: ; If an escapee is chased by multiple pursuers , and the number of pursuers of another escapee is less than 3, all pursuers will be subject to an additional global penalty of -10; if a pursuer crosses the border, an additional penalty of -5 will be imposed; Task completion reward function : ; When an escapee is successfully chased by at least one pursuer, the task completion reward is assigned to all pursuers involved in the target chase; At the same time, in order to encourage faster completion of the pursuit task, a time penalty : ; The time penalty takes effect when the pursuer adopts a conservative strategy.
[0018] Further, in order to improve the generalization ability of the model and speed up the learning speed, in the training stage, a curriculum learning method is adopted, three learning processes are set, and the task difficulty is gradually increased, and the learning process includes the following three: In the initial stage, a simple environment is set, including a small number of obstacles and robots, and a small perception radius; In the middle stage, the number of robots is gradually increased, and complex environmental disturbances are introduced; In the advanced stage, a complete large-scale environment is used, and more intelligent pursuit and escape methods are introduced.
[0019] To sum up, due to the adoption of the above technical solutions, the beneficial effects of the present application are: 1. The present application dynamically adjusts the key variable weight through the self-attention mechanism of the Transformer architecture, models the pursuit and escape, and makes the target selection of the robot more intelligent, so that the problem of multiple robots excessively concentrating on chasing the same target and other targets without robot tracking is avoided; thereby improving the efficiency of the pursuit.
[0020] 2、The three-layer reward mechanism is set, the safety reward avoids robot collision, and the stability of the multi-robot system is improved; the target allocation reward ensures that the robot reasonably allocates the pursuit target and improves the cooperation ability; the task completion reward encourages the robot to complete the task as soon as possible, and improves the overall capture efficiency. Through the three-layer reward mechanism, the robot learns a safer motion mode, avoids collision between robots or between robots and obstacles, while ensuring that the robot can efficiently allocate tasks, reduce task delays, and improve task completion efficiency.
[0021] 3、The method can be adapted to large-scale multi-robot tasks, and through the Transformer architecture and the training paradigm of centralized training and decentralized execution, the relationship between robots can be efficiently calculated, and the robots can still efficiently cooperate when expanded to a large scale, and the method can also adapt to more complex real scenes. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 The method flowchart of the present application is shown in the figure; Figure 2 The multi-robot multi-target pursuit and escape scene of the present application is shown in the figure; Figure 3 The pursuit-escape local situation map of embodiment 8 of the present application is shown in the figure. DETAILED DESCRIPTION
[0023] The present application will be described in detail below with reference to the accompanying drawings.
[0024] In order to make the purpose, technical scheme and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0025] The current market has the following defects in the same technical scheme, which limits its applicability in complex real scenes and affects its performance in large-scale, multi-target and safety.
[0026] (1) The existing research scene setting is relatively simple, which is difficult to adapt to multi-target task demand: at present, the research of multi-robot pursuit-escape problem is mostly concentrated in the simplified scene setting of multiple pursuers pursuing a single evader. This setting usually assumes that multiple slow-speed pursuers surround a fast-speed evader. However, in practical applications, multiple evading targets are often involved. The existing research fails to fully consider this real scene, resulting in that its scheme is difficult to be directly applied in multi-target tasks (such as unmanned aerial vehicle group encirclement, unmanned vehicle cooperative interception, etc.). In addition, in the absence of efficient target allocation and cooperation strategy, the robustness of the existing scheme is low, which is difficult to meet the demand of large-scale complex tasks.
[0027] (2) Lack of safety considerations and effective obstacle avoidance mechanisms: Safety is one of the key factors for the application of multi-robot systems in real-world environments. However, existing research lacks sufficient consideration of safety, mainly in the following aspects: 1. Many studies assume an ideal environment for robot operation and do not consider obstacles in complex environments, thus lacking effective obstacle avoidance strategies. In real-world scenarios, robots need to be able to autonomously identify and avoid obstacles to ensure smooth task execution and prevent equipment damage. 2. Existing strategies often focus only on how to optimally approach the target when pursuing, ignoring the relative position relationship with teammates, leading to a high probability of collision between robots, affecting task completion efficiency and even causing robot damage. 3. Without safety constraints, existing methods are difficult to deploy in real-world environments, especially in applications requiring high safety requirements such as urban environments, marine unmanned ship formations, and complex industrial scenarios. Therefore, how to improve obstacle avoidance ability while ensuring pursuit efficiency and reducing collisions between robots and obstacles is an important direction for existing schemes to optimize.
[0028] (3) Insufficient applicability of existing research in large-scale scenarios: Currently, multi-robot pursuit-escape research mostly focuses on small-scale scenarios, involving only a small number of pursuers and escape targets. However, in large-scale tasks such as security monitoring, drone swarm hunting, multi-vehicle cooperative interception, etc., robots need to face more complex interaction scenarios and make reasonable decisions in large-scale environments. Existing research has not fully explored this aspect, mainly with the following problems: 1. Incomplete target selection mechanism: When the number of targets is large, pursuers need to have a reasonable target allocation strategy to ensure the efficiency of the hunting. However, existing research lacks a clear target selection mechanism, leading to the situation of repeated pursuit of targets or missing targets, thus reducing the efficiency of task execution. 2. Lack of understanding of environmental relationships: In large-scale environments, the interaction relationships between targets, obstacles, and pursuers are more complex. Existing research often uses simple heuristic methods for decision-making, failing to fully utilize environmental information for global optimal strategy planning, thus affecting the overall performance of the system. High-frequency collision problem between robots: In large-scale scenarios, the number of robots increases and the interaction complexity rises, making it difficult for existing schemes to effectively coordinate the motion of multiple robots, leading to a significant increase in collision frequency, which in turn affects the smooth execution of tasks.
[0029] Therefore, this paper addresses the limitations of existing approaches to the multi-robot pursuit-escape problem and proposes a more efficient, secure, and large-scale solution: a reinforcement learning-based multi-robot, multi-target, large-scale pursuit-escape method. By optimizing target selection, obstacle avoidance strategies, and adaptability to large-scale tasks, this approach improves the collaborative capabilities, task completion efficiency, and practical feasibility of multi-robot systems.
[0030] The present invention optimizes the following problems existing in the existing multi-robot pursuit-escape scheme: 1. Difficulty in efficiently handling multi-target tasks: Current technologies mostly focus on single-target pursuits, ignoring the multi-target scenarios common in practical applications. This lack of a rational target allocation mechanism leads to inefficient task execution. This invention proposes an intelligent target selection and allocation strategy that enables pursuers to rationally allocate tasks and improve overall capture efficiency.
[0031] 2. Lack of effective obstacle avoidance strategies and insufficient safety: Existing technologies pay little attention to robot obstacle avoidance, making robots prone to collisions in complex environments and impacting mission completion efficiency. This invention significantly reduces collision risks and improves system stability by strengthening safety constraints and introducing autonomous obstacle avoidance and collaborative obstacle avoidance mechanisms with teammates.
[0032] 3. Poor adaptability to large-scale scenarios: Large-scale tasks involve numerous robots and targets. Existing solutions struggle to coordinate interactions between multiple robots, leading to frequent conflicts and reduced task execution efficiency. This invention, through global information modeling and a multi-level collaboration strategy, improves the system's adaptability to large-scale tasks, making it more feasible in real-world scenarios.
[0033] Example 1 like Figure 1 As shown, one embodiment of the present invention is a multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning, the pursuit and escape method includes the following: Initialize modeling, model key variables and tasks, and initialize variables; this includes setting scenarios and pursuit tasks. The scenario setting includes the location of obstacles and specific pursuit tasks. Pursuit and escape modeling: all robots are set within a fixed perception radius and divided into pursuers and escapers; Calculate the distance between the pursuer and the escapee and determine whether the pursuer has captured the escapee.
[0034] Each robot can observe the surrounding environment within a fixed perception radius. Based on the classification of pursuer and escaper robots, the pursuer confirms the pursuit target, calculates the distance between the pursuer and the escaper, and determines whether the escaper can be captured.
[0035] The embodiment sets multiple pursuers to track multiple escape targets in an uncertain environment, and realizes successful pursuit within a certain distance threshold, so as to overcome the limitations in the prior art of multi-robot pursuit-escape method, provide a more intelligent, higher safety and large-scale task applicable solution, and provide technical support for group hunting of unmanned aerial vehicles, unmanned ships and the like, intelligent security, automatic driving cooperative interception and the like.
[0036] Embodiment 2 Another embodiment of the present application is that the key variables include robot size, obstacle position, robot active area and its boundary; the task includes that the pursuer pursues the escapee; Each robot is modeled as an individual with uniform size and is subject to the same kinematic constraints; the obstacle position is initialized to be randomly distributed and is subject to dynamic disturbance; the boundary of the active area of the robot is set as a soft boundary, and the active space is a two-dimensional bounded space; The pursuer pursues the escapee to a set threshold, and the task is completed.
[0037] The robot itself is taken as the center point, and the sensing radius is set as the observable environment of the robot, so that the size of the robot itself needs to be modeled; the obstacle position is randomly set and can be adjusted at any time; the robot runs in a two-dimensional bounded space (2D bounded space), which is set by a soft boundary (Soft Boundary), that is, the robot is not determined to be collided when it is out of the boundary, but is subject to a certain penalty. Obstacles are randomly distributed in the environment and are subject to dynamic disturbance, for example, the influence of wind field or water flow vortex.
[0038] The task of the pursuer pursuing the escapee is adjusted at any time according to the actual situation, including the number of pursuers in the pursuit task of one escapee.
[0039] The embodiment sets key variables and tasks, and models the initial variables, so that the robot can learn.
[0040] Embodiment 3 Another embodiment of the present application is that the pursuer makes decisions by using a reinforcement learning strategy, and the escapee uses a predefined artificial potential field (APF) strategy to escape; and the pursuer and the escapee do not communicate with each other and do not master the strategy information of the other party.
[0041] The pursuer intelligently learns to select the pursuit target within its own sensing radius, and the pursuers can dynamically allocate tasks among each other; The escapee uses an APF strategy to guide the robot to avoid obstacles and reach the target point by simulating the interaction of gravity and repulsion. Specifically, the APF sets the pursuer as a gravity source and the obstacle as a repulsion source, and the robot moves in the virtual force field by force balance. The gravity drives the robot to move towards the target point, and the repulsion prevents it from colliding with the obstacle. This model can simplify the complex environment into a force interaction problem.
[0042] The adaptive target selection mechanism is adopted in the embodiment to ensure that the pursuer can dynamically adjust the target allocation, avoid repeated pursuit, and improve the overall efficiency of the multi-target task. At the same time, the environmental information modeling (such as obstacle position, robot activity area and boundary, etc.) is combined to enable the robot to perceive the complexity of the environment, reasonably formulate the task strategy, and enhance the adaptability of the system.
[0043] Embodiment 4 As shown in Figure 2 Another embodiment of the present application is to set the number of pursuers to N and the number of escapees to M. The judgment method of the pursuer chasing the escapee is as follows: the pursuer approaches the escapee to a set threshold distance , while avoiding collision with other robots or obstacles; When the Euclidean distance between an escapee and at least one pursuer satisfies the following conditions, the pursuer successfully chases the escapee : ; Wherein, is the i-th pursuer, is the i-th escapee, is the position of the pursuer, is the position of the escapee, is the set threshold distance for successful pursuit. When a large-scale pursuit scene is involved, in order to avoid collision, the condition for successful pursuit is set as the Euclidean distance , which can ensure the safety between robots and achieve the statistics of the task. The specific value of needs to be set according to the specific robot and the size of the scene.
[0044] Embodiment 5 Another embodiment of the present application is that since there is no communication between the pursuer and the escapee, each robot only relies on local observation information for decision-making. Therefore, the method of chasing the escapee is modeled as a partially observable Markov decision process, which is defined as follows:
[0045] Embodiment 5 Another embodiment of the present application is that since there is no communication between the pursuer and the escapee, each robot only relies on local observation information for decision-making. Therefore, the method of chasing the escapee is modeled as a partially observable Markov decision process, which is defined as follows: ; in, is the state space, which is used to describe the state information of the pursuer and the escaper; is the action space, which is used to define the pursuer's optional actions, including acceleration and angular velocity; is the state transition probability, which is used to describe the dynamic changes of the robot state over time; is a reward function used to motivate the robot to quickly complete the pursuit task and avoid collisions; The observation space is used to define the local environment information that each pursuer can observe; is the observation probability distribution, which describes the observation probability of the pursuer in different states; is a discount factor used to control the weight of future rewards.
[0046] At each time step , set each pursuer Accept a local observation , and select an action based on the current strategy , the environment is based on the state transition probability Update the status of all pursuers and escapers, and the pursuers receive reward signals To optimize the long-term cumulative return, the optimal strategy aims to maximize the discounted cumulative reward : ; in, It is the mathematical expectation used to calculate the expected value of the discounted cumulative reward and is the core indicator for measuring the long-term strategic benefits of the pursuer. Strategy is the decision rule obtained by the pursuer based on reinforcement learning training, which determines the logic of its action selection under different local observations. is the state at time step t, including dynamic information such as the position and velocity of the pursuer and the escaper, as well as static constraints such as obstacle distribution and environmental boundaries. It is a comprehensive representation of the environment at that moment.
[0047] At each time step In each pursuer Local observation of It consists of the following information: ; in, The status information of the pursuer itself; to perceive the observation information of other pursuers, i.e., pursuer teammates, within the perception range; to perceive the observation information of the evader; to perceive the obstacle information; The action space of each pursuer is defined as: ; wherein, is a linear acceleration used to control the change in velocity, is an angular velocity used to control the change in direction.
[0048] The pursuit and evasion method adopts a training paradigm (centralized training and decentralized execution) to improve the collaboration ability of robots while ensuring autonomy and real-time performance in actual execution. In the training phase, the pursuers access global information, including the positions and velocities of all pursuers and evaders, optimize the behavior strategy of robots during the training process, and learn more efficient collaborative pursuit methods. In the execution phase, global information is removed, and each pursuer makes decisions based only on its own local observations, ensuring that pursuers can still complete efficient multi-target pursuit tasks without communication. Through the training paradigm, pursuers can learn complex and adaptive strategies in the training phase, while still having the ability to make decentralized and autonomous decisions in the execution phase, ensuring the real-time performance and robustness of the system.
[0049] To encourage pursuers to avoid collisions (including with obstacles and other robots) while pursuing the evader, the embodiment proposes a distance-based reward function, the reward function includes a safety reward function , a target assignment reward function , and a task completion reward function ; The safety reward function specifically includes: ; ; ; wherein, is the minimum Euclidean distance between the pursuer and all observed teammates, evaders, and obstacles, and when is a collision penalty; is the distance from the pursuer to the nearest evader; is a threshold distance; is a safety threshold used to reduce the risk of collision; The safety reward function encourages the pursuers to keep a reasonable safety distance while approaching the targets, avoiding collisions between robots or with obstacles.
[0050] In the multi-robot pursuit task, the uneven distribution of targets (e.g., multiple pursuers chasing the same target while other targets are not being tracked) can reduce the overall pursuit efficiency. Therefore, the embodiment proposes a target distribution reward function , for the global reward mechanism to balance target selection and ensure effective pursuit strategies.
[0051] For each evader , the number of pursuers around it is calculated , then the number of pursuers within the relative distance is: ; If there is an evader being chased by multiple pursuers and the number of pursuers of another evader is less than 3, then all pursuers will be subject to an additional global penalty of -10; if a pursuer crosses the boundary, it will be subject to an additional penalty of -5; Task completion reward function is based on the reward for successful pursuit: ; When an evader is successfully pursued by at least one pursuer, i.e., , the task completion reward is assigned to all pursuers involved in the target pursuit; At the same time, in order to encourage faster completion of the pursuit task, a time penalty is imposed on all pursuers at each time step: ; The time penalty takes effect when the pursuers adopt a conservative strategy, i.e., negative pursuit, etc.
[0052] Embodiment 6 Another embodiment of the present invention is to use a curriculum learning method in the training phase, setting three learning processes to gradually increase the task difficulty, and the learning processes include the following three: In the initial stage, a simple environment is set up, including a small number of obstacles and robots, and a small perception radius; In the middle stage, the number of robots is gradually increased, and complex environmental disturbances are introduced, including wind fields, vortices, etc. In the advanced stage, a complete large-scale environment is used, and more intelligent pursuit and escape methods are introduced.
[0053] Through this training method, the model can gradually adapt to complex environments and learn more efficient multi-target pursuit strategies.
[0054] Embodiment 7 Another embodiment of the present application is a multi-robot multi-target large-scale pursuit and escape system based on reinforcement learning, which includes six modules. In the training process, the modules are optimized by the centralized training, decentralized execution (CTDE) strategy to form an end-to-end reinforcement learning framework. In the execution process, each robot independently runs the aforementioned modules to complete decentralized decision-making. The modules specifically include the following: The task modeling module models key variables such as robot state (including pursuers and escapees), environmental constraints (obstacles, boundaries), and action space, and defines the task goal (pursuers pursuing escapees to within a threshold distance) and sets the conditions that need to be met for task completion.
[0055] The task modeling module provides a basic mathematical framework for the reinforcement learning modeling module, defines the state space, observation space, and action space, etc.
[0056] The reinforcement learning modeling module converts the mathematical definitions of the task modeling module into a reinforcement learning framework POMDP, specifically including designing state transitions, decision-making strategies, reward functions, etc.; and using centralized training CTDE for optimization to enable the robot to learn pursuit strategies.
[0057] The reinforcement learning modeling module constructs training data and transmits feedback to the partially observable Markov decision process during the training process to reinforce good behavior patterns. After training is complete, it provides a strategy network for the observation processing, relationship extraction, target selection, and decision output modules.
[0058] The observation processing module receives local observation data from the robot and converts it into a high-dimensional feature vector. It independently encodes different types of observation information (self-state, teammate state, escapee state, obstacle state) and extracts useful feature representations through a multi-layer perception MLP embedding layer.
[0059] The observation processing module receives state definitions from the reinforcement learning modeling module and encodes the original observation data. It outputs the feature vector to the relationship extraction module for subsequent interaction information modeling.
[0060] The relationship extraction module uses a self-attention mechanism (Self-Attention) to calculate the interaction between the robot and surrounding entities (escapees, teammates, obstacles). It uses a deep learning model based on the Transformer architecture to calculate the relevance of the robot and other entities in the environment, generating an interaction relationship matrix for subsequent decision-making.
[0061] The relationship extraction module receives the feature representation from the observation processing module and calculates the interaction weight between entities, and delivers the interaction information to the target selection module for selecting the optimal pursuit target.
[0062] The target selection module selects the optimal target from multiple possible pursuit targets, calculates the target priority through attention mechanism, ensures that the pursuer focuses on the most critical escapee, combines with the teammate information to avoid excessive pursuit of the same target, and improves the global efficiency; the global environment information is extracted through max pooling and spliced with the individual state of the pursuer to form the final target representation.
[0063] The target selection module calculates the target priority through attention mechanism, ensures that the pursuer focuses on the most critical escapee, combines with the teammate information to avoid excessive pursuit of the same target, and improves the global efficiency.
[0064] The decision output module receives the target selection result and generates the actual executed action, adopts the Value-based reinforcement learning algorithm to estimate the expected value of each action, selects the optimal action output acceleration and angular velocity through the greedy strategy to control the motion of the robot.
[0065] The decision output module receives the target information from the target selection module, determines the target direction, and outputs to the physical environment for the robot to execute the corresponding motion.
[0066] The Transformer self-attention mechanism is used in this embodiment, which can dynamically calculate the relationship weight between the robot and other entities, thereby realizing intelligent target selection and improving the pursuit efficiency. At the same time, the centralized training and decentralized execution (CTDE) optimization strategy is adopted, so that the robot can access global information during training to learn a better cooperation strategy; each robot only relies on local information during execution to avoid information dependency problems, making the system scalable to a larger scale. A multi-layer reward mechanism is set up to guide the robot to learn a safer and more efficient pursuit strategy, including safety reward (avoiding collision), target allocation reward (balanced target selection), task completion reward (encouraging to complete the task as soon as possible), etc. Course learning is adopted to improve the generalization ability of the model, and the complexity of the environment is gradually increased during the training process, so that the robot can maintain high decision-making ability in different scale task scenarios.
[0067] This embodiment adopts a hierarchical architecture design, making the data flow more efficient, reducing unnecessary computational overhead, and improving the interpretability and stability of the system.
[0068] Information layering processing: each module focuses on a specific task, making the data flow clearer and improving processing efficiency. Efficient data transmission: the data flow from observation to decision is end-to-end, reducing the information redundancy problem commonly seen in traditional multi-robot tasks.
[0069] Example 8 Another embodiment of the present application is to set a scene, in a bounded two-dimensional plane environment, there are randomly distributed static obstacles and unknown environmental disturbances (such as dynamic disturbances such as ocean currents, wind field).
[0070] And set the environment: 60 pursuers (Pursuers), the initial position is randomly distributed; 20 evaders (Evaders), the initial position is randomly distributed, and an autonomous escape strategy is adopted.
[0071] The pursuer needs to make decisions only by local observation under the condition of no inter-robot communication, to ensure efficient pursuit of the evader, and to avoid collision during the pursuit process.
[0072] The success of the task is determined by the criterion: when the Euclidean distance between the pursuer and the target evader is less than the set threshold, it is considered that the evader is successfully pursued.
[0073] As shown in Figure 3 The actual running local situation of the method is: (a) Initial state: 60 pursuers start from random positions, 20 evaders in the environment are scattered in different areas, and some robots are affected by obstacles or environmental disturbances; (b) Pursuit process: the pursuer starts to disperse the target, dynamically adjusts the pursuit strategy, and maintains an efficient target allocation mechanism while avoiding obstacles; (c) Task completion: all evaders are successfully pursued, the task completion time is significantly shorter than the traditional method, and the number of robot collisions is significantly reduced.
[0074] The 60 pursuers make autonomous decisions, reasonable division of labor, dynamically adjust the pursuit target according to environmental information, avoid multiple pursuers concentrating on pursuing the same target, and reduce the overall efficiency. The pursuer uses the optimal path to approach the evader under the premise of ensuring safety, effectively avoids obstacles, avoids collision, and adapts to dynamic changes in the environment. The pursuer learns an optimized target selection strategy during training, which can adjust the priority according to distance, speed, teammate state and other information to ensure the optimal pursuit path.
[0075] The above only describes the preferred embodiments of the present application and does not limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning, characterized in that: Pursuit and escape methods include the following: Initialize modeling, model key variables and tasks, and initialize variables; Pursuit-escape modeling: All robots are set within a fixed perception radius and divided into pursuers and escapers. Through the self-attention mechanism of the Transformer architecture, the weights of key variables are dynamically adjusted to perform pursuit-escape modeling. Calculate the distance between the pursuer and the escapee and determine whether the pursuer has captured the escapee.
2. The multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning according to claim 1 is characterized in that: Key variables include robot size, obstacle locations, robot activity area and its boundaries; the task consists of a pursuer chasing a fugitive; Each robot is modeled as an individual of uniform size and subject to the same kinematic constraints. The obstacle positions are initialized to randomly distributed and are subject to dynamic disturbances. The boundaries of the robot's activity area are set to soft boundaries, and the activity space is set to a two-dimensional bounded space. The mission is completed when the pursuer chases the escapee to the set threshold.
3. The multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning according to claim 2 is characterized in that: The pursuer makes decisions through reinforcement learning strategies, and the evader uses a predefined artificial potential field (APF) strategy to escape. There is no communication between the pursuer and the evader, and neither party has information about the other's strategy.
4. The multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning according to claim 3 is characterized in that: Assume that there are N pursuers and M escapers. The judgment method of the pursuer chasing the escaper is: the pursuer approaches the escaper to the set threshold distance , while avoiding collisions with other robots or obstacles; When the Euclidean distance between a fugitive and at least one pursuer If the following conditions are met, the pursuer Successfully pursued the fugitive : ; in, For the A pursuer, For the A fugitive, For the pursuer location, For escapees location, The threshold distance for successful pursuit is set.
5. The multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning according to claim 4 is characterized in that: Since there is no communication between the pursuer and the evader, the pursuit-escape method is modeled as a partially observable Markov decision process, which is formally defined as follows: ; in, is the state space, which is used to describe the state information of the pursuer and the escaper; is the action space, which is used to define the pursuer's optional actions, including acceleration and angular velocity; is the state transition probability, which is used to describe the dynamic changes of the robot state over time; is a reward function used to motivate the robot to quickly complete the pursuit task and avoid collisions; The observation space is used to define the local environment information that each pursuer can observe; is the observation probability distribution, which describes the observation probability of the pursuer in different states; is a discount factor used to control the weight of future rewards.
6. The multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning according to claim 5 is characterized in that: At each time step , set each pursuer Accept a local observation , and select an action based on the current strategy , the environment is based on the state transition probability Update the status of all pursuers and escapers, and the pursuers receive reward signals To optimize the long-term cumulative return, the optimal strategy aims to maximize the discounted cumulative reward : ; in, is the mathematical expectation, which is used to calculate the expected value of the discounted cumulative reward. Strategy is the decision rule obtained by the pursuer based on reinforcement learning training, which determines the logic of its action selection under different local observations. is the state at time step t.
7. The multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning according to claim 6 is characterized in that: At each time step In each pursuer Local observation of It consists of the following information: ; in, The status information of the pursuer itself; The observation information of other pursuers within the sensing range, i.e. the pursuer's teammates; Observe information for escapees; Obstacle information; The action space of each pursuer is defined as: ; in, is the linear acceleration, used to control the speed change, is the angular velocity, which is used to control the direction change.
8. The multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning according to claim 7 is characterized in that: The pursuit-escape method adopts a training paradigm. During the training phase, the pursuer accesses global information, including the positions and velocities of all pursuers and escapers, to optimize the robot's behavioral strategy during training and learn more efficient collaborative pursuit methods. During the execution phase, global information is removed, and each pursuer makes decisions based only on its own local observations, ensuring that the pursuer can still complete efficient multi-target pursuit tasks without communication conditions.
9. The multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning according to claim 8, characterized in that: Reward Function Including safety reward function , target allocation reward function and the task completion reward function ; Safety Reward Function Specifically include: ; ; ; in, is the minimum Euclidean distance between the pursuer and all observed teammates, escapers, and obstacles; is the distance from the pursuer to the nearest escapee; is the threshold distance; It is a safety threshold used to reduce the risk of collision; Target allocation reward function is the global reward mechanism, for each escaper , calculate the number of pursuers around it , then at the relative distance The number of pursuers in is: ; If there is a fugitive being chased by multiple pursuers , and the number of pursuers of the other escapee If the number is less than 3, all pursuers will receive an additional global penalty of -10. If a pursuer crosses the line, they will receive an additional penalty of -5. Task completion reward function : ; When a fugitive When a target is successfully pursued by at least one pursuer, the mission completion reward is distributed to all pursuers who participated in the pursuit of the target; At the same time, in order to encourage faster completion of the pursuit task, a time penalty is imposed on all pursuers at each time step. : ; When the pursuer adopts a conservative strategy, a time penalty takes effect.
10. The multi-robot multi-target large-scale pursuit and escape method based on reinforcement learning according to claim 8, characterized in that: During the training phase, a curriculum learning method is used, with three learning processes set up to gradually increase the difficulty of the tasks. The learning process includes the following three steps: In the initial stage, a simple environment is set up, including a small number of obstacles and robots, and a small perception radius; In the mid-term stage, the number of robots will be gradually increased and complex environmental interference will be introduced; The advanced stage adopts a complete large-scale environment and introduces more intelligent pursuit and escape methods.