Multi-robot collaborative boarding method and system based on reinforcement learning
By introducing an adaptive grouping mechanism of reachable security area model and K-Medoids clustering, combined with the TrapNet decision network, the coordination problem in dynamic environment in multi-robot enclosure technology is solved, efficient and stable multi-objective capture is achieved, and the system's task execution capabilities and robustness are improved.
Patent Information
- Application Number
- CN202510975988.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-07-16
AI Technical Summary
The existing multi-robot enclosure technology has shortcomings in dealing with multi-objective coordination problems, obstacle utilization efficiency, strategy generalization ability, and dynamic grouping adjustment in dynamic environments, and it is difficult to meet the comprehensive requirements for efficiency, autonomy and stability in practical application scenarios.
Using a multi-robot collaborative encirclement method based on reinforcement learning, combined with adaptive grouping mechanism and policy network optimization design, the adaptive dynamic grouping and collaborative decision-making of robots in complex environments is realized by introducing accessible security area models, K-Medoids clustering and TrapNet decision-making networks.
It significantly improves the task execution capabilities of multi-robot systems in complex environments, improves the success rate of target capture, path length and time efficiency, and enhances the robustness and coordination efficiency of the system.
Smart Images

Figure CN120469431A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot collaborative control, and in particular to a multi-robot collaborative encirclement method and system based on reinforcement learning. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Multi-robot containment tasks, a key branch of multi-agent collaborative control, are widely used in high-risk scenarios such as disaster relief, military lockdowns, and security patrols. The goal is to effectively surround and restrict the movement of a target through the coordinated efforts of multiple robots. However, current mainstream containment methods can be broadly divided into traditional model-based control approaches and the recently emerging intelligent decision-making methods based on reinforcement learning. While the former possesses theoretical rigor, it suffers from significant shortcomings in adaptability to dynamic environments and efficient multi-agent coordination. While the latter demonstrates strong environmental generalization capabilities, it still faces numerous challenges in multi-target collaborative containment and structural design.
[0004] Traditional containment methods often rely on classical control theory, optimization algorithms, or heuristic rules, obtaining robot motion paths through the optimization of static policies or objective functions. While these methods can achieve relatively ideal containment results in structured, predictable environments, they often fail due to high model complexity and insufficient real-time responsiveness when faced with dynamic obstacles, nonlinear target escape behavior, and coordinated conflicts among multiple agents. For example, while methods based on Voronoi partitioning and Hungarian matching can achieve relatively good target allocation and path optimization in static scenarios, they suffer from significant drawbacks such as high computational overhead and rigid policies in real-world tasks with multiple targets, multiple obstacles, and high dynamics, making them difficult to meet real-time and scalability requirements. Furthermore, while bio-inspired algorithms possess certain adaptability and distributed execution characteristics, most remain at the heuristic level, lacking adjustable parameter tuning mechanisms, are prone to falling into local optima, lack global optimality guarantees, and perform poorly in complex, irregular environments.
[0005] With the development of deep reinforcement learning (RL), researchers have begun to explore its application to trapping tasks, leveraging its adaptive policy learning capabilities to enhance system intelligence. Reinforcement learning methods, which continuously adjust their strategies through interaction between the agent and the environment, possess strong self-optimization capabilities and show great potential in coping with dynamic target escapes and complex obstacle environments. However, existing RL trapping methods still face several prominent challenges: First, some studies focus on single-target trapping and fail to fully consider the task allocation and policy coordination issues in multi-target escape scenarios. This leads to uneven resource utilization and difficulty trapping some targets in complex tasks. Second, current methods generally neglect the strategic use of environmental obstacles, failing to design mechanisms that enable robots to proactively use obstacles to construct a containment zone, thus limiting trapping efficiency and policy flexibility. Third, there is a lack of effective coupling between individual strategies and collective coordination, making it difficult for reinforcement learning models to balance overall system collaboration and individual execution feasibility during training. In particular, the lack of dynamic grouping and task matching mechanisms in multi-target trapping tasks further exacerbates the instability of system coordination. Specifically, current reinforcement learning models often struggle to adjust grouping strategies in response to situations such as frequent target state changes and the intersecting escapes of multiple targets. This can cause some targets to remain unattended for extended periods, or for some robot resources to be concentrated in localized areas. This leads to fragmented system behavior and reduced collaborative efficiency, ultimately impacting the stability and sustainability of the containment effort. These issues are particularly pronounced in real-world multi-robot scenarios, such as security patrols, post-disaster search, and field hunting, where targets are often highly mobile and elusive, requiring real-time response to environmental changes and flexible strategy adjustments. Without an effective grouping mechanism and collaborative control, strategies trained using reinforcement learning struggle to achieve optimal solutions and may even be unable to effectively contain targets. Therefore, incorporating dynamic grouping strategies that are real-time, task-adaptive, and collaboratively consistent into multi-target containment becomes a key area for improving the practicality of reinforcement learning containment systems.
[0006] Therefore, the existing multi-robot encirclement technology still has many shortcomings in dealing with multi-target coordination problems in dynamic environments, obstacle utilization efficiency, strategy generalization capabilities, and dynamic grouping adjustment, and it is difficult to fully meet the comprehensive requirements of efficiency, autonomy, and stability in actual application scenarios. Summary of the Invention
[0007] In response to the shortcomings of the existing technology, the purpose of the present invention is to provide a multi-robot collaborative encirclement method and system based on reinforcement learning, which combines technical means such as adaptive grouping mechanism and strategy network optimization design to achieve efficient, stable and intelligent encirclement decision-making construction, and can significantly improve the task execution capability of the multi-robot system in complex environments.
[0008] In order to achieve the above object, the present invention is implemented through the following technical solutions: A first aspect of the present invention provides a multi-robot collaborative encirclement method based on reinforcement learning, comprising the following steps: A multi-robot trapping task model is established based on the target trapping task. The robots include a target robot and a pursuit robot. The pursuit robot traps the target robot and uses the reachable safe area to quantify the degree of restriction of the target robot. Based on the clustering algorithm, the pursuit robots in the siege task model are adaptively and dynamically grouped; Based on the grouping of the pursuit robots, a multi-robot encirclement decision network is used to mine environmental observation information, and the action decisions of the pursuit robots are optimized according to the environmental observation information.
[0009] Furthermore, each chasing robot acts as an intelligent agent, and the number of chasing robots is greater than the number of target robots.
[0010] Furthermore, the specific steps for establishing a multi-robot trapping task model based on the target trapping task are as follows: Determine the mission area based on the target encirclement mission; Set the motion space, observation space, and reward function for the pursuit robot and target robot; Determine the initial coordinated siege mechanism and the initial escape mechanism.
[0011] Furthermore, in the siege task model, the goal of the pursuit robot is to reduce the reachable safe area of all target robots, while the goal of the target robot is to expand the reachable safe area. When the area of the reachable safe area is smaller than the set threshold, it is considered that the corresponding target robot can no longer move, that is, the target robot has been captured by the pursuit robot.
[0012] Furthermore, based on the clustering algorithm, the specific steps for adaptively and dynamically grouping the pursuit robots in the siege task model are as follows: Using each target robot as the initial cluster center, each pursuit robot is assigned to the target cluster closest to it to complete the initial grouping. The total distance within the initial group is used as the optimization target, the pros and cons of the grouping are evaluated, and the grouping is adjusted according to the evaluation results to obtain the final grouping.
[0013] Furthermore, based on the grouping of the pursuit robots, the specific steps of using the multi-robot siege decision network to mine environmental observation information are as follows: Constructing a multi-robot trapping decision network and training the multi-robot trapping decision network, wherein the trapping decision network includes a policy network and a value network; The trained multi-robot enclosure decision network is used to mine environmental observation information, where the environmental observation information includes environmental observation information related to obstacles and environmental observation information related to the intelligent agent.
[0014] Furthermore, each hunting robot has its own independent policy network, which is used to make action decisions based only on its own observable information.
[0015] A second aspect of the present invention provides a multi-robot collaborative encirclement system based on reinforcement learning, comprising: a data initialization module configured to establish a multi-robot trapping task model based on the target trapping task, wherein the robots include a target robot and a pursuit robot, the pursuit robot traps the target robot, and uses a reachable safe area to quantify the degree of restriction of the target robot; The dynamic grouping module is configured to adaptively and dynamically group the pursuit robots in the siege task model based on a clustering algorithm; The action decision module is configured to mine environmental observation information based on the grouping of the pursuit robots using a multi-robot encirclement decision network, and optimize the action decision of the pursuit robots according to the environmental observation information.
[0016] The third aspect of the present invention provides a computer-readable storage medium storing a computer program, which is suitable for being loaded by a processor and executing the steps in the multi-robot collaborative encirclement method based on reinforcement learning as described in the first aspect of the present invention.
[0017] A fourth aspect of the present invention provides a computer device, comprising: a processor adapted to execute a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the multi-robot collaborative encirclement method based on reinforcement learning as described in the first aspect of the present invention is implemented.
[0018] One or more of the above technical solutions have the following beneficial effects: This paper discloses a multi-robot collaborative trapping method and system based on reinforcement learning. To address the problems of uneven target distribution and poor strategy generalization in multi-robot multi-target trapping tasks, a multi-robot collaborative trapping method is proposed that integrates adaptive task grouping with a centralized training distributed execution (CTDE) mechanism. This method introduces a reachable safe zone (ASZ) during the modeling phase to rationally constrain the spatial generation of trapping strategies and construct a trapping task model. By incorporating a dynamic grouping mechanism implemented through K-Medoids clustering, the robots can adaptively adjust based on the current environmental state and target distribution, addressing the low capture efficiency caused by static robot grouping.
[0019] This paper combines the CTDE multi-agent reinforcement learning framework in the strategy training process and introduces networks such as MLP and CNN to model the relationship between multi-agents and environmental observations, significantly improving the utilization of local information and the coordination of global decision-making. Experimental results show that the proposed method significantly improves the target capture success rate in multi-target scenarios compared to traditional strategies (such as artificial potential field methods), achieves shorter path lengths and times, and exhibits higher robustness in more complex environmental scenarios.
[0020] By coupling a multi-agent adaptive grouping mechanism with a reinforcement learning strategy, this invention enables robots to dynamically adjust their capture strategies and execution paths when faced with varying numbers and states of targets, achieving optimal resource allocation and behavioral coordination. This mechanism reduces the resource waste and target escape risk associated with traditional methods, and offers excellent scalability, making it suitable for large-scale robotic systems for tasks such as target capture in complex environments.
[0021] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0023] Figure 1 This is a flow chart of a multi-robot collaborative trapping method based on reinforcement learning in Example 1 of the present invention; Figure 2 This is a schematic diagram of a reachable safe area in the first embodiment of the present invention; Figure 3Schematic diagram of the local mask matrix of the robot in Example 1 of the present invention; Figure 4 Schematic diagram of adaptive dynamic grouping modeling based on K-Medoids clustering in Example 1 of the present invention; Figure 5 This is a diagram of the TrapNet strategy network structure in Example 1 of the present invention. DETAILED DESCRIPTION
[0024] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used in this embodiment have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0025] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations; The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0026] Example 1: The first embodiment of the present invention provides a multi-robot collaborative trapping method based on reinforcement learning. This method addresses the problems of imperfect modeling and insufficient collaborative strategies in existing methods for multi-target trapping tasks. The following three key technical issues are proposed and systematically addressed: First, existing research fails to fully consider the issues of task allocation and strategy coordination under multi-target escape behaviors, lacking an effective modeling mechanism for the trapping problem. When faced with multiple dynamic escape targets, it is difficult to rationally define the task boundary between the robot's motion space and the trapping targets, resulting in a lack of systematic and targeted trapping behavior. To address this issue, this paper introduces the reachable safe zone (ASZ) model to establish a unified modeling framework for multi-target trapping tasks. This systematically defines the robot's motion space, observation space, and reward function, enabling target-state-driven task modeling and trapping mechanism design, significantly improving the modeling accuracy and strategy implementability of the trapping task.
[0027] Secondly, existing methods lack flexible and efficient resource allocation and coordination mechanisms when faced with changes in target state or interference from obstacles, resulting in low target tracking efficiency and difficulty ensuring safety. This paper proposes an adaptive grouping strategy based on K-Medoids clustering. This strategy can dynamically adjust the grouping strategy based on the current target position and robot state, allowing each robot group to adaptively balance its forces around the target.
[0028] Finally, existing multi-robot systems generally suffer from insufficient perception capabilities in dynamic and complex environments. They find it difficult to fully perceive the environmental structure and target distribution when multi-source observation information is limited, thus affecting the effectiveness and stability of the entrapment strategy. To address this issue, the present invention designs a multi-robot entrapment decision network framework, TrapNet, that integrates the CTDE mechanism. By constructing a hierarchical strategy network structure consisting of an encoding module and a decision module, it effectively integrates the robot's own state information, the states of other intelligent agents, and the spatial distribution information of obstacles in the environment, thereby improving the ability to extract multi-source information and the accuracy of environmental perception. A centralized critic is introduced during network training to enhance global collaboration capabilities. During the execution phase, each robot performs action reasoning based on its own observations, achieving distributed collaborative entrapment that is stable in dynamic environments. This is particularly true in scenarios with dense obstacles or complex structures, where obstacles can be used to form natural entrapment boundaries, thereby improving the overall entrapment success rate.
[0029] like Figure 1 As shown, the specific steps include: Step 1: Establish a multi-robot trapping task model based on the target trapping task.
[0030] Among them, the robots include target robots and pursuit robots. The pursuit robots surround the target robots and use the reachable safe area to quantify the degree of restriction of the target robots.
[0031] Step 1.1: Determine the mission area based on the target encirclement mission.
[0032] In this embodiment, the task area of the target encirclement task is set as a two-dimensional free space S containing obstacles.
[0033] Step 1.2: Set the motion space, observation space, and reward function for the pursuit robot and target robot.
[0034] First, a multi-robot target encirclement task is modeled. The multi-robot encirclement task studied in this embodiment includes two types of robots that compete with each other, namely, the pursuit robot (number ) and target robots (number of ),and Each pursuing robot or target robot acts as an agent, and the number of pursuing robots is greater than the number of target robots. A unified second-order dynamics model is established for both the pursuing and target robots, with corresponding action and observation spaces constructed. Multi-agent target trapping rewards are set to incentivize the completion of the task.
[0035] Step 1.3: Determine the initial coordinated entrapment mechanism and the initial escape mechanism.
[0036] In the encirclement task of this embodiment, the initial collaborative encirclement mechanism is that the pursuit robot encircles the target robot that is close to it based on the initial positions of all robots. The initial escape mechanism is obtained through a heuristic method, and the target robot performs the escape task according to the obtained escape plan. Based on the initial collaborative encirclement mechanism and the initial escape mechanism, the pursuit robot should use obstacles in the environment to restrict the movement of the target robot until the target robot can no longer move, and the target robot needs to find a suitable path to break through the encirclement of the pursuit robot to escape. On this basis, this embodiment defines a concept called Accessible Safety Zone (ASZ), which represents the area that the target robot can safely reach, and is used to quantify the degree of restriction of the target robot in order to establish a mathematical expression for the encirclement process. As Figure 2 As shown in the figure, in the siege task model, the goal of the pursuit robot is to reduce the reachable safe area of all target robots, while the goal of the target robot is to expand the reachable safe area as much as possible. When the area of the reachable safe area is smaller than the set threshold, it is considered that the corresponding target robot can no longer move, that is, the target robot has been captured by the pursuit robot.
[0037] Specifically, assuming that in a two-dimensional free space (including fixed obstacles), there is A hunting robot and target robots. Moment, hunt robots The location is , target robot The location is The arrival functions of the pursuit robot and the target robot are expressed as and . Then the target robot exist The ASZ of a moment is defined as follows: .
[0038] in, Indicates the target robot At the moment The smaller the reachable safe area, the more dangerous the corresponding target robot is. represents a point in the motion space. In a multi-robot target capture task, the pursuit robot's goal is to reduce the ASZ of all target robots, while the target robot's goal is to maximize its ASZ. If the ASZ area is smaller than a certain threshold, rendering the corresponding target robot unable to move, the pursuit robot is considered to have captured the target robot.
[0039] Based on the above concepts, a complete mathematical definition of the multi-robot target trapping task is given. Assume A hunting robot and The target robot has a motion space of Initial moment When hunting robots and target robot The locations are and , corresponding to the target robot The initial ASZ is recorded as .exist time, and Hunting Robot and target robot The action control signal, is the action space. Let the time budget of the task be , the multi-robot target encirclement task can be described by the dynamic change of the ASZ area, which can be expressed as: .
[0040] in and are the motion equations of the pursuit robot and the target robot respectively. express Time target robot The optimization goal is to chase the robot from the perspective of ASZ. The purpose is to The size of the target robot's ASZ is minimized as much as possible to complete the encirclement mission more effectively.
[0041] It should be noted that this embodiment uses the second-order dynamics model and artificial potential field method in the encirclement action planning to achieve obstacle avoidance. Specifically, when setting the action space, the robot's action force and the virtual repulsion of the obstacle are set, and the robot moves through the combined force of the action force and the virtual repulsion of the obstacle. Among them, the action force is obtained by converting the discrete action value using the second-order dynamics model. Discrete action values include: forward, left turn, right turn, stop, and backward. The virtual repulsion of the obstacle is obtained by calculating the virtual repulsion of the surrounding obstacles using the artificial potential field method, which can improve the obstacle avoidance ability.
[0042] In some other implementations, a dynamic window method, nonlinear model predictive control, an obstacle avoidance strategy based on reinforcement learning, or a composite obstacle avoidance mechanism combining geometric methods and graph search may also be used to improve the ability to adapt to complex obstacles, dynamic targets, or multi-layer environments.
[0043] Step 2: Based on the clustering algorithm, the pursuit robots in the siege task model are adaptively and dynamically grouped, such as Figure 4 shown.
[0044] Step 2.1: Using each target robot as the initial cluster center, assign each pursuit robot to the target cluster with the closest distance to complete the initial grouping.
[0045] This embodiment, based on the core idea of the K-Medoids clustering algorithm, designs an adaptive dynamic grouping modeling method for multi-robot encirclement tasks. This method aims to solve the problem of reasonable grouping between multiple pursuit robots and multiple target robots. The core idea is to divide the pursuit robots into several groups and optimize the collaborative efficiency of each group when encircling the corresponding target. Considering that in the encirclement task, the spatial relationship between the pursuit robots and the target directly affects the encirclement effect, this embodiment will The value is set directly to the target quantity , the target is the target robot, to ensure that each target is assigned a group of pursuit robots to besiege. In the modeling process, calculate each pursuit robot With each target robot The Euclidean distance between them is used to construct a distance matrix , where a single distance is defined as ,in and Respectively represent the pursuit robot and target robot This embodiment uses each target robot as the initial cluster center and assigns each pursuit robot to the target cluster to which it is closest. This grouping is achieved by minimizing the sum of distances within all target clusters. This is recalculated every set round to achieve dynamic grouping.
[0046] Step 2.2: Take the total distance within the initial group as the optimization target, evaluate the quality of the group, adjust the grouping according to the evaluation results, and obtain the final grouping.
[0047] In order to quantify the quality of the current grouping, this embodiment introduces the total distance within the group as the optimization target and defines the The distance cost of a cluster of target robots is The overall optimization goal is to minimize the total distance sum within all target clusters, that is: .
[0048] in, is the overall optimization goal, Indicates that it is assigned to the target robot A collection of hunting robots.
[0049] It is important to note that in addition to using K-Medoids clustering to implement dynamic grouping strategies for multiple robots targeting targets, some other implementations can also substitute an intelligent adaptive grouping mechanism based on a neural network. This mechanism can integrate multi-dimensional information such as robot motion state and environmental characteristics, and utilize end-to-end deep reinforcement learning or attention mechanisms to dynamically optimize the grouping strategy to meet the needs of efficient collaboration in large-scale systems or highly dynamic environments. In addition, traditional methods such as DBSCAN and hierarchical clustering can also be considered to adjust groupings for different task scales and objectives, thereby ensuring system flexibility and robustness.
[0050] Step 3: Based on the grouping of the pursuit robots, the multi-robot siege decision network is used to mine the environmental observation information and optimize the action decision of the pursuit robots according to the environmental observation information.
[0051] Step 3.1: Construct a multi-robot encirclement decision network and train the multi-robot encirclement decision network.
[0052] This embodiment combines the theory of multi-agent reinforcement learning and the idea of centralized training and distributed execution (CTDE) under the reinforcement learning framework to design a multi-robot trapping decision network named TrapNet. Figure 5 As shown in the figure, this network is designed to fully exploit surrounding environmental observations and enhance temporal feature modeling capabilities, thereby improving the autonomous collaborative capture effectiveness of robot swarms in complex environments. The capture decision network comprises a policy network and a value network. The construction of the policy network (actor) and the value network (critic) is crucial in the overall framework design.
[0053] To more effectively process multi-source observation data and generate appropriate action decisions, the TrapNet policy network employs a two-level architecture: an encoding module and a decision module. The encoding module extracts and fuses features from all agent observations and the surrounding environment, while the decision module generates the final action output based on these extracted and fused features, ensuring the robot makes the optimal decision based on the environment.
[0054] Specifically, the encoding module is composed of a multi-layer perceptron (MLP) and a convolutional neural network (CNN), which respectively perform feature processing on different types of input information. Figure 5 As shown in FIG, in this embodiment, the environmental observation information includes environmental observation information related to obstacles and environmental observation information related to the agent, which is directly obtained from the robot's observation space. Specifically, each robot's environmental observation information related to obstacles can be represented as a local mask matrix centered on itself and rotating in real time with the robot's orientation, where each binary element in the matrix indicates whether there is an obstacle at its corresponding position, as shown in FIG. Figure 3 As shown. Through this local map, the robot can obtain the distribution of surrounding obstacles in real time, providing support for path planning and obstacle avoidance decisions. The environmental observation information related to the intelligent agent includes its own internal state, the position, speed and orientation of other robots, and other information. Combined with the design of observation information, MLP is mainly used to process observation information related to the intelligent agent, and can effectively extract features at the individual level. The CNN part is specifically used to process environmental observation information related to obstacles, using the powerful spatial feature extraction capability of the convolutional structure to capture the local spatial relationship of obstacles and the environmental distribution characteristics. After the two parts of feature extraction are completed, the output results are spliced and fused at the feature layer to provide a complete and rich environmental representation for the subsequent decision-making module.
[0055] Based on this, the decision module further receives the fused feature vectors and generates the robot's action decision. This module is composed of a combination of an MLP and a recurrent neural network (RNN). The MLP is used to perform nonlinear transformations and further refine the fused features, while the introduction of the RNN enables the network to process time series information. By modeling historical observation sequences, the RNN can capture the dynamic characteristics of the environmental state over time, significantly enhancing the model's temporal reasoning capabilities in continuous decision-making scenarios. Ultimately, the decision module outputs the specific actions that each agent should take at the current moment, effectively implementing the encirclement strategy.
[0056] To further enhance multi-agent collaboration, the CTDE framework (Centralized Training and Distributed Execution) was introduced during the training of TrapNet's multi-robot trapping decision network. This framework utilizes a centralized value assessment mechanism during training, where all agents share a centralized critic network. This critic network receives global observations and joint actions, and evaluates the value of the current policy, providing accurate global feedback for policy optimization.
[0057] Distributed execution means each hunting robot has its own independent policy network, which makes action decisions based solely on its own observable information. This design ensures the model's fully distributed nature and strong environmental adaptability during execution.
[0058] After each pursuit robot makes its own action decision, the Critic network can obtain global observation information and joint actions to evaluate the value of the current strategy and then make adjustments based on the evaluation.
[0059] Step 3.2: Use the trained multi-robot enclosure decision network to mine environmental observation information.
[0060] It should be noted that this embodiment currently uses the centralized training distributed execution (CTDE) architecture to implement high-level role selection. In some other implementations, it can also be replaced by a distributed reinforcement learning framework (such as VDN, QMIX, MADDPG) or a task scheduling method that integrates game theory, graph search and other mechanisms.
[0061] In addition, while this embodiment primarily uses networks such as MLP and CNN for multi-robot information fusion, other implementations could also substitute various information modeling approaches, such as attention mechanisms, autoencoder structures, or local collaborative aggregation algorithms. In scenarios where communication bandwidth is limited or data is incomplete, communication mechanisms such as compression coding and round-robin sharing could be designed to improve communication efficiency and collaborative accuracy.
[0062] This embodiment can be further extended to irregular, dynamic, or high-dimensional scenarios, such as multi-story buildings, complex mixed indoor and outdoor environments, and dynamic obstacle interference fields. To adapt to the interference and uncertainty in more realistic scenarios, three-dimensional map modeling, spatiotemporal fusion perception mechanisms, or simulation of sensor errors in real environments can be introduced to improve the robustness of the algorithm and its engineering applicability.
[0063] In order to cope with the computing resource pressure caused by the expansion of the number of robots and the complexity of tasks, real-time guarantee can be achieved during system deployment through model compression, knowledge distillation, lightweight networks, etc., to promote the engineering application and rapid deployment of the method of the present invention in actual robot systems.
[0064] Example 2: A second embodiment of the present invention provides a multi-robot collaborative encirclement system based on reinforcement learning, comprising: a data initialization module configured to establish a multi-robot trapping task model based on the target trapping task, wherein the robots include a target robot and a pursuit robot, the pursuit robot traps the target robot, and uses a reachable safe area to quantify the degree of restriction of the target robot; The dynamic grouping module is configured to adaptively and dynamically group the pursuit robots in the siege task model based on a clustering algorithm; The action decision module is configured to mine environmental observation information based on the grouping of the pursuit robots using a multi-robot encirclement decision network, and optimize the action decision of the pursuit robots according to the environmental observation information.
[0065] Example 3: Embodiment 3 of the present invention provides a computer-readable storage medium storing a computer program, which is suitable for being loaded by a processor and executing the steps of the multi-robot collaborative encirclement method based on reinforcement learning as described in embodiment 1 of the present invention.
[0066] Example 4: A fourth embodiment of the present invention provides a computer device, comprising: a processor adapted to execute a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the steps of the multi-robot collaborative encirclement method based on reinforcement learning as described in the first embodiment of the present invention are implemented.
[0067] The steps involved in the above embodiments 2, 3 and 4 correspond to those in the method embodiment 1. For the specific implementation methods, please refer to the relevant description part of the embodiment 1.
[0068] Those skilled in the art will appreciate that the units and algorithmic steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data processing device such as a server or data center that integrates one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)). The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any technical object of a person skilled in the art that can be easily conceived of within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A multi-robot collaborative encirclement method based on reinforcement learning, characterized in that: The following steps are involved: A multi-robot trapping task model is established based on the target trapping task. The robots include a target robot and a pursuit robot. The pursuit robot traps the target robot and uses the reachable safe area to quantify the degree of restriction of the target robot. Based on the clustering algorithm, the pursuit robots in the siege task model are adaptively and dynamically grouped; Based on the grouping of the pursuit robots, a multi-robot encirclement decision network is used to mine environmental observation information, and the action decisions of the pursuit robots are optimized according to the environmental observation information.
2. The multi-robot collaborative trapping method based on reinforcement learning according to claim 1, characterized in that: Each hunting robot acts as an intelligent agent, and the number of hunting robots is greater than the number of target robots.
3. The multi-robot collaborative trapping method based on reinforcement learning according to claim 1, characterized in that: The specific steps for establishing a multi-robot trapping task model based on the target trapping task are as follows: Determine the mission area based on the target encirclement mission; Set the motion space, observation space, and reward function for the pursuit robot and target robot; Determine the initial coordinated siege mechanism and the initial escape mechanism.
4. The multi-robot collaborative trapping method based on reinforcement learning according to claim 1, characterized in that: In the siege task model, the goal of the pursuit robot is to reduce the reachable safe area of all target robots, while the goal of the target robot is to expand the reachable safe area. When the area of the reachable safe area is smaller than the set threshold, it is considered that the corresponding target robot can no longer move, that is, the target robot has been captured by the pursuit robot.
5. The multi-robot collaborative trapping method based on reinforcement learning according to claim 1, characterized in that: Based on the clustering algorithm, the specific steps for adaptive dynamic grouping of the pursuit robots in the siege task model are as follows: Using each target robot as the initial cluster center, each pursuit robot is assigned to the target cluster closest to it to complete the initial grouping. The total distance within the initial group is used as the optimization target, the pros and cons of the grouping are evaluated, and the grouping is adjusted according to the evaluation results to obtain the final grouping.
6. The multi-robot collaborative trapping method based on reinforcement learning according to claim 1, characterized in that: Based on the grouping of the pursuit robots, the specific steps of mining environmental observation information using the multi-robot trapping decision network are as follows: Constructing a multi-robot trapping decision network and training the multi-robot trapping decision network, wherein the trapping decision network includes a policy network and a value network; The trained multi-robot enclosure decision network is used to mine environmental observation information, where the environmental observation information includes environmental observation information related to obstacles and environmental observation information related to the intelligent agent.
7. The multi-robot collaborative trapping method based on reinforcement learning according to claim 6, characterized in that: Each hunting robot has its own independent policy network, which is used to make action decisions based only on its own observable information.
8. A multi-robot collaborative siege system based on reinforcement learning, characterized in that: include: a data initialization module configured to establish a multi-robot trapping task model based on the target trapping task, wherein the robots include a target robot and a pursuit robot, the pursuit robot traps the target robot, and uses a reachable safe area to quantify the degree of restriction of the target robot; The dynamic grouping module is configured to adaptively and dynamically group the pursuit robots in the siege task model based on a clustering algorithm; The action decision module is configured to mine environmental observation information based on the grouping of the pursuit robots using a multi-robot encirclement decision network, and optimize the action decision of the pursuit robots according to the environmental observation information.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is suitable for being loaded by a processor and executing the multi-robot collaborative encirclement method based on reinforcement learning according to any one of claims 1 to 7.
10. A computer device, characterized in that: include: a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein the computer program, when executed by the processor, implements the multi-robot collaborative encirclement method based on reinforcement learning as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-target capturing method for cooperative operation of swarm robots in complex non-convex environment
CN111240333A
Unmanned aerial vehicle cooperative pursuit method of multi-degree-of-freedom model based on multi-agent reinforcement learning
CN116225065A
Multi-agent self-organizing synergistic hunting method in non-convex environment
CN117574950A
Water surface target collaborative hunting method based on multi-agent reinforcement learning
CN117806318A
Unmanned aerial vehicle cluster dynamic task re-planning framework for real-time decision
CN120066067A
Cited By
Unmanned cluster multi-target encircle resource perception type task allocation method, device and equipment and medium
CN121979618A
Unmanned cluster multi-target encirclement and capture resource perception type task allocation method, device, equipment and medium
CN121979618B