A deep reinforcement learning system using asymmetric self-play for robust multi-robot flocking
The deep reinforcement learning system with an asymmetric self-game framework and auxiliary training module enhances multi-robot swarm control in dynamic environments, addressing the limitations of existing methods by improving adaptability and robustness through two-stage training and attention mechanisms.
Patent Information
- Authority / Receiving Office
- HK · HK
- Patent Type
- Applications
- Current Assignee / Owner
- LINGNAN UNIVERSITY
- Filing Date
- 2026-05-14
- Publication Date
- 2026-07-17
AI Technical Summary
Existing multi-robot swarm control methods struggle to navigate safely and efficiently in complex, dynamic environments with adversarial obstacles, often relying on static obstacle assumptions that limit their generalization and robustness.
A deep reinforcement learning system using an asymmetric self-game framework with a two-stage training paradigm and an auxiliary training module enhances environmental perception and adaptability, incorporating feature-level and agent-level attention mechanisms to handle dynamic obstacles and adversarial interference.
The proposed system significantly improves policy robustness and generalization ability, enabling effective navigation and collision avoidance in complex scenarios, as demonstrated by extensive experiments and practical deployments.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202511476089.3 (22) Application Date 2025.10.16 (30) Priority Data 18 / 919,653 2024.10.18 US (71) Applicant Lingnan University Address Tuen Mun, New Territories, Hong Kong, China (72) Inventors Kwong Tak-hu, Jia Yun-jie (74) Patent Agency Shenzhen Yibao Intellectual Property Agency (General Partnership) 44588 Patent Attorney Wang Qin, Cao Yu-cun (51) Int.Cl. B25J 11 / 00 (2006.01) G06N 3 / 045 (2023.01) G06N 3 / 092 (2023.01) G06N 5 / 04 (2023.01) (54) Invention Title: Robust Multi-Robot Swarm Based on Asymmetric Self-Game Theory (57) Abstract: This paper presents a task robot for a deep reinforcement learning system that can achieve robust multi-robot swarming using an asymmetric self-game algorithm. The task robot includes a reinforcement learning control module and an auxiliary training module. The control module can dynamically adjust the robot's behavior to optimize decision-making, and its features include a target navigation model, a swarm behavior maintenance model, and a collision avoidance model. The auxiliary training module enhances environmental perception capabilities, enabling the robot to predict dynamic changes. The auxiliary training module includes a local environment grid estimation model for generating small-scale maps and a motion prediction model for predicting the trajectories of the robot and obstacles. Claims 3 pages, Description 14 pages, Drawings 10 pages, CN 121893303 A 2026.04.21 CN 1 21 89 33 03 A 1. A task robot for a deep reinforcement learning system, the system utilizing asymmetric self-game to achieve robust multi-robot swarming, characterized in that the task robot comprises: a reinforcement learning control module, used to dynamically adjust the robot's behavior, help it make optimal decisions, and continuously learn and improve through a real-time feedback mechanism, wherein the reinforcement learning control module comprises: a target navigation model, used to guide the task robot toward a predetermined target area, continuously calculate the distance between the task robot and the target area, and generate action commands based on the calculated information to guide the task robot to shorten the distance to the target area; a swarm behavior maintenance model, used to maintain swarm behavior and monitor the relative positions and distances of other robots relative to the task robot, such that the swarm behavior maintenance model adjusts the movement of the task robot to maintain coordination within the robot group; and a collision avoidance model, used to avoid collisions, and the collision avoidance model adjusts the trajectory of the task robot according to the relative positions of obstacles and interfering robots; andAn auxiliary training module is used to enhance the environmental perception capability of the task robot, enabling the task robot to predict dynamic changes around it. The auxiliary training module includes: a local environment mesh estimation model for generating a small-scale environment mesh map at each time step to display the states of obstacles, interfering objects, and other robots around the task robot, providing contextual information for the task robot's subsequent decisions; and a motion prediction model for analyzing and predicting the trajectories of other robots and the attributes of obstacles and interfering objects based on the generated environment mesh map. 2. According to claim 1, the reinforcement learning control module is further used to optimize action learning and value learning in its neural network by using a proximal policy optimization (PPO) algorithm combined with a convolutional neural network (CNN) and a multilayer perceptron (MLP) architecture. 3. The task robot according to claim 2, wherein the framework adopted by the reinforcement learning control module can use a feature-level attention mechanism in the action learning and an agent-level attention mechanism in the value learning, wherein the action learning allows the assigned robot to make independent decisions through repeated training in a distributed execution mode, while the value learning is carried out in a centralized training process, wherein the assigned robot shares environmental data with other robots in the centralized training process, thereby realizing collective information sharing. 4. The task robot according to claim 1 further includes a reward function module for rewarding the task robot for performing beneficial actions during training, wherein the reward function module includes: a reward model for providing feedback to the task robot in the form of target approach reward, formation maintenance reward, and obstacle avoidance reward when the task robot approaches the target area; wherein, when the task robot approaches the target area, the reward function module provides a positive reward based on the reward model, thereby encouraging the robot to continue moving forward; wherein, when the task robot maintains an appropriate position in the robot swarm, the task robot receives a reward from the reward model, thereby encouraging cooperation; wherein, when the task robot successfully avoids collisions with the interfering robot and the static obstacle, the task robot receives a reward to promote safe navigation. 5. The task robot according to claim 1, further comprising: a two-stage self-adversarial training module, used to enhance the adaptability of the task robot in complex environments through adversarial training, and the two-stage self-adversarial training module consists of a first stage and a second stage; wherein, in the first stage, the task robot and the interference robot are trained simultaneously through the two-stage self-adversarial training module and interact through the two-stage self-adversarial training module, wherein the interference robot...The network parameters are periodically stored in the model pool, enabling the task robot to be trained for different levels of adversarial intelligence. In the second stage for policy testing and optimization, the two-stage self-adversarial training module introduces complex adversarial models to test and optimize the task robot's policies, and the interference models are sampled from the model pool in different combinations. 6. A deep reinforcement learning system based on asymmetric self-game theory for achieving robust multi-robot swarming, characterized in that the deep reinforcement learning system comprises: multiple interfering robots; multiple static obstacles; a starting region and a target region; and multiple task robots, activated to move from the starting region to the target region and bypass the interfering robots and the static obstacles; wherein each task robot is equipped with a physical communication device to achieve real-time data transmission and information sharing, enabling coordination and cooperation with other task robots; wherein each task robot comprises: a reinforcement learning control module for dynamically adjusting the robot's behavior to help it make optimal decisions and continuously learn and improve through a real-time feedback mechanism, wherein the reinforcement learning control module comprises: a target navigation model for guiding the task robot toward a predetermined target region, continuously calculating the distance between the task robot and the target region, and generating action commands based on the calculated information to guide the task robot to shorten the distance to the target region; and a swarm behavior maintenance model for maintaining swarm behavior and monitoring the relative positions and distances of other robots relative to the task robot, such that the swarm behavior maintenance model adjusts the movement of the task robot. To maintain coordination within the robot group; and a collision avoidance model for avoiding collisions, wherein the collision avoidance model adjusts the trajectory of the task robot based on the relative positions of the static obstacle and the interfering robot; and an auxiliary training module for enhancing the environmental perception capability of the task robot, enabling the task robot to predict dynamic changes around it, wherein the auxiliary training module includes: a local environment mesh estimation model for generating a small-scale environment mesh map at each time step to display the state of the interfering robot, the static obstacle, and other robots around the task robot, providing contextual information for the task robot's subsequent decisions; and a motion prediction model for analyzing and predicting the trajectories of other robots and the attributes of the static obstacle and the interfering robot based on the generated environment mesh map. 7. The deep reinforcement learning system according to claim 6, wherein the reinforcement learning control module is further configured to use a proximal policy optimization (PPO) algorithm, combined with a convolutional neural network (CNN) and a multilayer perceptron (MLP) architecture, to optimize action learning and value learning in its neural network.8. The deep reinforcement learning system according to claim 7, wherein the framework adopted by the reinforcement learning control module can use a feature-level attention mechanism in the action learning and an agent-level attention mechanism in the value learning, wherein the action learning allows the assigned robot to make independent decisions through repeated training in a distributed execution mode, while the value learning is carried out in a centralized training process, wherein all the task robots share environmental data and provide collective information sharing. 9. The deep reinforcement learning system according to claim 6, wherein each task robot further includes a reward function module for rewarding the task robot for performing beneficial actions during training, and includes: a reward model for providing feedback to the task robot in the form of target approach reward, formation maintenance reward, and obstacle avoidance reward when the task robot approaches the target area; wherein, when the task robot approaches the target area, the reward function module provides a positive reward based on the reward model, thereby encouraging the robot to continue moving forward; wherein, when the task robot maintains an appropriate position in the swarm of task robots, the task robot receives a reward from the reward model to encourage cooperation; wherein, when the task robot successfully avoids collisions with the interfering robot and the static obstacle, the task robot receives a reward to facilitate safe navigation. 10. The deep reinforcement learning system of claim 6, wherein each task robot further comprises a two-stage self-adversarial training module for enhancing the adaptability of the task robot in complex environments through adversarial training, and the two-stage self-adversarial training module consists of a first stage and a second stage, wherein in the first stage, the task robot and the interfering robot are trained simultaneously through the two-stage self-adversarial training module and interact through the two-stage self-adversarial training module, wherein the network parameters of the interfering robot are periodically stored in a model pool, enabling the task robot to train for different levels of adversarial intelligence; wherein in the second stage for policy testing and optimization, the two-stage self-adversarial training module introduces complex adversarial models to test and optimize the policy of the task robot, and the interfering models are sampled from the model pool in different combinations. 11. The deep reinforcement learning system of claim 6, wherein each interfering robot is equipped with an interfering reward model, which provides a mechanism to encourage the interfering robot to approach the task robot, and when any of the interfering robots successfully collides with one of the task robots, the interfering robot receives a reward from the interfering reward model. 12. The deep reinforcement learning system of claim 6, wherein each task robot and the interfering robot...The robot is used to execute steering commands that include linear velocity and angular velocity. Each task robot has four linear velocity options in its motion space: 0 m / s, 0.15 m / s, 0.3 m / s, and 0.5 m / s, and nine angular velocity options: -2 rad / s, -1.2 rad / s, -0.8 rad / s, -0.3 rad / s, 0 rad / s, 0.3 rad / s, 0.8 rad / s, 1.2 rad / s, and 2 rad / s. For each interfering robot, its angular velocity options are the same as those of the task robot, while its linear velocity options are 0 m / s, 0.15 m / s, 0.3 m / s, and 0.38 m / s. 13. The deep reinforcement learning system according to claim 6, wherein the deep reinforcement learning system serves as a training platform, and through two-stage swarm training, the robot is trained to navigate from the starting area to the target area while maintaining the swarm movement pattern, wherein the scene size of the first stage is 15m × 15m, including five swarms of the task robots, three of the interference robots, and two of the static obstacles, wherein the scene size of the second stage is 25m × 25m, including eight swarms of the task robots, six of the interference robots, and five of the static obstacles. Claims 3 / 3 Page 4 CN 121893303 A Deep Reinforcement Learning System for Robust Multi-Robot Swarms Based on Asymmetric Self-Game Theory Technical Field
[0001] This invention relates to the control technology of robust multi-robot swarm methods; and specifically to a deep reinforcement learning system based on asymmetric self-game, used to realize robust multi-robot swarm control. Background Art
[0002] Swarm control is a key technology for realizing safe and sustainable navigation of robot groups, enabling robots to approach each other and reach the target area without collision. However, most existing swarming solutions are based on the assumption that obstacles behave statically or constantly, which limits their ability to cope with malicious physical attacks. Therefore, there is a need to develop swarming strategies that can cope with dynamic disturbances in adversarial environments, thereby initiating safe cooperative navigation. A large body of literature aims to achieve swarming by following three principles: (1) collision avoidance, (2) velocity matching, and (3) swarm centering.
[0003] In some traditional approaches, methods such as those based on artificial potential fields (APF) and model predictive control (MPC) have been proposed to address these principles. However, while adopting manually designed rules can be intuitive, they often fall into local minima traps and their performance degrades significantly when the scene changes. Furthermore, these methods typically assume full access to environmental information and involve cumbersome procedures.The specific scenario design limits the feasibility and robustness in complex dynamic environments.
[0004] In recent years, deep reinforcement learning (DRL) has shown excellent performance in many multi-robot system tasks, and it can provide an alternative solution for swarm control in complex scenarios. Some DRL-based frameworks have shown the ability to avoid collisions in static obstacle environments. In addition, some methods have been proposed to deal with dynamic obstacles. However, in these works, the behavior of obstacles will always remain static, which also limits their impact on policy performance. Therefore, the swarm model trained under such conditions may overfit the behavior of fixed obstacles, thereby reducing its generalization ability to unknown scenarios.
[0005] Therefore, it is currently necessary to develop a robust swarm control method that can effectively handle real and dynamic scenarios to improve the performance and reliability of multi-robot systems operating in complex environments. Summary of the Invention
[0006] As mentioned above, swarm control is a key method to ensure the continuous navigation of multi-robot systems and is widely used in logistics, service delivery and search and rescue. However, the real world environment is often complex, dynamic and even hostile, posing a significant risk to the safety of swarm robots.
[0007] To address this challenge, an Asymmetric Self-play-empowered Flocking Control (ASFC) framework based on deep reinforcement learning is proposed. In this invention, the swarming robot can be trained together with a behavior-learnable adversarial disruptor to promote the development of more intelligent swarming strategies. A two-stage self-play training paradigm can be used to improve the robustness and generalization ability of the model. In addition, an auxiliary training module focused on transition dynamics is added, which can significantly improve the adaptability to environmental uncertainty. The framework uses feature-level and agent-level attention mechanisms for action and value generation, respectively. Furthermore, a large number of comparative experiments and practical deployments have been conducted, all of which demonstrate the superiority and practicality of the proposed framework.
[0008] According to a first aspect of the present invention, a task robot for a deep reinforcement learning system is provided, wherein the deep reinforcement learning system described in the specification (page 1 / 14, CN 121893303 A) can achieve robust multi-robot swarming using asymmetric self-play. The task robot includes a reinforcement learning control module and an auxiliary training module. The reinforcement learning control module dynamically adjusts the robot's behavior to help it make optimal decisions and continuously learns and improves through a real-time feedback mechanism. The module includes a target navigation model, a swarm behavior maintenance model, and a collision avoidance model. The target navigation model guides the task robot toward a predetermined target area.The distance between the task robot and the target area is continuously calculated, and action commands are generated based on the calculated information to guide the task robot to shorten the distance to the target area. A swarm behavior maintenance model is used to maintain swarm behavior and monitor the relative position and distance of other robots relative to the task robot, so that the swarm behavior maintenance model adjusts the motion of the task robot to maintain coordination within the robot group. A collision avoidance model is used to avoid collisions, and it adjusts the trajectory of the task robot according to the relative position of obstacles and interfering robots. An auxiliary training module is used to enhance the environmental perception ability of the task robot, enabling the task robot to predict dynamic changes around it. The auxiliary training module includes: a local environment mesh estimation model and a motion prediction model. The local environment mesh estimation model is used to generate a small-scale environmental mesh map at each time step to show the state of obstacles, interfering objects and other robots around the task robot, providing contextual information for the task robot's subsequent decisions. The motion prediction model is used to analyze and predict the trajectories of other robots and the properties of obstacles and interfering objects based on the generated environmental mesh map.
[0009] According to a second aspect of the present invention, a deep reinforcement learning system is provided that can perform robust multi-robot swarming based on asymmetric self-game. The deep reinforcement learning system comprises multiple interfering robots, multiple static obstacles, a starting region, a target region, and multiple task robots. Task robots can be activated and move from the starting region to the target region, bypassing interfering robots and static obstacles. Each task robot is equipped with physical communication devices, enabling real-time data transmission and information sharing, allowing for coordination and cooperation with other task robots. Each task robot includes a reinforcement learning control module and an auxiliary training module. The reinforcement learning control module dynamically adjusts the robot's behavior to help it make optimal decisions and continuously learns and improves through a real-time feedback mechanism. The reinforcement learning control module includes a target navigation model, a swarm behavior maintenance model, and a collision avoidance model. The target navigation model guides the task robots to move towards the target region, continuously calculates the distance between the task robots and the target region, and generates action commands based on the calculation results to guide the task robots to shorten the distance to the target region. The swarm behavior maintenance model maintains swarm behavior and monitors the relative positions and distances of other robots relative to the task robots, allowing the swarm behavior maintenance model to adjust the movement of the task robots to maintain consistent coordination within the robot group. The collision avoidance model avoids collisions by adjusting the trajectory of the task robots based on the relative positions of interfering robots and static obstacles. The auxiliary training module enhances the environmental perception capabilities of the task robot, enabling it to predict dynamic changes in its surroundings. The module includes a local environment mesh estimation model and a motion prediction model. The local environment mesh estimation model generates a small-scale environmental mesh at each time step.The graph shows the state of interfering robots, static obstacles, and other robots around the task robot, providing contextual information for the task robot's subsequent decisions. A motion prediction model is used to analyze and predict the trajectories of other robots and the properties of interfering robots and static obstacles based on the generated environmental mesh map.
[0010] Based on the above configuration, the proposed solution in this paper is a swarm control framework (ASFC) based on asymmetric self-game and deep reinforcement learning (DRL), designed to enhance policy robustness in dynamic scenarios. In this framework, dynamic obstacles are treated as adversarial interference sources, and learnable policies evolve synchronously with the swarm policy. This asymmetric adversarial training process is like a natural curriculum, enabling the swarm policy to gradually improve the framework performance by coping with increasingly complex interference sources. Furthermore, the proposed framework can also assume the existence of communication-constrained environments to enhance its practical applicability.
[0011] The contributions of this invention are summarized as follows:
[0012] (I): A novel swarming framework based on two-stage asymmetric self-game is proposed, which can significantly improve policy robustness and generalization ability, laying an important foundation for the application of swarm control in complex scenarios.
[0013] (II): An innovative self-supervised auxiliary training module is integrated into the deep reinforcement learning (DRL)-based architecture to enhance adaptability to environmental uncertainty.
[0014] (III): Feature-level and agent-level attention mechanisms are introduced for action and value generation, respectively, making the proposed method scalable for swarms of arbitrary sizes.
[0015] (IV): Extensive ablation studies and comparative experiments with state-of-the-art methods demonstrate the superiority of the proposed framework. Furthermore, physical experiments on a robotic platform further verify the practicality of the proposed framework.
[0016] The embodiments of the present invention will be described in more detail below with reference to the accompanying drawings, wherein:
[0017] Figure 1 shows a partial observation diagram of the robot;
[0018] Figure 2 shows a network architecture diagram of ASFC proposed according to an embodiment of the present invention;
[0019] Figure 3 shows a two-stage cluster training algorithm 1 with auxiliary loss proposed according to the present invention;
[0020] Figure 4 shows Table I, which contains the parameters used during training;
[0021] Figure 5 shows the reward performance diagram of each model in the same test scenario, and the model is saved and evaluated every 2000 training rounds during training;
[0022] Figure 6 shows Table II, which is used to record the corresponding indicators;
[0023] Figure 7 shows the trajectory diagram of the ASFC method in three scenarios, wherein the RAND scenario includes (a-1), (a-2), (a-3) DWA scenarios include (b-1), (b-2), and (b-3), and HEUR scenarios include (c-1), (c-2), and (c-3);
[0024] Figure 8 shows snapshots of four swarm robots using ASFC and two heuristic adversarial interferences, including parts (a), (b), (c), and (d);
[0025] Figure 9 shows a schematic structural diagram of a deep reinforcement learning system according to an embodiment of the present invention, which can achieve robust multi-robot swarming based on asymmetric self-game; and
[0026] Figure 10 shows a schematic diagram of the architecture of a task robot according to an embodiment of the present invention. Detailed Description
[0027] In the following description, a deep reinforcement learning system for achieving robust multi-robot swarming based on asymmetric self-game will be described by preferred examples. Those skilled in the art will understand that various modifications, including additions and / or substitutions, can be made without departing from the scope and spirit of the present invention. Specific details may be omitted so as not to make the present invention difficult to understand; however, this disclosure is intended to enable those skilled in the art to practice the teachings herein without excessive experimentation.
[0028] To facilitate understanding of the technical content of the present invention, the following description first provides an introduction to related work in the field, including an introduction to swarm control based on deep reinforcement learning (DRL) and asymmetric self-game techniques.
[0029] Related Work – (A): Swarm Control Based on Deep Reinforcement Learning (DRL)
[0030] Deep reinforcement learning (DRL) has become a potential solution for achieving swarm control in a variety of situations. Q-learning-based methods have been proposed to enable a group of unmanned aerial vehicles (UAVs) to learn to swarm in a leader-follower topology. In other works, single-agent deep reinforcement learning (DRL) and local situation maps considering collision risk are applied to generate collision-free policies. Learning-based frameworks have also been proposed to adapt to scenarios with static obstacles. Furthermore, demonstrations of non-expert prior policies can be used to improve sample efficiency. In some works, methods based on sequential attention mechanisms and behavioral reasoning have also been introduced to handle random communication environments. In addition, some specifications, page 3 / 14, CN 121893303 A, design a swarming algorithm based on deep reinforcement learning (DRL) to eliminate communication dependencies. To solve dynamic obstacles, current techniques have also developed graph attention-based networks, but the motion patterns of the dynamic obstacles they propose are still quite simple, which limits the model's generalization ability. In contrast, in this invention, under the assumption of non-communication and allowing the dynamic obstacles to behave in a complex and aggressive manner, an effective swarming framework that exhibits robustness to environmental disturbances is proposed.
[0031] Related work – (B): Asymmetric self-game
[0032] Self-game mechanisms have become a training method adopted by many well-known systems, such as AlphaGo and Google Football. This mechanism can stimulate the intelligence of an agent by training it against a learnable opponent. Asymmetric self-games mean that the opponent and the agent have different abilities and goals. In some works, the existing schemes include: visual trackers are trained using learnable targets that they are trying to escape, thereby improving tracking strategies. In addition, in other works, the existing schemes include: using a cooperative team consisting of one target and multiple distractors to compete against the tracker to enhance its robustness. In multi-agent environments, learnable base runners and heterogeneous capture systems can be trained together to improve capture skills.
[0033] Based on these related works, this disclosure proposes a novel two-stage self-game training paradigm: the first stage focuses on enhancing model performance, while the second stage aims to promote model generalization.
[0034] Next, the problem formula and system model are provided. First, a detailed statement of the identified problem and its settings is given, followed by a detailed description of the observations and operations.
[0035] Problem Formula and System Model – (A): Problem Statement
[0036] This disclosure aims to enable a group of differential wheeled robots to maintain a swarm motion pattern and navigate from a starting area to a target area. The application scenario includes not only static obstacles but also adversarial and learnable disruptors represented by another group of robots. These robots are not allowed to communicate with each other and must make distributed decisions based solely on their local observations.
[0037] For each robot, the sub-objectives of swarm control are as follows: (1) minimize navigation time; (2) minimize distance from the swarm center and reduce heading deviation relative to other robots; (3) avoid collisions with neighbors, static obstacles, and disruptors. In contrast, in adversarial training, the disruptor's objective is to collide with the robots, thereby disrupting the swarm behavior. It is noteworthy that both the robot team and the disruptor are scalable, which means that the control strategy must be swarm-invariant. In the context of reinforcement learning (RL), the objective problem can be formulated as a decentralized partially observable Markov decision process (Dec-POMDP).
[0038] Problem Statement and System Model – (B): Observation Space
[0039] In this paper, two problems will be discussed, including: (1) robot observation and (2) observer observation.
[0040] (1) Robot Observation: Figure 1 illustrates a local observation map of robot i. In addition to the vector format input, a local grid map is used to represent the surrounding environment centered on robot i. From the perspective of robot i, as shown in Figure 1, the observation can be represented asWhere si:=(vi,ω i) are its linear velocity and angular velocity, and pi:=(ρi, αi) represents the target's relative position (i.e., distance ρi and angle αi). It is important to note that ρi is the only non-local component in the observation that can be obtained through a global positioning system. Based on a widely used grid-like sensor, represents a three-channel local grid map that can represent the surrounding environment centered on itself for the most recent three time steps. Here, is the most recently updated grid map, and Ng is the input length. The resolution is 0.12 meters. In each channel of , each entity category (i.e., free space, static obstacles, robots, and interfering objects) is represented by 0, 1, 2, and 3 respectively (represented by different colored blocks in Figure 1). and represent the observable states of other clusters of robots (Ni) and interfering objects (Mi) within the field of view, respectively. Specifically, for a movable entity j (robot or interfering object) within the field of view, ξi,j is defined as ξi,j:=(di,j,ψi,j,φi,j), where di,j,ψi,j and φi,j represent the relative distance, relative angle and azimuth difference between entity j and robot i, respectively.
[0041] It is worth noting that the given observations are distributed, localized and self-centered configurations, where all quantities except velocity are relative in order to generalize to different maps. It is also worth mentioning that the number of elements in sets Ni and Mi is variable, and can even be zero, indicating that the network used should be able to handle a variable number of neighbors and interfering sources.
[0042] 2) Observation of interfering sources: For interfering source k, its observation is represented as where the meaning of the first two elements is the same as that of the reference robot's observation. The robot set and interfering source set of interfering source k contain all robots and interfering sources in the global scope rather than the local scope, meaning that the observable range of the interfering source is greater than the observable range of the robot.
[0043] Problem Statement and System Model – (C): Action Space
[0044] In the proposed system, both the robot and the disruptor execute steering commands consisting of linear and angular velocities. The action space is designed as a set of velocity options. Specifically, for each robot, the action space contains four linear velocity options: {0, 0.15, 0.3, 0.5} m / s, and nine angular velocity options: {-2, -1.2, -0.8, -0.3, 0, 0.3, 0.8, 1.2, 2} rad / s. For the disruptor, the angular velocity options are the same as those for the robot, while the linear velocity options are {0, 0.15, 0.3, 0.38} m / s. The maximum velocity difference between the robot and the disruptor is to compensate for the difference in difficulty of their respective tasks. Therefore, the action spaces of both the robot and the disruptor contain 36 different velocity combinations.
[0045] The execution method will be discussed next. Figure 2 shows the network architecture diagram of ASFC proposed according to an embodiment of the present invention. Feature-level and agent-level attention mechanisms are used for action and value generation, respectively, thereby realizing the scalability of ASFC. In addition, an auxiliary training module is also designed to predict the next local map (see the left part). It should be noted that the left and right parts are only used during training to simplify policy learning, and are inactive during execution.
[0046] This section provides details of ASFC, including the proposed network architecture (refer to Figure 2), auxiliary training module, reward function and the design of the two-stage training paradigm. Here, there are four key points, including: (A) action and value learning; (B): auxiliary training module; (C): reward function; and (D): two-stage self-game training.
[0047] (A) Action and value learning
[0048] The proposed network design takes into account two factors. The first is that it should be able to handle a variable number of robots and interfering entities; the second is that it is expected to meet the Centralized Training and Decentralized Execution (CTDE) paradigm, because integrating additional information into the critique network has been shown to effectively simplify training. For the first consideration, an attention mechanism is adopted as its population and permutation invariance. For the second consideration, an egocentric global agent-level communication module is designed to enhance the information of the critique network.
[0049] 1) Action learning: The action learning of the proposed architecture can be shown in the middle part of Figure 2. For robot i, the local segmentation map and (si,pi) are embedded by a convolutional neural network (CNN) and a multi-layer perceptron (MLP), respectively, and then the outputs are concatenated as self-feature qi. Next, based on the query-key-value mechanism, the information of robot feature set and interfering entity feature set within its perception range is aggregated using robot i's self-feature qi as the query source. The robot feature embedding and the interfering object feature embedding based on the self-feature qi of robot i can be calculated as follows:
[0050] Where and are trainable parameters, and the coefficients and are normalized importance scores, which can be expressed as:
[0051] Where and are scaling factors; are trainable parameters; σ is the Softmax function.
[0052] Then, the robot feature vector, the interfering object feature vector, and the self-feature qi are concatenated to form local features ei, which serve as the input source for the actor network implemented by a two-layer multilayer perceptron (MLP). The action ai can be based on the actor...The output probability of the last layer of the network is sampled using the SoftMax function. In addition, the network architecture of the interferer is basically the same as that above. The main difference is that the input of the interferer k is global, and its features ek are respectively passed through the policy branch and the criticism branch implemented by the two-layer multilayer perceptron (MLP) to generate the interferer's actions and values.
[0053] 2) Value learning: In order to enhance value learning, all robots can exchange agent-level information with each other to generate values. It should be noted that during the execution, the policy network still follows the no-communication assumption. As shown on the right side of Figure 2, an agent-level attention mechanism is introduced, which enables the robot to aggregate global features by learning the importance of other robots. Specifically, from the perspective of robot i, the local total features are obtained by connecting its local features ei and the cluster features fi. is the cluster-specific information, which consists of the relative distance and relative angle between the cluster center point and robot i. Then, robot i can receive local features from other robots and calculate the relative importance of any robot as follows:
[0054] where is the scaling factor, WQ is the query parameter, and WK is the key parameter. The aggregated information Ci of robot i can be obtained according to the following algorithm:
[0055] where WV is the value parameter.
[0056] Ci and can be connected and then the value vi can be generated through the evaluation network implemented by a two-layer multilayer perceptron (MLP). It is worth noting that ei in the policy network and the evaluation network are obtained from networks with the same structure but different parameters. The proximal policy optimization (PPO) algorithm can be used to optimize the action learning neural network and the value learning neural network, where the corresponding learning rates are ηp and ηc, respectively.
[0057] (B) Auxiliary training module
[0058] In some embodiments, at each time step, the swarm robot and the disruptor can simultaneously apply distributed actions to the environment to determine the transition to a new global state. This results in the environment of each robot being highly uncertain, so it is necessary to help the robot improve its understanding of the scene dynamics. Inspired by adversarial modeling, an auxiliary training module can be designed to predict the local mesh map of the next time step, as shown on the left side of Figure 2. Auxiliary signals and reinforcement learning signals can be combined to jointly optimize the policy network.
[0059] Specifically, the local features ei of the policy network, in addition to being used to generate actions ai (as mentioned above), can also be used to generate estimates of the next local grid map. The decoder branch, which can be composed of an MLP and multiple deconvolutional layers, is constructed based on (ei, ai). Since the local map contains four categories, similar to semantic segmentation techniques,The construction loss can be defined as the cross-entropy loss:
[0060] where is the true category of the label at position (x,y) (which can be obtained from the simulator), C is the category number, and I(·) is the binary indicator function.
[0061] The total loss during training can be composed of the above loss and the relevant loss of PPO, and the network is jointly optimized in each mini-batch iteration of the PPO algorithm, as shown below: Ltotal=βpL pol+βvL val -βeL ent+βaL ce …(6) Lent=- Σ πθ(oi)log(πθ(oi))…(9)
[0062] Where βp, βv, βe and βa are the corresponding coefficients; Lpol is the policy loss; Lval is the value loss; Lent is the policy entropy; γt is the discount factor; tmax represents the length of a game round; V(oi) is the state value function; ζ(θ) represents the policy change rate; ∈ is the boundary hyperparameter of ζ(θ); represents the estimate of generalized advantage; πθ(oi) represents the policy with parameter θ.
[0063] In this way, the robot can continuously learn to predict the movement of surrounding robots and interfering objects, thereby mitigating the impact of environmental uncertainty, which can be demonstrated by the provided ablation analysis. This module itself is scalable because it can run in a fixed-size local representation and is independent of the number of agents. In addition, the auxiliary training module and the value learning component are only used during the training phase and are used to assist policy learning. They are removed during execution to ensure that the system settings are not affected.
[0064] (C) Reward Function
[0065] Here, two key points need to be discussed: (1) robot reward and (2) interferer reward.
[0066] (1) Robot reward:
[0067] The design of the reward function in reinforcement learning is highly related to the task objective. Based on the various sub-objectives of swarm control described in the above description, corresponding reward functions are proposed to guide reinforcement learning training:
[0068] The first term is used to encourage the robot to move towards the target area and is defined as: Specification 7 / 14 Page 11 CN 121893303 A
[0069] where is the distance from robot i to the center of the target area; Dtar determines whether navigation is completed; λtar and λapp are parameters.
[0070] The second term is introduced to incentivize robots to maintain swarm behavior:
[0071] where and represent the rewards for swarm center localization and swarm orientation consistency, respectively; is the distance between robot i and the swarm center; is the average deviation of robot i's orientation from the orientations of other robots; Dctr and Φori are the thresholds for swarm center localization and orientation deviation, respectively; λctr and λori are both positive parameters. The design indicates that the robot will only receive a reward if its position and orientation meet expectations during the swarming process.
[0072] The third termFor collision avoidance, it is defined as:
[0073] where λcoll, λint and λobs are parameters; the second and third rows are used to indicate that the robot maintains a safe distance from the interfering object and the obstacle, respectively; and represent the shortest distance between robot i and the obstacle and the interfering object, respectively, and the corresponding thresholds are Dint and Dobs, respectively.
[0074] 2) Interferer reward: The design of the interferer reward is relatively simple. Its principle is to encourage the interferer to approach the swarm robot and collide with it. The reward of the interferer k can be defined as follows:
[0075] where represents the distance between the interferer k and robot i; λad, λad and λna are parameters.
[0076] (D) Two-stage self-game training
[0077] In the self-game training of the technical solution proposed in this invention, there are two objectives. First, the interferer should have sufficient intelligence to stimulate the robot's skill learning; second, the swarm strategy needs to have generalization ability. For this purpose, a novel two-stage self-game training strategy is proposed. In the first stage, to enhance skills, both the robot and the disruptor are trained simultaneously, and the network parameters ωT of the disruptor are added to the model pool W every NT rounds. In the second stage, to enhance generalization ability, the swarm strategy obtained in the first stage is trained on the entire disruptor model pool W containing different levels of adversarial intelligence.
[0078] In the second training stage, two techniques are used to enhance environmental diversity and integrate it into the curriculum learning. The first technique is agent sampling of the disruptor models. Assuming there are Nk disruptors in the team and W has NW models, the number of all possible combinations of the disruptor team is greatly increased, which greatly increases the diversity of the environment. The second advantage is that the disruptor models are sampled according to periodically updated weights, that is, gradually focusing on stronger models. Specifically, the disruptor team is updated every Ns rounds using new weighted samples in W. The probability of a model being selected depends on its performance when it was last selected, that is, its average cumulative reward
[0079] where Rmin and Rmax are the lower and upper bounds, respectively. In the second stage W, the initial reward for each model is assigned the value Rmin. Intuitively, this iterative weighted sampling can gradually increase the difficulty of the clustering scenario. Algorithm 1 in Figure 3 summarizes the two-stage clustering training method with auxiliary loss proposed in this invention. Algorithm 1 is presented in code.
[0080] The experimental results are as follows: simulation experiments in various scenarios demonstrate the superiority of the proposed framework, and physical experiments verify its practicality.
[0081] Experiment - (A) Experimental setup
[0082] 1. Configuration: Simulation was performed on a server with configured CPU and GPU using the 3D simulator PyBullet. In the two-stage...In the swarm training, the scene size in the first stage is 15m×15m, containing 5 swarm robots, 3 interfering objects and 2 static obstacles; the scene size in the second stage is 25m×25m, containing 8 robots, 6 interfering objects and 5 static obstacles. The robots and interfering objects are all modeled as disks with a radius of 0.18 meters. The obstacle energy collection is composed of a cylinder with a diameter of 1 meter and a cube with a side length of 1 meter. In each scene, the swarm robots are initialized within 1 meter of a randomly selected starting point near the edge of the scene, while the target point is the symmetrical point of the starting point about the center of the scene. The interfering objects and obstacles are randomly initialized throughout the scene. When the robot moves to within 1 meter of the target point, it is considered to have completed a scene. The training parameters are shown in Table 1 in Figure 4.
[0083] 2. Ablation baseline: In order to verify the contribution of the ASFC key components, the following three variants are compared. First, to demonstrate the advantages of self-game training, an ablation method, denoted ASFC-S, was developed, in which the interfering player is controlled by a heuristic adversarial strategy (denoted as HEUR strategy), while the swarm strategy is configured in the same way as ASFC. In this way, each interfering player will actively move towards the nearest swarm robot. Second, to demonstrate the effectiveness of two-stage training, a method without a second stage was used, called ASFC-G. Throughout the process, training can be performed simultaneously with the interfering player, and the total number of training rounds is the same, where the environment is configured in the same way as the second stage of ASFC. In addition, to demonstrate the contribution of auxiliary loss, the implementation of ASFC-A does not use an auxiliary training module. The remaining configurations are the same as ASFC.
[0084] 3. Ablation Baseline: Performance Metrics:
[0085] The following metrics were used to evaluate the model from multiple perspectives: 1. Collision Avoidance Metrics: Collision rate (CR) and private space disturbance rate (PR) were used to evaluate the collision avoidance effect. CR represents the average probability of a robot colliding with other entities, and PR represents the average probability of a robot being less than 0.6 meters from the nearest interfering object. 2. Swarming Behavior Metrics: Flocking centering distance (FCD) and flocking orientation deviation (FOD) are introduced to evaluate swarming performance. For each robot, FCD is its average distance to the swarm center, and FOD is its average orientation deviation from other robots. 3. Target Achievement Metrics: Average speed (AS) and extra distance rate.The target arrival rate (EDR) is used to evaluate the target arrival rate. AS is obtained by dividing the trajectory length by the navigation time. EDR represents the percentage of redundant distance in the entire training round trajectory. When the EDR value is large, it means that the robot has traveled a long distance.
[0086] Experiment - (B) Training Details
[0087] For all models, each stage of the two-stage training had 10,000 training rounds. To demonstrate the learning process, the model was evaluated every 2,000 training rounds in the same test scenario. The test scenario was 25m × 25m in size and contained 8 robots, 6 interfering objects and 5 static obstacles. For each saved model, 100 training rounds were run for each evaluation. For each test training round, the strategy for each interfering object was randomly selected from three strategies, including: HEUR strategy, Dynamic Window Approach (DWA) (DWA strategy) and random motion strategy (RAND strategy) (RAND strategy). The latter two strategies are non-adversarial. The DWA strategy only considers the existence of entities other than robots, which means that the responsibility of avoiding collisions is handed over to the swarm robots. Figure 5 shows the reward performance of each model in the same test scenario, and the model is saved and evaluated every 2000 training rounds during training. Figure 5 shows the average reward of the swarm.
[0088] First, it can be observed that ASFC obtained the highest reward, proving its superiority. Second, compared with ASFC-S, ASFC obtained a higher reward due to the use of the self-game training mechanism. The reward of ASFC-S hardly increased after 12000 training rounds, indicating that the interference of fixed behavior has limited effect on the improvement of swarm performance. In addition, ASFC outperformed ASFC-G, further proving that the proposed two-stage training mechanism enhances the generalization ability to complex scenarios. In addition, the effectiveness of the auxiliary training module was also confirmed in the comparison between ASFC and ASFC-A.
[0089] Experiment - (C) Comparative Experiments in Different Scenarios
[0090] This section evaluates the performance of each model under different interference strategies. Three scenarios were named RAND, DWA, and HEUR, respectively, with corresponding interference strategies of RAND, DWA, and HEUR. Each scenario was 25m × 25m in size and contained 10 swarm robots, 8 interfering objects, and 5 static obstacles. Each evaluation consisted of 100 tests. Furthermore, to demonstrate the superiority of the proposed method, it was compared with three state-of-the-art methods: 1) MADDPG; 2) MAAC; and 3) APF. The first two learning-based methods were trained using the same self-game mechanism to ensure fairness. APFThis is a traditional method where the target point and cluster center have attractive potential fields, while the robot, static obstacles, and interfering objects have repulsive potential fields. Figure 6 shows Table II, which records the corresponding metrics mentioned above (best performance is highlighted in bold).
[0091] The results show that ASFC has the lowest CR and the smallest FCD in all scenarios, demonstrating its advantages. In addition, ASFC also has the lowest EDR, which is attributed to the robot's ability to make decisions that are both safe and not overly conservative. Meanwhile, the results show that the auxiliary loss used to model environmental dynamics helps to improve the robustness of the cluster. Furthermore, ASFC achieves better results even when using only local observations compared to APF which uses global information. In addition, the ASFC sampling behavior for each test scenario is provided, as shown in Figure 7. Figure 7 shows the trajectory plots of the ASFC method in three scenarios, where the RAND scenario includes (a-1), (a-2), and (a-3), the DWA scenario includes (b-1), (b-2), and (b-3), and the HEUR scenario includes (c-1), (c-2), and (c-3). Each scenario presents trajectories for three time frames. The swarm robot FR is represented by a trajectory, and the interfering object IF is also represented by a trajectory. The results show that the robot can avoid the influence of various interfering objects while effectively maintaining the desired swarm behavior. In addition, it is noted that once away from the interference, the robots gradually form a denser and more consistent swarm.
[0092] Experiment - (D) Real Experiment Demonstration
[0093] In addition to the simulation experiment, the proposed ASFC can be deployed in a physical robot to verify the practicality of the proposed method. Specifically, a 4m×4m scenario can be constructed, which includes four robots, two interfering objects, and two static obstacles. An AgileX LIMO mobile robot with a four-wheel differential steering mode is used as the robot and the interfering object. The ASFC trained in the simulation can be used as the swarm strategy for the real experiment, while the interfering object follows the HEUR strategy. A wireless communication system built by a router connects the robot as a mobile node to the control station. The control station is equipped with a CPU and GPU as data transceivers for all robots. The control station receives inference-related data from the robot, calculates actions, and then transmits them to the relevant robot. This setup does not violate the distributed execution design employed, as the robot's onboard processor can be modified for deep learning inference.
[0094] The employed model runs on the robot platform, and the process is recorded using a camera. Figure 8 shows snapshots of four swarming robots employing ASFC and two heuristic adversarial interferences, including portions (a), (b), (c), and (d). Figure 8 provides snapshots of training rounds, showing that the robots are able to avoid interference while maintaining swarming until reaching the target area.Domain. Furthermore, the computation time of the model used is 2.25 milliseconds per run, far less than the system's 200 millisecond decision interval. These experiments demonstrate that the proposed framework meets both the requirements for simulation-to-real-world deployment and real-time requirements, indicating its potential for practical application.
[0095] Figure 9 shows a schematic structural diagram of a deep reinforcement learning system 100 according to an embodiment of the present invention, which can achieve robust multi-robot swarming based on asymmetric self-games. The deep reinforcement learning system can serve as a platform for training robots whose task is to maintain a swarm movement pattern and navigate from a starting region to a target region. The deep reinforcement learning system includes multiple task robots 110, multiple interfering robots 120, and static obstacles 130. In the deep reinforcement learning system, a target region 140 can be set, and the task robots 110 are configured to navigate from the starting region (i.e., their initial location) to the target region 140 while still maintaining a swarm movement pattern. The arrangement of items in Figure 9 is for illustrative purposes only and is intended to show the types of objects that can be placed during training.
[0096] Each task robot 110 is equipped with physical communication devices to enable real-time data transmission and information sharing, thereby ensuring coordination and cooperation with other task robots 110. In various embodiments, the task robot 110 is equipped with sensors, including lidar, cameras, and ultrasonic sensors, which enable the robot to acquire real-time information about its surrounding environment, facilitating obstacle detection and location recognition, thereby enhancing its autonomous navigation and decision-making capabilities. Furthermore, all task robots 110 are equipped with modules or models for learning to adapt to changing conditions. The modules and models used will be further described below.
[0097] The interference robot 120 is configured to dynamically interfere with the movement of the task robot 110, thereby creating challenges in its environment. Static obstacles 130 also pose a challenge to the movement of the interference robot 120. The task robot 110 can be trained to adapt and find better navigation methods. This interaction helps the task robot 110 improve its navigation and decision-making skills, enabling it to complete its tasks even in the face of interference. By learning to avoid interfering robots 120 and static obstacles 130, task robots 110 can become more efficient and flexible in the real world.
[0098] As shown in FIG2, the ASFC network architecture can be applied to the modules and models equipped in task robots 110, enabling them to function during the training phase. Specifically, FIG10 shows a schematic diagram of the architecture of a task robot 110 according to an embodiment of the present invention. Each task robot 110 includes a reinforcement learning control module 112, an auxiliary training module 114, a reward function module 116, and a two-stage self-adversarial training module 118. These modules communicate with each other to realize the transmission, reception, and exchange of information.Furthermore, these modules are connected to multiple sensors of the task robot 110, enabling it to receive environmental data in real time and adjust the robot's navigation strategy accordingly.
[0099] The reinforcement learning control module 112 is used to dynamically adjust the robot's behavior to help it make optimal decisions. The reinforcement learning control module 112 can continuously learn and improve through a real-time feedback mechanism and includes a target navigation model, a swarm behavior maintenance model, and a collision avoidance model.
[0100] The target navigation model is responsible for guiding the task robot 110 to move towards a predetermined target area (e.g., target area 140). Through the target navigation model, the task robot 110 continuously calculates its distance from the target and generates action commands based on the calculated information, thereby guiding the task robot 110 to shorten the distance from the target to ensure that it always moves towards the target during task execution.
[0101] The swarm behavior maintenance model is responsible for maintaining swarm behavior. In order to ensure that multiple task robots 110 can coordinate their movements, the swarm behavior maintenance model monitors the relative positions and distances between the task robots 110. For task robot 110, the swarm behavior maintenance model adjusts its movement to maintain coordination within the group.
[0102] The collision avoidance model is responsible for avoiding collisions. Task robot 110 perceives its surroundings, identifies potential obstacles and disturbances through the collision avoidance model, and generates a collision avoidance strategy. The collision avoidance model adjusts the trajectory of task robot 110 based on the relative positions of obstacles (e.g., disturbance robot 120 and static obstacle 130) to avoid collisions and ensure safety.
[0103] Furthermore, reinforcement learning control module 112 uses the proximal policy optimization (PPO) algorithm to optimize action learning and value learning in the neural network, and combines a convolutional neural network (CNN) and multilayer perceptron (MLP) architecture. The framework adopted by reinforcement learning control module 112 can use a feature-level attention mechanism for action learning and also uses an agent-level attention mechanism for value learning. Action learning allows each task robot 110 to make independent decisions through repeated training, accumulate experience and improve its strategy in a distributed execution mode, thereby reducing dependence on central control. Simultaneously, value learning occurs during centralized training, where all task robots 110 can share environmental data and policy feedback to optimize decision-making. Using this collective information-sharing model helps the task robots 110 better understand their surroundings and improve coordination. By integrating these learning methods, the reinforcement learning control module 112 enhances collaboration and autonomous decision-making in complex environments.
[0104] The auxiliary training module 114 enhances the environmental perception capabilities of the task robot 110, enabling it to better predict dynamic changes around the task robot 110. The auxiliary training module 114 includes a local environment grid estimation model, a motion prediction model, and a decision optimization model.
[0105] The local environment mesh estimation model generates a small-scale environment mesh map at each time step, which can display obstacles, disturbances, and the status of other task robots 110 around the robot itself. The environment mesh map can provide important background information for subsequent decision-making.
[0106] The motion prediction model can analyze and predict the trajectories of other task robots 110 and the behavior of obstacles and disturbances based on the generated environment mesh map. This prediction allows the task robot 110 to formulate strategies in advance, thereby minimizing unexpected events and ensuring the smooth execution of the task.
[0107] The decision optimization model can provide environmental information and pass the environmental information data to the reinforcement learning control module 112 to assist in decision optimization. By predicting environmental changes, the task robot 110 can better adapt to different situations, thereby improving the overall task efficiency.
[0108] The reward function module 116 is designed to reward the task robot 110 for performing beneficial actions during training and establishes multiple reward mechanisms to encourage the task robot 110 to achieve specific goals. In one embodiment, the reward function module 116 includes a reward model.
[0109] When the task robot 110 meets certain conditions, the reward model provides feedback to it in the form of target approach reward, formation maintenance reward, and obstacle avoidance reward. For example, when the task robot 110 approaches the target area 140, the reward function module 116 provides a positive reward according to the reward model to encourage the robot to continue moving forward. Similarly, when the task robot 110 maintains a proper position in the swarm, the robot receives a reward from the reward model, thereby encouraging cooperation. In addition, when the task robot 110 successfully avoids collisions with the interfering robot 120 and the static obstacle 130, the robot receives a reward, thereby promoting safe navigation.
[0110] Furthermore, in one embodiment, each interfering robot 120 is equipped with an interference reward model, which provides a mechanism to encourage the interfering robot to approach the swarm robots (e.g., the task robot 110). When the interfering robot 120 successfully collides with the swarm robots (e.g., the task robot 110), the robot receives a reward. This reward is designed to incentivize the interfering robot 120 to engage in combat with the task robot 110, thereby providing a more challenging training scenario.
[0111] The two-stage self-adversarial training module 118 is designed to enhance the adaptability of the task robot 110 in complex environments through adversarial training, and it comprises two stages.
[0112] The first stage is skill enhancement training. In this stage, the task robot 110 and the interfering robot 120 can train simultaneously and interact through the two-stage self-adversarial training module 118. This interaction helps the task robot 110 improve its skills.Its ability to cope with unexpected situations and complete tasks under interference. The network parameters of the interference robot 120 are periodically stored in the model pool, enabling the task robot 110 to be trained for different levels of adversarial intelligence.
[0113] The second stage is policy testing and optimization. In this stage, the two-stage self-adversarial training module 118 introduces increasingly complex adversarial models to test and optimize the policy of the task robot 110. By simulating more intelligent interference or challenging environments, the task robot 110 is forced to continuously adjust and improve its policy to improve its generalization ability. To enhance environmental diversity, interference models can be sampled from the model pool in different combinations and applied to the course learning. This process can gradually increase the complexity of the scenario, focusing on stronger interference factors to challenge the robot.
[0114] In some embodiments, the auxiliary training module 114 can assist the reinforcement learning control module 112 and the two-stage self-adversarial training module 118 by providing additional focused training or skill enhancement to improve the robot's capabilities. The reward function module 116 can help shape the decision-making of the reinforcement learning control module 112 by reinforcing the beneficial behaviors learned during interaction with the two-stage self-adversarial training module 118. The two-stage self-adversarial training module 118 can introduce dynamic and progressively increasing adversarial conditions to prompt the reinforcement learning control module 112 to adapt and improve, and this process can be based on input information from the reward function module 116 and the auxiliary training module 114.
[0115] While the above describes configuring modules / models on the task robot itself, in other embodiments, these modules / models can be hosted on a central server with storage capabilities. The task robot can be trained or perform tasks through communication with the central server, thereby allowing remote operation and control of the modules / models without direct installation on the task robot. This configuration allows for more efficient processing and centralized management of training or task execution. Similarly, the central server can be equipped with storage capabilities and communicate with the interfering robot to control its actions or behaviors. This setup allows the interfering robot to receive instructions or make adjustments remotely during training or task execution, thereby improving the efficiency of managing and coordinating the interfering robot without directly installing the modules / models on each component unit.
[0116] This disclosure proposes Adaptive Swarm Control (ASFC) and focuses on implementing swarm control in highly dynamic and communication-constrained environments. The disclosed two-stage self-game training method can continuously enhance swarm strategies through interaction with learnable and adversarial disruptors. Furthermore, the auxiliary training module can mitigate the impact of environmental uncertainty, ensuring more robust control. The effectiveness and practicality of the ASFC framework have also been verified through extensive simulations and real-world experiments. The proposed framework represents a significant advancement in swarm control and holds promise for application in complex and challenging environments.
[0117] In this disclosure, the two-stage self-adversarial training method enhances the robot's adaptability to complex environments by employing a structured training workflow. This training method enables the robot to improve its response to adversarial conditions and optimize its strategies without being overwhelmed by the complexity of the task. Furthermore, the method improves the overall robustness of the training process, enabling the robot to effectively manage resources and adapt to dynamic environments, thereby maximizing its performance and operational efficiency. Specification 13 / 14 pages 17 CN 121893303 A
[0118] The functional units and modules of the apparatus and methods according to embodiments of the present disclosure can be implemented using computing devices, computer processors, or electronic circuits, including but not limited to application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), microcontrollers, and other programmable logic devices configured or programmed according to the teachings of this disclosure. Those skilled in the art of software or electronics can readily write computer instructions or software code that execute in computing devices, computer processors, or programmable logic devices based on the teachings of this disclosure.
[0119] All or part of the methods according to the embodiments can be executed in one or more computing devices, including server computers, personal computers, laptops, mobile computing devices (e.g., smartphones and tablets).
[0120] Embodiments may include computer storage media, transient and non-transient memory devices storing computer instructions or software code, which can be used to program or configure computing devices, computer processors, or electronic circuits to perform any of the processes of the present invention. Storage media, transient and non-transient memory devices may include, but are not limited to, floppy disks, optical disks, Blu-ray discs, DVDs, CD-ROMs, magneto-optical disks, ROMs, RAMs, flash memory devices, or any type of medium or device suitable for storing instructions, code, and / or data.
[0121] Each functional unit and module according to the various embodiments may also be implemented in a distributed computing environment and / or cloud computing environment, wherein all or part of the machine instructions are executed in a distributed manner by one or more processing devices interconnected by a communication network, such as an intranet, wide area network (WAN), local area network (LAN), the Internet, and other forms of data transmission media.
[0122] The above description of the invention is provided for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations will be apparent to those skilled in the art.
[0123] These embodiments were chosen and described in order to best explain the principles of the invention and its practical application, thereby enabling those skilled in the art to understand the various embodiments of the invention and various modifications suitable for the particular intended use. Specification 14 / 14 pages 18 CN 121893303 A Figure 1 Specification Drawings 1 / 10 pages 19CN 121893303 A Figure 2, Appendix 2 / 10, Page 20 CN 121893303 A Figure 3, Appendix 3 / 10, Page 21 CN 121893303 A Figure 4, Appendix 4 / 10, Page 22 CN 121893303 A Figure 5, Appendix 5 / 10, Page 23 CN 121893303 A Figure 6, Appendix 6 / 10, Page 24 CN 121893303 A Figure 7, Appendix 7 / 10, Page 25 CN 121893303 A Figure 8, Appendix 8 / 10, Page 26 CN 121893303 A Figure 9, Appendix 9 / 10, Page 27 CN 121893303 A Figure 10, Appendix 10 / 10, Page 28 CN 121893303 A Title: A DEEP REINFORCEMENT LEARNING SYSTEM USING ASYMMETRIC SELF-PLAY FOR ROBUST MULTI-ROBOT FLOCKING Abstract: A tasked robot for a deep reinforcement learning system using asymmetric self-play for robust multi-robot flocking is provided. The tasked robot includes a reinforcement learning control module and an auxiliary training module. The control module dynamically adjusts the robot's behavior to optimize decisions, featuring a target navigation model, a cluster behavior maintenance model, and a collision avoidance model. auxiliary training module enhances environmental perception, enabling the robot to predict dynamicchanges. The auxiliary training module includes a local environment grid estimation model to generate small-scale maps and a motion prediction model to forecast robot and obstacle trajectories. Abstract
Claims
1. A task robot for a deep reinforcement learning system, the system utilizing asymmetric self-game to achieve robust multi-robot swarming, characterized in that, The task robot includes: A reinforcement learning control module is used to dynamically adjust the robot's behavior, helping it make optimal decisions, and continuously learn and improve through a real-time feedback mechanism. The reinforcement learning control module includes: A target navigation model is used to guide a task robot toward a predetermined target area, continuously calculate the distance between the task robot and the target area, and generate action commands based on the calculated information to guide the task robot to shorten the distance to the target area; A swarm behavior maintenance model is used to maintain swarm behavior and monitor the relative positions and distances of other robots with respect to the task robot, enabling the swarm behavior maintenance model to adjust the movement of the task robot to maintain coordination within the robot group; and A collision avoidance model is used to avoid collisions, and the collision avoidance model adjusts the trajectory of the task robot based on the relative positions of obstacles and interfering robots; and An auxiliary training module is used to enhance the environmental perception capability of the task robot, enabling the task robot to predict dynamic changes in its surroundings. The auxiliary training module includes: A local environment mesh estimation model is used to generate a small-scale environment mesh map at each time step to display the state of obstacles, interference, and other robots around the task robot, providing contextual information for the task robot's subsequent decisions; and A motion prediction model is used to analyze and predict the trajectories of other robots and the properties of obstacles and interference based on the generated environmental mesh map.
2. The task robot according to claim 1, wherein the reinforcement learning control module is further configured to use the proximal policy optimization (PPO) algorithm, combined with a convolutional neural network (CNN) and a multilayer perceptron (MLP) architecture, to optimize action learning and value learning in its neural network.
3. The task robot according to claim 2, wherein the framework adopted by the reinforcement learning control module can use a feature-level attention mechanism in the action learning and an agent-level attention mechanism in the value learning, wherein the action learning allows the assigned robot to make independent decisions through repeated training in a distributed execution mode, while the value learning is carried out in a centralized training process, wherein the assigned robot shares environmental data with other robots in the centralized training process, thereby realizing collective information sharing.
4. The task robot according to claim 1, further comprising a reward function module for rewarding the task robot for performing beneficial actions during training, wherein the reward function module comprises: A reward model is used to provide feedback to the task robot in the form of target approach reward, formation maintenance reward and obstacle avoidance reward when the task robot approaches the target area; When the task robot approaches the target area, the reward function module provides a positive reward based on the reward model, thereby encouraging the robot to continue moving forward. When the task robot maintains an appropriate position within the robot swarm, it receives a reward from a reward model, thereby encouraging collaboration. When the task robot successfully avoids collisions with the interfering robot and the static obstacle, the task robot will be rewarded to promote safe navigation.
5. The task robot according to claim 1, further comprising: A two-stage self-adversarial training module is used to enhance the adaptability of the task robot in complex environments through adversarial training, and the two-stage self-adversarial training module consists of a first stage and a second stage. In the first stage, the task robot and the interference robot are trained simultaneously through the two-stage self-adversarial training module and interact with each other through the two-stage self-adversarial training module. The network parameters of the interference robot are periodically stored in the model pool, enabling the task robot to be trained for different levels of adversarial intelligence. In the second phase for strategy testing and optimization, the two-phase self-adversarial training module introduces complex adversarial models to test and optimize the strategy of the task robot, and the interference models are sampled from the model pool in different combinations.
6. A deep reinforcement learning system based on asymmetric self-game theory for achieving robust multi-robot swarms, characterized in that, The deep reinforcement learning system includes: Multiple interfering robots; Multiple static obstacles; Starting region and target region; and Multiple task robots are activated to move from the starting area to the target area and bypass the interfering robot and the static obstacle; Each of the task robots is equipped with a physical communication device to enable real-time data transmission and information sharing, thereby enabling coordination and cooperation with other task robots. Each of the aforementioned task robots includes: A reinforcement learning control module is used to dynamically adjust the robot's behavior, helping it make optimal decisions, and continuously learn and improve through a real-time feedback mechanism. The reinforcement learning control module includes: A target navigation model is used to guide a task robot toward a predetermined target area, continuously calculate the distance between the task robot and the target area, and generate action commands based on the calculated information to guide the task robot to shorten the distance to the target area; A swarm behavior maintenance model is used to maintain swarm behavior and monitor the relative positions and distances of other robots with respect to the task robot, enabling the swarm behavior maintenance model to adjust the motion of the task robot. To maintain coordination within the robot group; and A collision avoidance model is used to avoid collisions, and the collision avoidance model adjusts the trajectory of the task robot based on the relative positions of the static obstacle and the interfering robot; and An auxiliary training module is used to enhance the environmental perception capability of the task robot, enabling the task robot to predict dynamic changes in its surroundings. The auxiliary training module includes: A local environment mesh estimation model is used to generate a small-scale environment mesh map at each time step to show the state of the interfering robots, static obstacles, and other robots around the task robot. Provide contextual information for the robot's subsequent decisions; and A motion prediction model is used to analyze and predict the trajectories of other robots, as well as the properties of the static obstacles and the interfering robots, based on the generated environmental mesh map.
7. The deep reinforcement learning system according to claim 6, wherein the reinforcement learning control module is further configured to use the proximal policy optimization (PPO) algorithm, combined with a convolutional neural network (CNN) and a multilayer perceptron (MLP) architecture, to optimize action learning and value learning in its neural network.
8. The deep reinforcement learning system according to claim 7, wherein the framework adopted by the reinforcement learning control module can use a feature-level attention mechanism in the action learning and an agent-level attention mechanism in the value learning, wherein the action learning allows the assigned robot to make independent decisions through repeated training in a distributed execution mode, while the value learning is carried out in a centralized training process, wherein all the task robots share environmental data and provide collective information sharing in the centralized training process.
9. The deep reinforcement learning system according to claim 6, wherein each task robot further comprises a reward function module for rewarding the task robot for performing beneficial actions during training, and includes: A reward model is used to provide feedback to the task robot in the form of target approach reward, formation maintenance reward and obstacle avoidance reward when the task robot approaches the target area; When the task robot approaches the target area, the reward function module provides a positive reward based on the reward model, thereby encouraging the robot to continue moving forward. Wherein, when the task robot maintains an appropriate position in the cluster of task robots, the task robot receives a reward from the reward model to encourage cooperation; When the task robot successfully avoids collisions with the interfering robot and the static obstacle, the task robot receives a reward to promote safe navigation.
10. The deep reinforcement learning system according to claim 6, wherein each task robot further comprises a two-stage self-adversarial training module for enhancing the adaptability of the task robot in complex environments through adversarial training, and the two-stage self-adversarial training module consists of a first stage and a second stage. in, In the first stage, the task robot and the interference robot are trained simultaneously through the two-stage self-adversarial training module and interact with each other through the two-stage self-adversarial training module. The network parameters of the interference robot are periodically stored in the model pool, enabling the task robot to train against different levels of adversarial intelligence. In the second phase for strategy testing and optimization, the two-phase self-adversarial training module introduces complex adversarial models to test and optimize the strategy of the task robot, and the interference models are sampled from the model pool in different combinations.
11. The deep reinforcement learning system of claim 6, wherein each of the interfering robots is equipped with an interference reward model that provides a mechanism to encourage the interfering robot to approach the task robot, and the interfering robot receives a reward from the interference reward model when any of the interfering robots successfully collides with one of the task robots.
12. The deep reinforcement learning system of claim 6, wherein each of the task robot and the interfering robot is used to execute steering commands including linear velocity and angular velocity, and each of the task robots has four linear velocity options in its action space, including: The linear velocity options are 0 m / s, 0.15 m / s, 0.3 m / s, and 0.5 m / s, and nine angular velocity options, including -2 rad / s, -1.2 rad / s, -0.8 rad / s, -0.3 rad / s, 0 rad / s, 0.3 rad / s, 0.8 rad / s, 1.2 rad / s, and 2 rad / s. For each of the interfering robots, the angular velocity options are the same as those of the task robot, while the linear velocity options are 0 m / s, 0.15 m / s, 0.3 m / s, and 0.38 m / s.
13. The deep reinforcement learning system according to claim 6, wherein the deep reinforcement learning system serves as a training platform, and through two-stage swarm training, the robot is trained to navigate from the starting area to the target area while maintaining a swarm movement pattern, wherein the scene size of the first stage is 15m × 15m, including five swarms of the task robot, three of the interference robots, and two of the static obstacles, wherein the scene size of the second stage is 25m × 25m, including eight swarms of the task robot, six of the interference robots, and five of the static obstacles.