Multi-robot cluster hunting method and system based on potential field enhancement reinforcement learning

Through the potential field enhancement reinforcement learning method designed by local perception and bionics, the problem of insufficient coordination of multi-robot clusters in complex environments is solved, efficient dynamic path planning and task execution are achieved, and learning efficiency and adaptability are improved.

CN120406469APending Publication Date: 2025-08-01SUN YAT SEN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510608459.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing multi-robot clusters have problems such as low learning efficiency, poor generalization ability and insufficient coordination in rounding up tasks in complex environments. Especially when unknown environments and communication are limited, traditional methods are difficult to perform tasks efficiently and collaboratively.

Method used

Using a potential field enhancement reinforcement learning method based on local perception, combined with bionics principles and cluster intelligent control technology, dynamic path planning and identity trade-off strategies are carried out through local perceptual information. The robot cluster works collaboratively in complex environments, adjusts motion trajectory and role allocation in real time, avoids collisions and efficiently pursues targets.

Benefits of technology

In complex environments with unknown or limited communication, robot clusters can efficiently coordinate tasks, improve learning efficiency and generalization capabilities, ensure task consistency and stability, and adapt to dynamic environment changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120406469A_ABST
    Figure CN120406469A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-robot cluster hunting method and system based on potential field enhanced reinforcement learning, and aims to solve the problems of low learning efficiency, poor generalization ability, insufficient coordination and the like of hunting tasks of a multi-robot cluster in a complex environment in the prior art. According to the method, by combining a potential field method and reinforcement learning, local sensing information between robots is utilized, attraction and repulsive force are calculated in real time, a motion track is dynamically adjusted, and efficient cooperation of clusters is ensured. Through a potential field enhancement mechanism, the robot can optimize path planning under guidance of obstacles and preys, collision is avoided, and a hunting task is completed. An identity tradeoff strategy is introduced, member roles in the cluster are dynamically adjusted according to task requirements and robot states, and autonomy and flexibility of the cluster are improved. Experimental results show that the method can effectively execute tasks in dynamic and complex environments, and has high adaptability, robustness and expandability. The method and the system can be widely applied to the fields of unmanned aerial vehicle formation, automatic chasing, intelligent monitoring and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to multi-robot cooperation and autonomous control technologies, and particularly to a multi-robot swarm hunting method and its application system based on potential field enhanced reinforcement learning. This method is applicable to multi-robots performing hunting tasks in complex environments, such as fields of automated pursuit, search and rescue, and UAV formations. Background Art

[0002] In multi-robot systems, swarm hunting tasks are widely applied in fields such as military, UAV formations, and driverless. Currently, swarm hunting problems usually adopt rule-driven strategies or game theory methods, but these methods have problems of low efficiency and poor adaptability in complex environments. Although existing reinforcement learning (RL) methods can solve such problems, they often face low learning efficiency and weak generalization ability, especially lacking effective initial strategy guidance during the training process, resulting in a slow learning process.

[0003] Therefore, how to effectively introduce external expert knowledge in reinforcement learning to improve learning efficiency and enhance the algorithm's generalization ability in different environments is the key challenge in solving multi-robot swarm hunting problems. Summary of the Invention

[0004] To solve the problems of low learning efficiency, poor generalization ability, and insufficient coordination in the hunting tasks of multi-robot swarms in complex dynamic environments in the prior art, the object of the present invention is to provide a multi-robot swarm hunting method and system based on potential field enhanced reinforcement learning. This method combines bionics, reinforcement learning, and swarm intelligent control technologies, and can ensure that multi-robot swarms can efficiently cooperate to perform hunting tasks in unknown environments, areas with dense obstacles, or communication-limited situations.

[0005] The present invention designs a potential field enhancement mechanism based on local perception through bionics principles, mimicking the movement patterns of animal groups in nature, enabling robots to efficiently perform collaborative operations in complex and dynamic environments. Especially when communication is interrupted or the task environment is unknown, they can still perform effective swarm behavior planning based on local perception information.

[0006] Specifically, the present invention proposes the design of a bionic vision projection field. This projection field collects environmental information in real time through the robot's own vision perception system (such as cameras, lidar, etc.), and adjusts the relative positions of robots within the swarm according to this information. Different from traditional global path planning methods, the system of the present invention emphasizes dynamic path planning based on local perception. Each robot calculates the positions of its surrounding environment and teammates in real time, and adjusts its movement trajectory to adapt to the constantly changing environment.

[0007] One of the core technologies of the present invention is to optimize the path planning and cooperative control of a robot swarm through a potential field enhancement mechanism. Specifically, robots adjust their movement directions by calculating the combined resultant force of gravitational and repulsive forces based on the positions of obstacles and prey in the surrounding environment. In particular, the movement of robots not only depends on the attraction of prey but also takes into account the repulsive force from obstacles and the relative positions of teammates, thus ensuring that the robots in the swarm can cooperate in an orderly manner, avoid collisions and effectively pursue the target.

[0008] To further enhance the autonomy and flexibility of the swarm, the present invention proposes an identity trade-off strategy. Under this strategy, each robot dynamically adjusts its role in the swarm (such as leader, follower or collaborator) according to task requirements, current state and environmental information. For example, during the critical stage of a task, the robot responsible for pursuit is given the identity of a leader, while other robots choose to follow or cooperate according to task requirements. During the execution of the task, if a robot is in a dangerous area or after the task is completed, its identity can be adjusted in real time, thus ensuring the continuity and coordination of the swarm task.

[0009] The present invention also proposes a dynamic path adjustment method based on local information, which solves the problem that traditional global path planning methods cannot cope with dynamic environmental changes. Each robot adjusts its path in real time according to its local perception information (such as the positions of surrounding obstacles and other robots), and ensures the coherence and stability of the overall movement of the swarm through cooperative control algorithms. Especially when facing obstacles (such as complex terrains like ruins and forests), robots can flexibly adjust their paths while ensuring the overall task goal, avoiding swarm splitting or task failure.

[0010] Path planning process: Local perception: Each robot obtains real-time information about the surrounding environment through a visual perception system, including information about prey, obstacles, teammates, etc., and makes local decisions based on this information.

[0011] Dynamic adjustment: During flight or movement, the swarm dynamically adjusts its path according to environmental changes (such as changes in obstacle positions or communication interruptions). Each robot ensures the consistency and stability of the overall path of the swarm through information exchange and control rules.

[0012] Crossing obstacles: When the swarm encounters complex obstacles (such as ruins, forests, etc.), robots adjust their movement trajectories based on local perception information while maintaining the overall movement state of the swarm to avoid task failure.

[0013] Key technological innovations: Dynamic path adjustment based on local perception: Through real-time local information acquisition and decision-making, robots can automatically adjust their paths according to the changing environment, avoiding over-reliance on global path planning.

[0014] Identity weighing strategy: Dynamically adjust the roles and task priorities of robots in the cluster, enabling the robot cluster to flexibly adjust task allocation and coordination methods in the event of emergencies or changes in task requirements.

[0015] Potential field enhancement mechanism: Through the potential field method introduced by bionics, robots can work together naturally and overcome various challenges in the environment, such as obstacle avoidance and cluster collaboration.

[0016] The present invention also provides a multi-robot cluster pursuit system based on potential field enhanced reinforcement learning, including multiple robot platforms, a visual perception module, an identity weighing decision module, a local path planning module, and a cluster collaborative control module. Each module ensures the efficient collaboration and autonomous task execution of the cluster in a complex environment through local information interaction and bionic algorithms.

[0017] Functions of each module: Visual perception module: Each robot is equipped with a visual perception system to collect real-time information about the surrounding environment, such as the positions of obstacles, prey, and teammates.

[0018] Identity weighing decision module: Dynamically adjust the role of each robot (such as leader, follower, or collaborator) according to task priorities, positions, environmental information, and interaction feedback.

[0019] Local path planning module: Combine local perception information, potential field enhancement algorithm, and task requirements to plan the motion trajectory of each robot in real time.

[0020] Cluster collaborative control module: Ensure the coordination and consistency among robots through simplified cluster collaborative control rules, and guarantee the stability and safety of cluster motion.

[0021] Environmental perception and information collection: Each robot obtains real-time information about its surrounding environment through the visual perception module, including the positions of prey, obstacles, and teammates.

[0022] Identity weighing and role assignment: According to task priorities, environmental information, and interactions among robots, the identity weighing module assigns appropriate roles to each robot and dynamically adjusts the task allocation of robots in the cluster.

[0023] Path planning and collaborative control: Each robot plans its path based on local perception information and cluster control rules. The robot will exchange information with neighboring robots to ensure the consistency and stability of the overall path of the cluster.

[0024] Cluster crossing and task execution: During task execution, the robot updates its position and state according to the bionic visual projection field, and dynamically adjusts the path when crossing a complex environment to ensure the efficient completion of tasks by the cluster.

[0025] The beneficial effects of the present invention are as follows: By combining the potential field enhancement mechanism and the identity weighing strategy, the present invention provides a multi-robot cluster hunting method with strong adaptability and high robustness in a dynamic environment. This method can efficiently cooperate to execute tasks in unknown, complex or communication-limited environments, breaking through the limitations of traditional path planning and cluster collaboration models, and has broad application prospects. Brief Description of the Drawings

[0026] Figure 1 is the flowchart of the multi-robot cluster hunting method based on potential field enhanced reinforcement learning

[0027] Figure 2 is the trend chart of the attractive force amplitude changing with distance

[0028] Figure 3 is the trend chart of the repulsive force amplitude changing with distance

[0029] Figure 4 is the trend chart of the amplitude of the interaction force between individuals changing with distance

[0030] Figure 5 is the schematic diagram of the problem that the temporary obstacle is unreachable in the artificial potential field method

[0031] Figure 6 is different value's influence on the cluster behavior

[0032] Figure 7 is the local minimum problem caused by the pursuer being repelled by the obstacle

[0033] Figure 8 is the "wall-following" behavior rule.

[0034] Figure 9 is the local minimum problem caused by the pursuer being repelled by its teammates

[0035] Figure 10 is the overall process of the multi-robot cluster hunting method based on potential field enhanced reinforcement learning

[0036] Figure 11 is the training environment of the multi-robot cluster hunting method based on potential field enhanced reinforcement learning

[0037] Figure 12 is the learning curve of the multi-robot cluster hunting and control method based on potential field enhanced reinforcement learning

[0038] Figure 13 is the average time used for each hunting of the multi-robot cluster hunting and control method based on potential field enhanced reinforcement learning

[0039] Figure 14The number of samples collected by each method per round

[0040] Figure 15 The verification scenario for verifying the generalization ability of each method to different obstacles Detailed implementation manners

[0041] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the purpose of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0042] Phase One: Observation and Perception

[0043] Each robot uses its sensors to perceive the current environment and collect the following information: Prey position: The relative position between the robot and the prey.

[0044] Obstacle position: The position of the obstacle closest to the robot.

[0045] Teammate position: The position of the teammates that the current robot can perceive.

[0046] Based on this data, the robot can make a basic judgment on the environment.

[0047] Phase Two: Embedding of Observation Information

[0048] To solve the problem of the variable length of the observation information between robots, an average embedding method is adopted to standardize the input. The position information of each teammate is processed by an embedding neural network to generate a fixed-length vector, and then the average value of all teammate information is taken to form a unified input vector, avoiding the problem of inconsistent input lengths caused by the change in the number of robots.

[0049] Phase Three: Calculation of Reinforcement Learning Model

[0050] The processed observation information is input into the reinforcement learning policy network, and the network outputs two key parameters (η, λ): η (Attraction): Determines the strength of the robot's movement towards the prey. The stronger the attraction, the closer the robot gets to the prey.

[0051] λ (Repulsion): Determines the strength of the robot's avoidance of obstacles. The stronger the repulsion, the stronger the robot's ability to avoid obstacles.

[0052] These parameters determine the robot's next action decision.

[0053] Phase Four: Action Execution and Collaboration

[0054] Based on the potential field parameters output by reinforcement learning, the robot will calculate and execute corresponding motion actions. The robot will move in the direction of the resultant force, while considering the relative position with its teammates, maintaining a reasonable formation, and ensuring the coordination of the entire cluster.

[0055] When encountering an obstacle, if the directions of the attractive force and the repulsive force are opposite, the robot will enter the "wall-following" mode to avoid being stagnant due to the repulsive force of the obstacle. In this mode, the robot will adjust its own direction, bypass the obstacle, and continue to pursue the prey.

[0056] Phase Five: Collaboration and Strategy Adjustment

[0057] During the process of task execution, the robots will learn from each other through a shared experience pool and update their strategies based on the feedback after each task execution (such as collisions, successful encirclement, etc.). Each robot dynamically adjusts its own strategy according to the local information it perceives and the task requirements, ensuring the collaboration and efficiency of the entire cluster.

[0058] Phase Six: Task Completion and Round Termination

[0059] When the robot successfully encircles the prey or completes the specified task, the task ends. According to the actual environmental conditions, the system can dynamically adjust the termination conditions of the task to ensure the successful execution of the task.

[0060] In the simulation environment, a scenario including prey, obstacles, and multiple robots (pursuers) is designed. The speed difference between each pursuer and the prey is set reasonably, and the obstacles are randomly arranged to increase the environmental complexity. Through experiments, it is verified that the multi-robot cluster encirclement method based on potential field enhanced reinforcement learning can effectively improve the learning efficiency and encirclement success rate of robots in complex dynamic environments.

[0061] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0062] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0063] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A multi-robot swarm hunting method based on potential field enhanced reinforcement learning, characterized in that, Comprising: Construct an average-embedded observation space module for converting variable-length observation information into fixed-length network inputs; Design a combined module of a reinforcement learning policy network and the artificial potential field method, taking the artificial potential field method as an expert policy and incorporating the parameter space of the artificial potential field method into the action space of reinforcement learning; Formulate a "wall-following" rule module for solving the local minimum problem of the artificial potential field method in path planning; Run a method process module based on a specific reinforcement learning paradigm, adopting an independent reinforcement learning paradigm and a weight sharing method, training and making decisions based on D3QN, processing empirical data, and updating network weights.

2. The multi-robot swarm hunting method based on potential field enhanced reinforcement learning according to claim 1, characterized in that, The constructed average-embedded observation space module includes: dividing the agent's observation information into prey and obstacle information with fixed vector lengths and observable teammate information with variable lengths; using a single-layer fully connected neural network to perform high-dimensional embedding on the information of each teammate, then averaging the high-dimensional embedding vectors of all teammates to obtain a representation vector containing all teammate information, and splicing it with the prey and obstacle information to form an environmental feature vector that can be used as the input of the reinforcement learning policy network.

3. The multi-robot cluster hunting method based on potential field enhanced reinforcement learning according to claim 1, characterized in that The designed combined module of the reinforcement learning policy network and the artificial potential field method includes: defining the attraction of the prey to the pursuer, the repulsion of the obstacle to the pursuer, and the inter-individual force of the teammate on the pursuer, and taking the vector sum and direction of these forces as the expected heading of the pursuer; reinforcement learning outputs a scalar for updating the relevant parameters in the artificial potential field method based on the local observation information of the pursuer, and substitutes the updated parameters into the artificial potential field method to calculate the resultant force.

4. The multi-robot swarm hunting method based on potential field enhanced reinforcement learning according to claim 1, wherein, The formulated "wall-following" rule module includes: when the included angle between the resultant force of the attraction and repulsion calculated by the pursuer exceeds 90°, switch to the "wall-following" mode, and select the orthogonal direction of the repulsion as the expected heading according to whether there are teammates and the relevant direction included angle situation; if a teammate enters the prey neighborhood, treat this teammate as a virtual obstacle.

5. The multi-robot swarm hunting method based on potential field enhanced reinforcement learning according to claim 1, characterized in that The run method process module based on a specific reinforcement learning paradigm includes: adopting D3QN as the underlying reinforcement learning method to discretize the parameter space of the artificial potential field method; in each episode, the pursuer receives initial local observation information, performs average embedding on the observation information, selects a pair of parameters of the artificial potential field method according to the Q-network, calculates various forces, determines the expected heading to move, and stores empirical data; after each episode ends, randomly sample historical experiences from the experience pool and update the weights of the Q-network according to the learning rules of D3QN.

6. A multi-robot swarm hunting system based on potential field enhanced reinforcement learning, characterized in that, Comprising: An average-embedded observation space unit for performing the operations of the constructed average-embedded observation space module as described in claim 2; A combined unit of reinforcement learning and artificial potential field for performing the operations of the designed combined module of the reinforcement learning policy network and the artificial potential field method as described in claim 3; A "wall-following" rule execution unit for performing the operations of the formulated "wall-following" rule module as described in claim 4; A method process control unit for performing the operations of the run method process module based on a specific reinforcement learning paradigm as described in claim 5.

7. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the multi-robot swarm hunting method based on potential field enhanced reinforcement learning according to any one of claims 1 - 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the multi-robot swarm hunting method based on potential field enhanced reinforcement learning according to any one of claims 1 - 5.

Citation Information

Cited By

  • Foot type robot cluster surrounding control method based on interaction force matrix

    CN121300473A

  • A legged robot cluster hunting control method based on interaction matrix

    CN121300473B