Multi-robot cooperative search and obstacle avoidance method and system based on deep reinforcement learning
Patent Information
- Application Number
- CN202611139649.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-30
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-30
AI Technical Summary
(1)缺乏对动态障碍物未来状态的预测模块,策略表现仍属于被动响应;
本公开的基于深度强化学习的多机器人协同搜索避障方法,通过设计由定位、环境感知、通信和任务信息组成的全维状态空间,构建包括预测和覆盖的多目标奖励函数,并构建集障碍物轨迹预测、去中心化图注意力通信与参数共享决策网络为一体的核心决策端,在动态复杂场景下实现多机器人分散协同搜索、主动障碍预测及安全避障,从而显著提升搜索效率与避障成功率,以解决传统方法中被动应对、通信僵化和异构训练困难的问题。
Smart Images

Figure CN122653226B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of robot intelligent control technology, specifically to a multi-robot cooperative search and obstacle avoidance method and system based on deep reinforcement learning. Background Technology
[0002] The statements in this section are merely background information relating to this disclosure and do not necessarily constitute prior art.
[0003] Currently, robotics technology has been widely applied in fields such as tunnel drilling and blasting, emergency search and rescue, industrial inspection, and unmanned reconnaissance. However, basic single-robot operations are increasingly unable to meet engineering needs. Therefore, the deep development and evolution of robots from single-robot operations to multi-robot collaboration is an inevitable trend, and the challenges faced by multi-robot systems in dynamic unstructured environments are becoming increasingly prominent.
[0004] Traditional control frameworks based on reactive obstacle avoidance and path planners based on search rely excessively on sensor data for instantaneous responses and lack the ability to actively reason about the movement trends of dynamic obstacles. They have a high failure rate in obstacle avoidance scenarios with dense and high-speed obstacles. Some studies have introduced methods such as Kalman filtering for short-term prediction, but due to their heavy reliance on accurate prior kinematic models, they are difficult to generalize to the complex interactive environments in reagent engineering. In terms of communication architecture, existing multi-machine systems mostly adopt centralized communication or fixed adjacency topology, which cannot dynamically adjust communication objects and information granularity in real time according to environmental changes, and will also cause information redundancy or missing key information.
[0005] Existing deep reinforcement learning control schemes achieve autonomous decision-making and real-time control in dynamic environments by introducing maximum entropy frameworks, lightweight networks, and security constraint modules. However, the following technical limitations still exist: (1) Lacking a module for predicting the future state of dynamic obstacles, the strategy performance remains a passive response; (2) Without a decentralized dynamic communication mechanism, it is difficult for robots to coordinate efficiently; (3) Heterogeneous robots with different motion models and sensor types cannot achieve parameter sharing and unified control through the same policy network. Summary of the Invention
[0006] To address the aforementioned issues, this disclosure proposes a multi-robot cooperative search and obstacle avoidance method and system based on deep reinforcement learning. This method extends deep reinforcement learning robot control technology from single-robot to multi-robot cooperative scenarios, introducing active prediction, decentralized dynamic communication, and heterogeneous parameter sharing mechanisms. This significantly improves the decision-making ability and cooperative search efficiency of multiple robots in complex dynamic environments.
[0007] According to some embodiments, the present disclosure adopts the following technical solutions: Multi-robot cooperative search and obstacle avoidance methods based on deep reinforcement learning include: Constructing a robot kinematics model for a multi-robot cooperative search scenario; Define a multi-objective dynamic reward function that includes prediction accuracy reward and collaborative coverage reward. With the goal of maximizing the sum of the cumulative reward of the strategy and the entropy regularization term, construct a multi-robot collaborative search and obstacle avoidance decision model that includes an obstacle trajectory prediction module, a decentralized graph attention communication module, and a parameter sharing decision module. The obstacle trajectory prediction module generates the future obstacle trajectory distribution online based on the VAE-LSTM world model. The decentralized graph attention communication module performs graph attention calculation and exchanges compressed features. The parameter sharing decision module adopts a unified lightweight Actor-Critic network to output the original action. The multi-robot collaborative search and obstacle avoidance decision model is based on the real-time point cloud computing safety speed limit to obtain the final execution instructions to realize multi-robot collaborative search and obstacle avoidance.
[0008] Furthermore, the construction of the robot kinematics model in the multi-robot cooperative search scenario includes: In a multi-robot collaborative search scenario, the state space and continuous motion space of each robot are predefined; the state space of each robot is composed of four parts: localization state information, environmental perception state information, communication state information, and task state information; the continuous motion space is defined as the continuous three-dimensional movement velocity components in the robot's base coordinate system. The localization status information includes three-dimensional position coordinates and three-axis orientation; the environmental perception status information includes local obstacle distance information and dynamic obstacle relative coordinates and velocity; the communication status information is the attention-weighted feature vector from other robots; and the task status is the absolute coordinates of the target point.
[0009] Furthermore, the multi-objective dynamic reward function is composed of a weighted sum of multiple sub-reward functions, including a target proximity reward function, a distance reward function, a collision penalty function, an energy consumption reward function, a motion smoothness reward function, a prediction accuracy reward function, and a cooperative coverage reward function. Among them, the reward term of the prediction accuracy reward function is the negative log-likelihood of the predicted future obstacle location distribution and the actual observation; the cooperative coverage reward function is constructed based on the current spatial distribution dispersion of all robots or the area of unexplored regions, which encourages decentralized search.
[0010] Furthermore, the obstacle trajectory prediction module includes a world model based on a variational autoencoder; the encoder of the world model based on the variational autoencoder uses an LSTM network to process the local observation sequence of the past K time steps, mapping it to a Gaussian distribution of latent variables; the decoder samples from this distribution and outputs the mean and variance of the dynamic relative positions of obstacles in the next T time steps.
[0011] Furthermore, in the decentralized graph attention communication module, each robot linearly projects its local observation features through a learnable query, key, and value matrix; for each neighboring robot within the communication range, the attention weight is calculated using scaled dot product attention, and the weighted sum of the value vectors is used as the communication feature vector; finally, the communication feature vector is concatenated with the local feature vector and then input into the decision network.
[0012] Furthermore, the parameter-sharing decision module adopts a unified lightweight Actor-Critic network. The Actor network uses parallel parameters and includes a hierarchical feature extraction and fully connected layer structure. Different types of information in the input state vector, including state information, localization state information, task state information, and communication state information, are directly input into the fully connected layer. The local obstacle distance information in the environmental perception state information is extracted through a one-dimensional depthwise separable convolution module, and the dynamic obstacle information is input into the fully connected layer. The features extracted from each path are concatenated and then fused through the fully connected layer. Finally, the mean and logarithmic standard deviation of the Gaussian policy are output through the output layer to generate a unified original action command and output it.
[0013] Furthermore, the multi-robot collaborative search and obstacle avoidance decision-making model, based on the real-time point cloud computing security speed limit, obtains the final execution instructions to realize multi-robot collaborative search and obstacle avoidance, including: The original actions of the Actor network are output through the differentiable safety constraint module and are limited to the safety boundary calculated based on the current perception data. Specifically, based on the real-time LiDAR point cloud state, the distance to the nearest obstacle in each direction of motion is calculated, the maximum allowable speed in that direction is defined, and finally, the components of the original actions are clipped to the set range through the clipping function to obtain the safe action command and execute it.
[0014] According to some embodiments, the present disclosure adopts the following technical solutions: A multi-robot cooperative search and obstacle avoidance system based on deep reinforcement learning includes: Predefined modules are used to build robot kinematic models for multi-robot cooperative search scenarios; The decision model construction module is used to define a multi-objective dynamic reward function that includes prediction accuracy reward and collaborative coverage reward. With the goal of maximizing the sum of the cumulative reward of the strategy and the entropy regularization term, it constructs a multi-robot collaborative search and obstacle avoidance decision model that includes an obstacle trajectory prediction module, a decentralized graph attention communication module, and a parameter sharing decision module. The obstacle trajectory prediction module generates the future obstacle trajectory distribution online based on the VAE-LSTM world model. The decentralized graph attention communication module performs graph attention calculation and exchanges compressed features. The parameter sharing decision module adopts a unified lightweight Actor-Critic network to output the original action. The multi-robot collaborative search and obstacle avoidance decision model is based on the real-time point cloud computing safety speed limit to obtain the final execution instructions to realize multi-robot collaborative search and obstacle avoidance.
[0015] According to some embodiments, the present disclosure adopts the following technical solutions: A non-transitory computer-readable storage medium is provided for storing computer instructions, which, when executed by a processor, implement the aforementioned multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning.
[0016] According to some embodiments, the present disclosure adopts the following technical solutions: An electronic device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning.
[0017] Compared with the prior art, the beneficial effects of this disclosure are as follows: This disclosed multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning designs a full-dimensional state space composed of localization, environmental perception, communication, and task information. It constructs a multi-objective reward function including prediction and coverage, and builds a core decision-making end that integrates obstacle trajectory prediction, decentralized graph attention communication, and parameter sharing decision network. In dynamic and complex scenarios, it realizes multi-robot distributed cooperative search, active obstacle prediction, and safe obstacle avoidance, thereby significantly improving search efficiency and obstacle avoidance success rate, and solving the problems of passive response, rigid communication, and heterogeneous training difficulties in traditional methods.
[0018] This disclosed multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning introduces an obstacle trajectory prediction module based on variational autoencoder, enabling the robot to predict the future obstacle location distribution based on historical observation data, thus eliminating the dependence on precise kinematic models. At the same time, the prediction accuracy is incorporated into the reward function and guided learning, enabling the robot to plan ahead in high-speed dynamic obstacle scenarios and effectively reduce the risk of collision.
[0019] The multi-robot cooperative search and obstacle avoidance method disclosed herein, based on deep reinforcement learning, adopts a decentralized graph attention communication mechanism, which realizes dynamic adaptation of communication topology and significant reduction of communication bandwidth. It overcomes the single point of failure problem, bandwidth bottleneck significantly affected by the number of robots, and information redundancy problem of fixed topology in centralized schemes, enabling multi-robot systems to maintain efficient cooperation even when deployed on a large scale.
[0020] This disclosure presents a multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning. In the deep reinforcement learning framework, a heterogeneous parameter sharing and differentiable safety constraint module is introduced: the Actor network adopts a parallel hierarchical feature extraction structure and processes the LiDAR sequence with depthwise separable convolution, which can be adapted to embedded computing power; the safety constraint module calculates the upper limit of safe speed based on real-time point cloud dynamics to ensure that the output command meets the robot control safety boundary; the parameter sharing mechanism design means that robots with different kinematic configurations do not need to be trained separately, which greatly reduces training cost and deployment difficulty. Attached Figure Description
[0021] The accompanying drawings, which form part of this disclosure, are used to provide a further understanding of this disclosure. The illustrative embodiments of this disclosure and their descriptions are used to explain this disclosure and do not constitute an undue limitation of this disclosure.
[0022] Figure 1 This is a diagram illustrating the overall architecture of the multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning, as described in this embodiment. Figure 2 This is a schematic diagram of the deep reinforcement learning training process and the deployment of real robots according to an embodiment of this disclosure; the upper part of the diagram shows the multi-robot centralized training and experience playback mechanism in the simulation environment, and the lower part shows the deployment process of embedded real-time inference on real robots after model transfer. Detailed Implementation
[0023] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.
[0024] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0025] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this disclosure. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0026] Example 1 One embodiment of this disclosure provides a multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning, the method steps of which include: Step 1: Construct a robot kinematics model for a multi-robot cooperative search scenario; Step 2: Define a multi-objective dynamic reward function that includes prediction accuracy reward and collaborative coverage reward. With the goal of maximizing the sum of the cumulative reward of the strategy and the entropy regularization term, construct a multi-robot collaborative search and obstacle avoidance decision-making model that includes an obstacle trajectory prediction module, a decentralized graph attention communication module, and a parameter sharing decision module. The obstacle trajectory prediction module generates the future obstacle trajectory distribution online based on the VAE-LSTM world model. The decentralized graph attention communication module performs graph attention calculation and exchanges compressed features. The parameter sharing decision module adopts a unified lightweight Actor-Critic network to output the original action. The multi-robot collaborative search and obstacle avoidance decision model is based on the real-time point cloud computing safety speed limit to obtain the final execution instructions to realize multi-robot collaborative search and obstacle avoidance.
[0027] As one embodiment, the multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning disclosed herein constructs a maximum entropy deep reinforcement learning framework that integrates active prediction, dynamic communication, and heterogeneous parameter sharing. It also inherits lightweight networks and differentiable safety constraint design to achieve cooperative search and autonomous obstacle avoidance by multiple robots in dynamic and complex environments, while possessing embedded real-time operation capabilities. The specific implementation process is as follows: Step 1: Construct a robot kinematics model for a multi-robot cooperative search scenario; Specifically, the target search environment is a typical dynamic unstructured scenario, widely found in engineering applications such as tunnel construction, post-disaster ruins, and industrial plants. Static obstacles are entities with fixed positions in the scene, including but not limited to tunnel sidewalls and faces, support structures, ventilation ducts, cable trays, stockpiled materials, collapsed walls, abandoned equipment, and building structural columns. Dynamic obstacles are entities whose position or velocity changes over time, including but not limited to mobile construction vehicles, walking workers, other collaborative robots, sudden rockfalls, collapsed debris, and mobile construction machinery. Each robot's control objective is to start from its own starting point, collaboratively traverse or search to the target area, avoiding collisions with static and dynamic obstacles and preventing mutual interception. Each robot is equipped with a 16-line LiDAR, IMU, wheeled odometer, and wireless communication module.
[0028] In a predefined multi-robot cooperative search scenario, each robot's state space and continuous motion space are defined. Each robot's state space is composed of four parts: localization state information, environmental perception state information, communication state information, and task state information. The continuous motion space is defined as the continuous three-dimensional movement velocity components in the robot's base coordinate system. Specifically, the localization state information includes three-dimensional position coordinates and three-axis orientation; the environmental perception state information includes local obstacle distance information and dynamic obstacle relative coordinates and velocities; the communication state information is the attention-weighted feature vector from other robots; and the task state is the absolute coordinates of the target point.
[0029] Furthermore, the specific definition of the state space consists of four parts, represented as follows: Among them, the positioning status Including three-dimensional position coordinates and three-axis orientation, forming a 6-dimensional vector. .in, This represents the horizontal and vertical position coordinates of the robot along the x-axis in the world coordinate system. Indicates along y Horizontal lateral position coordinates along the axis. Indicates along z The vertical height coordinate along the axis, along with the other two coordinates, together determine the robot's three-dimensional spatial position, all in meters. Indicates the robot circles x The roll angle of the axis of rotation, Indicates circling y The pitch angle of the axis of rotation. Indicates circling z The yaw angle of the axis rotation, along with the other two axes, defines the robot's three-axis spatial orientation in Euler angles, all in radians. Environmental perception state. It includes local obstacle distance information and dynamic obstacle relative information. Local obstacle distance information is obtained by dividing the 180° fan-shaped area in front into... For each sector, the shortest distance of the laser point cloud within each sector is calculated to form a 36-dimensional vector; the relative information of dynamic obstacles is the nearest detected obstacle. The relative coordinates and relative velocities of each dynamic obstacle are represented in 12 dimensions. Total 48 dimensions. Communication status. The attention-weighted feature vectors from other robots are 32-dimensional and generated by the communication module. Task state. Let be the absolute coordinates of the target point, which are 3D vectors. The final concatenation yields a state with the following dimension: Dimension. Action space is defined as continuous three-dimensional movement speed commands. Including linear velocity and angular velocity .
[0030] Step 2: Define a multi-objective dynamic reward function that includes prediction accuracy reward and collaborative coverage reward; Specifically, the multi-objective dynamic reward function is composed of a weighted sum of multiple sub-reward functions, including: target proximity reward function, distance reward function, collision penalty function, energy consumption reward function, motion smoothness reward function, prediction accuracy reward function, and cooperative coverage reward function. The multi-objective dynamic reward function adopts a multi-objective weighted summation form, as follows:
[0031] in, A one-time high reward upon reaching the target area; A reward for each step that shortens the target distance; Penalties based on relative velocity when colliding with obstacles or other robots; Set a small penalty for each step to encourage efficient paths; Path smoothness is assessed by the degree of jitter in linear velocity and angular velocity; The collaborative coverage reward function is constructed based on the current spatial distribution dispersion of all robots or the area of unexplored regions. It adopts the negative absolute value of the deviation between the average nearest neighbor distance and the ideal distance or is designed based on the regional search coverage to encourage decentralized search. The dynamic weight coefficients corresponding to each sub-reward function are used to adjust the contribution ratio of different optimization objectives to the total reward: The weight of the reward for achieving the goal is determined by controlling the intensity of the one-time positive incentive the robot receives when it reaches the target area, thus guiding the strategy to form a clear goal orientation. The weight of the distance reward is adjusted to control the positive feedback amplitude of each step of movement that shortens the target distance, so that the robot can continuously approach the target during the movement. The weight of the collision penalty controls the intensity of the negative penalty when colliding with obstacles or other robots, guiding the strategy to actively avoid dangerous areas; The weight of the energy consumption reward is adjusted to regulate the fixed energy consumption penalty applied for each step, encouraging the strategy to complete the task with the shortest path and the fewest steps, avoiding redundant wandering; The weight of motion smoothness reward is adjusted to improve the contribution of path smoothness evaluation to the total reward. By penalizing drastic changes in motion between adjacent time steps, the robot's motion is made smoother and more fluid. To weight the reward for prediction accuracy, the influence of the prediction accuracy of the obstacle trajectory prediction module on the overall reward is controlled, thereby guiding the world model to improve its ability to infer the future location distribution of obstacles. To weight the collaborative coverage reward, the contribution of the spatial distribution dispersion of multiple robots or the area of unexplored regions to the total reward is adjusted, encouraging robots to conduct collaborative searches within the search area. The aforementioned weight coefficients can be dynamically adjusted according to the task stage and environmental complexity, providing flexible feedback guidance for policy learning.
[0032] further, The reward for prediction accuracy is the negative log-likelihood of the predicted future obstacle location distribution with respect to the actual observations, defined as:
[0033] in, Indicates from time At that time The actual observed sequence of dynamic obstacle positions; For encoders from the past Latent variables extracted from local observation sequences at each time step. This reward encourages the obstacle trajectory prediction module to output a predicted distribution that matches the actual observations; the more accurate the prediction, the higher the reward value, thereby guiding the world model to continuously optimize its ability to infer future obstacle movement trends.
[0034] Step 3: Aiming at maximizing the sum of the cumulative reward and entropy regularization term of the strategy, a multi-robot collaborative search and obstacle avoidance decision-making model is constructed, comprising an obstacle trajectory prediction module, a decentralized graph attention communication module, and a parameter sharing decision-making module. The obstacle trajectory prediction module generates a probability distribution of future obstacle positions based on historical observation sequences, enabling proactive prediction. The decentralized graph attention communication module dynamically selects communication objects and exchanges compressed feature vectors. The parameter sharing decision-making module allows all heterogeneous robots to share the same Actor-Critic network. After the Actor network outputs the original actions, a differentiable safety constraint module trims them into safe speed commands. During deployment, each robot independently runs the trained lightweight network, acquiring safe action commands in real time based on local sensor data and neighborhood communication information, thus achieving multi-robot collaborative search and obstacle avoidance. Specific details are as follows: First, training is performed by maximizing an objective function, which is the sum of the standard expected cumulative reward and the policy's entropy regularization term, expressed as:
[0035] in, As a strategy, State-action access distribution under the strategy; This is a discount factor used to balance the importance of current rewards and future rewards; Indicates the policy in the state Entropy is used to quantify the randomness and exploratory ability of a strategy; This represents the entropy coefficient, used to balance the importance of reward maximization and policy entropy. This is the total reward function.
[0036] Furthermore, a decision model is constructed that includes an obstacle trajectory prediction module, a decentralized graph attention communication module, and a parameter sharing decision module. The goal is to maximize the sum of the cumulative reward of the policy and the entropy regularization term, and a centralized training-decentralized execution paradigm is adopted for training.
[0037] The obstacle trajectory prediction module includes a world model based on a variational autoencoder. The encoder uses a two-layer Long Short-Term Memory (LSTM) network, one of which stores past events. Local observation sequence of steps Mapping to latent variables Gaussian distribution parameters; decoder sampling Then, the future output is processed through another LSTM layer. The mean and variance of the relative positions of dynamic obstacles in each step. The training loss is a combination of prediction error and KL divergence:
[0038] in, As a latent variable, For encoder distribution, For decoder distribution, The prior distribution of latent variables is usually taken as the standard normal distribution. , The weighting coefficients for the KL divergence term; prior. Assuming a standard normal distribution , Control the regularization strength.
[0039] The output of the obstacle trajectory prediction module effectively expands the robot's state space, thereby enabling forward-looking predictions.
[0040] Furthermore, the decentralized graph attention communication module works as follows: each robot linearly projects its local observation features using a learnable query, key, and value matrix; for each neighboring robot within the communication range, it calculates attention weights using scaled dot product attention, and uses the weighted sum of the value vectors as the communication feature vector, specifically including: The decentralized graph attention communication module enables each agent to... First, local observation features Perform linear projection and calculate attention weights:
[0041] in, For robots Neighborhood robots The attention weights are calculated using a scaled dot product attention mechanism; and Robots and Local observation features, , For learnable query and key projection matrices, The dimension of the key vector is used to scale the dot product to prevent gradient vanishing; the denominator is the product of all neighboring robots. The exponential terms are summed to normalize the weights.
[0042] This weight is used to perform weighted aggregation of neighborhood value vectors, enabling the robot to dynamically focus on the neighbor information within the communication range that is most valuable for the current decision, thus achieving decentralized adaptive communication. For robots The set of communicating neighbors; and thus the robot Communication feature vector for:
[0043] in, , , For learnable projection matrices, The dimension is the key vector. Communication features. As state components After being concatenated with local features, the data is input into the subsequent decision network to achieve dynamic topology and only exchange compressed features (32 dimensions), resulting in a communication volume much smaller than that of the original observations.
[0044] Furthermore, the parameter-sharing decision module adopts a centralized training-distributed execution paradigm, with all heterogeneous robots sharing one Actor network and two Critic networks. The Actor network employs parallel parameters, including hierarchical feature extraction and a fully connected layer structure. Different types of information in the input state vector, including state information, localization state information, task state information, and communication state information, are directly input into the fully connected layer. Local obstacle distance information in the environmental perception state information is extracted using a one-dimensional depthwise separable convolution module, while dynamic obstacle information is input into the fully connected layer. The features extracted from each pathway are concatenated, then fused through the fully connected layer, and finally output through the output layer the mean and logarithmic standard deviation of the Gaussian policy to generate a unified original action command.
[0045] Specifically, the Actor network structure continues the lightweight design of single-robot path planning: the LiDAR range sequence is input into a one-dimensional depthwise separable convolution, and spatial features are extracted using a kernel of size 3 and 16 channels; localization state... Task status With communication status The input is directly fed into the fully connected layer; after the three features are concatenated, they are fused through the fully connected layer, and the output layer generates the action mean. and logarithmic standard deviation This allows for the sampling of normalized original actions. Each component in Interval.
[0046] At the Actor network output, the differentiable safety constraint module outputs the Actor network's original actions, confining them within a safety boundary calculated based on the current perception data. Specifically, it calculates the nearest obstacle distance for each motion direction based on the real-time LiDAR point cloud status. Then the maximum permissible speed in that direction .
[0047] Specifically, the safety limit for each direction of motion is calculated based on real-time laser point cloud computing. :
[0048] in, The distance to the nearest obstacle in the lidar point cloud along the current direction of motion. To establish a safe distance, This is the time constant. This formula indicates that when the obstacle distance is greater than the safe buffer distance, the permissible speed is proportional to the margin exceeding the safe distance; when the obstacle distance is less than or equal to the safe buffer distance, the permissible speed drops to zero, prohibiting the robot from continuing to move in that direction, thus ensuring operational safety from the source of the command.
[0049] The final safety action is given by the following formula:
[0050] in, The raw action instructions output by the Actor network. The safe speed limits in each direction are calculated based on the current sensing data. For the clipping function, restrict each component of the original motion to... Within the specified range. This operation directly embeds security constraints into the policy network output, ensuring that all executed instructions remain within the current perceived security boundary, thus mitigating collision risks from the decision-making source.
[0051] The Critic network employs a dual-network design, sharing an input layer and a first hidden layer. It then branches into two evaluation heads: a first online commenting network, a second online commenting network, and a corresponding target commenting network. The first and second online commenting networks use a parameter-sharing design, sharing the input layer and the first hidden layer. They then branch into two independent lightweight evaluation heads. These lightweight evaluation heads are primarily used to concatenate the shared backbone output features and all robot motion vectors, and then map them to... Q Value; output two The values are adjusted to mitigate overestimation errors. The target comment network has the same structure as the online comment network, and its parameters are synchronized through periodic soft updates based on the online comment network.
[0052] The target comment network parameters are updated softly on average via exponential shifting. Training employs a maximum entropy reinforcement learning framework, trained by maximizing the objective function.
[0053] As one implementation, the multi-robot cooperative search and obstacle avoidance decision-making model is trained using a course-based learning strategy. A completely static, simple environment is initially set up, and dynamic obstacles are gradually introduced, with their speed and number increasing. The experience replay pool capacity... Batch size 1024, discount factor Soft update coefficient The prediction module, communication module, and decision network are updated alternately until convergence.
[0054] Finally, the trained Actor network, prediction module, and communication module optimized for TensorRT are deployed to a robot embedded board such as the Nvidia Jetson AGX Orin. During runtime, the following loop is executed at a frequency of 10Hz: sensor data acquisition → stitching state → running the prediction module to obtain future obstacle distribution → running the communication module to obtain neighborhood compression features → Actor network forward inference to obtain... →Safety constraint pruning→Finally, the speed command is sent to the underlying controller. The entire execution process does not rely on a central node, enabling fully distributed collaborative search and active obstacle avoidance.
[0055] Example 2 One embodiment of this disclosure provides a multi-robot cooperative search and obstacle avoidance system based on deep reinforcement learning, including: Predefined modules are used to build robot kinematic models for multi-robot cooperative search scenarios; The decision model construction module is used to define a multi-objective dynamic reward function that includes prediction accuracy reward and collaborative coverage reward. With the goal of maximizing the sum of the cumulative reward of the strategy and the entropy regularization term, it constructs a multi-robot collaborative search and obstacle avoidance decision model that includes an obstacle trajectory prediction module, a decentralized graph attention communication module, and a parameter sharing decision module. The obstacle trajectory prediction module generates the future obstacle trajectory distribution online based on the VAE-LSTM world model. The decentralized graph attention communication module performs graph attention calculation and exchanges compressed features. The parameter sharing decision module adopts a unified lightweight Actor-Critic network to output the original action. The multi-robot collaborative search and obstacle avoidance decision model is based on the real-time point cloud computing safety speed limit to obtain the final execution instructions to realize multi-robot collaborative search and obstacle avoidance.
[0056] Example 3 One embodiment of this disclosure provides a non-transitory computer-readable storage medium for storing computer instructions. When these computer instructions are executed by a processor, they implement the multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning.
[0057] Example 4 One embodiment of this disclosure provides an electronic device, including: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning.
[0058] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0059] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0060] While the specific embodiments of this disclosure have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of this disclosure are still within the scope of protection of this disclosure.
Claims
1. A multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning, characterized in that, include: Constructing a robot kinematics model for a multi-robot cooperative search scenario; Define a multi-objective dynamic reward function that includes prediction accuracy reward and collaborative coverage reward. With the goal of maximizing the sum of the cumulative reward of the policy and the entropy regularization term, construct a multi-robot collaborative search and obstacle avoidance decision model that includes an obstacle trajectory prediction module, a decentralized graph attention communication module, and a parameter sharing decision module. The multi-objective dynamic reward function is composed of a weighted sum of multiple sub-reward functions, including a target proximity reward function, a distance reward function, a collision penalty function, an energy consumption reward function, a motion smoothness reward function, a prediction accuracy reward function, and a cooperative coverage reward function. Each weight coefficient can be dynamically adjusted according to the task stage and environmental complexity; Among them, the reward term of the prediction accuracy reward function is the negative log-likelihood of the predicted future obstacle location distribution and the actual observation; the cooperative coverage reward function is constructed based on the current spatial distribution dispersion of all robots or the area of unexplored regions, which encourages decentralized search. The obstacle trajectory prediction module generates the future obstacle trajectory distribution online based on the VAE-LSTM world model, the decentralized graph attention communication module performs graph attention calculation and exchanges compressed features, the parameter sharing decision module adopts a unified lightweight Actor-Critic network to output the original action, and the multi-robot collaborative search and obstacle avoidance decision model is based on the real-time point cloud computing safety speed limit to obtain the final execution instruction to realize multi-robot collaborative search and obstacle avoidance. The obstacle trajectory prediction module includes a world model based on a variational autoencoder; the encoder of the world model based on the variational autoencoder uses an LSTM network to process the local observation sequence of the past K time steps, which is mapped to a Gaussian distribution of latent variables; the decoder samples from this distribution and outputs the mean and variance of the dynamic relative positions of obstacles in the next T time steps. In the decentralized graph attention communication module, each robot linearly projects its local observation features through a learnable query, key, and value matrix. For each neighboring robot within the communication range, the attention weight is calculated using scaled dot product attention, and the weighted sum of the value vectors is used as the communication feature vector. Finally, the communication feature vector is concatenated with the local feature vector and then input into the decision network. The parameter-sharing decision module adopts a unified lightweight Actor-Critic network. The Actor network uses parallel parameters and includes a hierarchical feature extraction and fully connected layer structure. Different types of information in the input state vector, including state information, localization state information, task state information, and communication state information, are directly input into the fully connected layer. The local obstacle distance information in the environmental perception state information is extracted through a one-dimensional depthwise separable convolution module, and the dynamic obstacle information is input into the fully connected layer. The features extracted from each path are concatenated and then fused through the fully connected layer. Finally, the mean and logarithmic standard deviation of the Gaussian policy are output through the output layer to generate a unified original action command and output it. The multi-robot collaborative search and obstacle avoidance decision-making model is based on the real-time point cloud computing safety speed limit, and obtains the final execution instructions to realize multi-robot collaborative search and obstacle avoidance, including: The original actions of the Actor network are output through the differentiable safety constraint module and are limited to the safety boundary calculated based on the current perception data. Specifically, based on the real-time LiDAR point cloud state, the distance to the nearest obstacle in each direction of motion is calculated, the maximum allowable speed in that direction is defined, and finally, the components of the original actions are clipped to the set range through the clipping function to obtain the safe action command and execute it.
2. The multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning as described in claim 1, characterized in that, The construction of the robot kinematics model in the multi-robot cooperative search scenario includes: In a multi-robot collaborative search scenario, the state space and continuous motion space of each robot are predefined; the state space of each robot is composed of four parts: localization state information, environmental perception state information, communication state information, and task state information; the continuous motion space is defined as the continuous three-dimensional movement velocity components in the robot's base coordinate system. The localization status information includes three-dimensional position coordinates and three-axis orientation; the environmental perception status information includes local obstacle distance information and dynamic obstacle relative coordinates and velocity; the communication status information is the attention-weighted feature vector from other robots; and the task status is the absolute coordinates of the target point.
3. A multi-robot cooperative search and obstacle avoidance system based on deep reinforcement learning, characterized in that, include: Predefined modules are used to build robot kinematic models for multi-robot cooperative search scenarios; The decision model construction module is used to define a multi-objective dynamic reward function that includes prediction accuracy reward and collaborative coverage reward. With the goal of maximizing the sum of the cumulative reward of the policy and the entropy regularization term, it constructs a multi-robot collaborative search and obstacle avoidance decision model that includes an obstacle trajectory prediction module, a decentralized graph attention communication module, and a parameter sharing decision module. The multi-objective dynamic reward function is composed of a weighted sum of multiple sub-reward functions, including a target proximity reward function, a distance reward function, a collision penalty function, an energy consumption reward function, a motion smoothness reward function, a prediction accuracy reward function, and a cooperative coverage reward function. Each weight coefficient can be dynamically adjusted according to the task stage and environmental complexity; Among them, the reward term of the prediction accuracy reward function is the negative log-likelihood of the predicted future obstacle location distribution and the actual observation; the cooperative coverage reward function is constructed based on the current spatial distribution dispersion of all robots or the area of unexplored regions, which encourages decentralized search. The obstacle trajectory prediction module generates the future obstacle trajectory distribution online based on the VAE-LSTM world model, the decentralized graph attention communication module performs graph attention calculation and exchanges compressed features, the parameter sharing decision module adopts a unified lightweight Actor-Critic network to output the original action, and the multi-robot collaborative search and obstacle avoidance decision model is based on the real-time point cloud computing safety speed limit to obtain the final execution instruction to realize multi-robot collaborative search and obstacle avoidance. The obstacle trajectory prediction module includes a world model based on a variational autoencoder; the encoder of the world model based on the variational autoencoder uses an LSTM network to process the local observation sequence of the past K time steps, which is mapped to a Gaussian distribution of latent variables; the decoder samples from this distribution and outputs the mean and variance of the dynamic relative positions of obstacles in the next T time steps. In the decentralized graph attention communication module, each robot linearly projects its local observation features through a learnable query, key, and value matrix. For each neighboring robot within the communication range, the attention weight is calculated using scaled dot product attention, and the weighted sum of the value vectors is used as the communication feature vector. Finally, the communication feature vector is concatenated with the local feature vector and then input into the decision network. The parameter-sharing decision module adopts a unified lightweight Actor-Critic network. The Actor network uses parallel parameters and includes a hierarchical feature extraction and fully connected layer structure. Different types of information in the input state vector, including state information, localization state information, task state information, and communication state information, are directly input into the fully connected layer. The local obstacle distance information in the environmental perception state information is extracted through a one-dimensional depthwise separable convolution module, and the dynamic obstacle information is input into the fully connected layer. The features extracted from each path are concatenated and then fused through the fully connected layer. Finally, the mean and logarithmic standard deviation of the Gaussian policy are output through the output layer to generate a unified original action command and output it. The multi-robot collaborative search and obstacle avoidance decision-making model is based on the real-time point cloud computing safety speed limit, and obtains the final execution instructions to realize multi-robot collaborative search and obstacle avoidance, including: The original actions of the Actor network are output through the differentiable safety constraint module and are limited to the safety boundary calculated based on the current perception data. Specifically, based on the real-time LiDAR point cloud state, the distance to the nearest obstacle in each direction of motion is calculated, the maximum allowable speed in that direction is defined, and finally, the components of the original actions are clipped to the set range through the clipping function to obtain the safe action command and execute it.
4. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement the multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning as described in any one of claims 1-2.
5. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to perform the multi-robot cooperative search and obstacle avoidance method based on deep reinforcement learning as described in any one of claims 1-2.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle formation efficient cooperative control method
CN120428742A
Unmanned aerial vehicle collaborative search method based on multi-agent reinforcement learning
CN121325927A