Multi-vehicle homogeneous robot cluster self-organizing formation operation cooperative scheduling method and system

By optimizing task allocation and path planning for multi-robot isomorphic robot clusters through deep reinforcement learning and model predictive control, the problem of high computational and communication overhead in multi-robot systems is solved, achieving efficient and collaborative task scheduling and improving the efficiency and flexibility of logistics warehousing.

CN120949822BActive Publication Date: 2026-04-21UNIV OF SCI & TECH BEIJING
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UNIV OF SCI & TECH BEIJING
Filing Date
2025-10-20
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as high computational and communication overhead, local optima, large computational load and slow convergence in multi-robot task allocation and path planning, making it difficult to achieve efficient and collaborative operation of multi-robot isomorphic robot clusters.

Method used

Deep reinforcement learning algorithms are used for task matching, dynamic window algorithms and leader-follower methods are combined for path planning, and model predictive control is used to optimize robot speed, so as to realize the self-organized formation operation of multi-vehicle isomorphic robot clusters.

Benefits of technology

It improves the overall working efficiency and adaptability of multi-vehicle isomorphic robot clusters, and enhances the system's adaptability to complex environments and the safety and stability of transporting large and bulky goods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120949822B_ABST
    Figure CN120949822B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for self-organizing platooning and collaborative scheduling of multi-vehicle isomorphic robot swarms, belonging to the field of intelligent logistics. It includes: constructing a representation function of the spatiotemporal characteristics of cargo handling; robot scheduling based on position-task objective decision-making; self-organizing collaborative path planning for multi-vehicle isomorphic robots; and collaborative optimization of the platooning speed of multi-vehicle isomorphic robots. Employing the technical solution of this invention effectively improves the efficiency and flexibility of logistics warehousing, enhances the system's adaptability to complex environments, and improves the safety and stability of transporting large and bulky goods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent logistics technology, and in particular relates to a method and system for self-organizing formation and collaborative scheduling of multi-vehicle isomorphic robot clusters. Background Technology

[0002] With the significant growth in e-commerce transactions and the rapid increase in express delivery volume, traditional warehousing urgently needs to transform into intelligent warehousing, and automated logistics technology has also achieved rapid development. Warehouse picking, as a key link in the logistics process, directly impacts overall logistics efficiency and service quality. With technological advancements, the traditional "person-to-goods" picking model is gradually being replaced by a "goods-to-person" picking model. In multi-robot scheduling in logistics warehousing, multi-robot task allocation and multi-robot path planning occupy a core position in the multi-robot system, crucial for achieving efficient and collaborative operations. Current research on multi-robot collaborative scheduling shows that task allocation and path planning are essential for achieving efficient and collaborative operations. However, in terms of multi-robot task allocation, the market-based approach incurs high computational and communication overhead when tasks are complex and there are many robots; the behavioral approach, while dynamically adjustable, is prone to getting trapped in local optima; and the optimization approach is suitable for small-scale tasks but is sensitive to initial conditions and parameters. Regarding path planning, traditional algorithms have high computational and communication overhead in large-scale logistics environments, making real-time path planning difficult; biomimetic algorithms have advantages in handling complex problems, but suffer from high computational load and slow convergence in large-scale environments. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method and system for self-organizing formation and collaborative scheduling of multi-vehicle isomorphic robot clusters, which can effectively improve the overall working efficiency of multi-vehicle isomorphic mobile robot clusters and has strong adaptability to cope with complex and changing working scenarios.

[0004] To achieve the above objectives, the present invention adopts the following technical solution:

[0005] A method for self-organizing formation and collaborative scheduling of multi-vehicle isomorphic robot swarms includes:

[0006] Construct a representation function for the spatiotemporal characteristics of cargo handling;

[0007] The optimal decision-making strategy for task matching of swarm robots based on deep reinforcement learning is used to determine the target scheduling objects of swarm handling robots.

[0008] The selected cluster handling robot adopts The algorithm performs path planning in the self-organizing assembly phase and combines it with a dynamic window algorithm to achieve local obstacle avoidance. It also performs self-organizing formation path planning using a cluster-based leader-follower method.

[0009] The Model Predictive Control (MPC) method is used to adjust the speed and acceleration of each robot in real time, thereby achieving collaborative optimization of the speed of multi-robot isomorphic robot formation.

[0010] This invention also provides a self-organizing formation and collaborative scheduling system for multi-vehicle isomorphic robot clusters, comprising:

[0011] The first processing module is used to construct the representation function of the spatiotemporal characteristics of cargo handling;

[0012] The second processing module is used to determine the target scheduling object of the cluster robot task matching optimal decision strategy based on deep reinforcement learning.

[0013] The third processing module is used to select the cluster handling robot. The algorithm performs path planning in the self-organizing assembly phase and combines it with a dynamic window algorithm to achieve local obstacle avoidance. It also performs self-organizing formation path planning using a cluster-based leader-follower method.

[0014] The fourth processing module is used to adjust the speed and acceleration of each robot in real time using the Model Predictive Control (MPC) method, so as to achieve collaborative optimization of the speed of multi-robot isomorphic robot formation driving.

[0015] This invention uses a deep reinforcement learning algorithm to assign tasks to robots. The algorithm performs global path planning and combines it with a dynamic window algorithm to achieve local obstacle avoidance. By designing a leader-follower mechanism, the robot's collaborative formation speed is optimized. This invention effectively improves the efficiency and flexibility of logistics warehousing, enhances the system's adaptability to complex environments, and improves the safety and stability of transporting large and bulky goods. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0017] Figure 1 This is a flowchart of the self-organizing formation and collaborative scheduling method for multi-vehicle isomorphic robot clusters according to an embodiment of the present invention;

[0018] Figure 2 This is a spatial positional relationship diagram of robot formation driving provided in an embodiment of the present invention; Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] Example 1:

[0022] like Figure 1 As shown, this embodiment of the invention provides a method for self-organizing formation and collaborative scheduling of multi-vehicle isomorphic robot clusters, including:

[0023] Step S1: Construct a representation function for the spatiotemporal characteristics of cargo handling.

[0024] First, based on the robot's travel speed and the requirements of cargo attributes for movement speed, a cargo handling speed characteristic function is constructed based on the cargo movement speed and robot travel speed under the constraints of cargo transportation attributes. The specific form is as follows:

[0025] ;

[0026] in, The maximum speed allowed for transporting goods is determined based on the characteristics of the goods. For example, liquid goods are more sensitive to speed and have a maximum speed requirement. This is the maximum speed allowed for the robot to move. This is a speed proportionality coefficient, allowing for speed adjustment.

[0027] Using the distance traveled during transportation as a known quantity, a time characteristic function is constructed. The specific form of the time function during the robot's cargo transportation process is as follows:

[0028] ;

[0029] in, It is the distance between the robot and the destination of the goods. This is the robot's maximum speed.

[0030] Based on the speed and spatial characteristics of cargo movement, the representation function of the spatiotemporal characteristics of cargo handling is as follows:

[0031] ;

[0032] The spatiotemporal characteristic function of cargo handling represents the maximum speed that limits the travel speed of the handling robot during cargo handling, thereby providing constraints for the robot's state variables in selection and speed optimization.

[0033] Step S2: Robot scheduling based on position-task objective optimization decision

[0034] Construct a task matching strategy for cluster robots based on deep reinforcement learning algorithms to determine the target scheduling objects for cluster transport robots;

[0035] The construction of decision-making strategies for swarm robots based on deep reinforcement learning includes: state space construction, action space construction, and reward function construction.

[0036] state space Covering task information Robot status Environmental information (Map), State Space The format is as follows:

[0037] ;

[0038] Among them, for the task Task information Including cargo weight Cargo length Cargo width Cargo location information Task Information The format is as follows:

[0039] ;

[0040] Among them, for robots Robot status Including robot location Robot speed v ;

[0041] ;

[0042] The requirements for the robot's travel speed are as follows:

[0043] ;

[0044] Among them, environmental information This refers to map information of the storage area, represented as a set in the form of a two-dimensional matrix. Represents coordinates in the environment map The state of being, This indicates that the location is passable. This indicates that there is an obstacle at this location;

[0045] ;

[0046] The action space (A) selects whether each robot participates in the task (i.e., selects the robot set);

[0047] ;

[0048] The reward function (R) is used to guide the model to optimize the objective. For the first The time it takes for a robot to travel to the location of the goods. For the first The energy consumption of a robot traveling to the location of goods. For the first The cost of a robot traveling to the location of goods. , and The reward function, with the corresponding weighting coefficients, is as follows:

[0049] ;

[0050] According to the designed state space S Action space A and reward function R A task matching model for the robot was established using the DQN deep reinforcement learning algorithm. Ultimately, the selected robots for cluster handling were determined as follows:

[0051] ;

[0052] Step S3: Self-organizing collaborative path planning for multi-vehicle isomorphic robots

[0053] The self-organizing collaborative path planning of multi-vehicle isomorphic robots is divided into two stages: self-organizing assembly path planning and self-organizing formation path planning.

[0054] Self-organizing assembly phase path planning, adopting The algorithm generates an initial path for each robot from its starting position to a common target position based on a static environment map, and combines this with the Dynamic Window (DWA) method to adjust the robot's motion in real time to address the issue of robots relying solely on a static environment map. The algorithm addresses the dynamic obstacle problem, reducing the risk of collisions and thus ensuring the global optimality and local real-time performance of the path.

[0055] Self-organizing swarm path planning: After the robots reach the cargo location, they form a swarm to transport the cargo to the destination. A swarm-based leader-follower method is used for path planning. To ensure stable cargo transportation, the robot closest to the cargo's center of gravity is selected as the leader robot. A map of the robot swarm's operating environment is pre-built on the server, using information about other robots outside the swarm as dynamic obstacle information. When the leader robot receives the destination location information, it sends a path planning request (including the leader robot's position and the target destination location) to the server. The server performs global path planning, generating the optimal path based on the leader robot's path planning request and the Floyd algorithm, and sends it to the leader robot. The leader robot, based on the optimal path and its kinematic model, generates a path deviation from the optimal path based on environmental information, and then performs real-time deviation correction (distance and angle correction) on the actual path, performing local planning to determine the local route and desired speed. The following robots calculate their own travel routes and desired speeds based on the leader robot's planned local routes and desired speeds, ensuring swarm travel during the journey.

[0056] In particular, a dynamic window path planning method that considers the robot's dynamic performance is adopted in the local planning of formation control. When the distance to the obstacle perceived by the robot system is less than a threshold, When the time comes, switch to formation obstacle avoidance mode and execute dynamic window algorithm for obstacle avoidance; threshold Select settings:

[0057] ;

[0058] in, To determine the shortest distance that allows the robot to avoid obstacles and successfully complete its path planning. This refers to the maximum obstacle avoidance distance when the robot reaches its maximum speed. For the robot's driving speed, The maximum speed the robot can achieve, a constant. ;

[0059] The navigator robot collects information about the surrounding environment, including static environmental information (obstacles such as infrastructure) and dynamic environmental information (robots outside the moving formation). It uses the real-time status and future trends of the environment around the formation as predictive information and combines the running status of the formation to perform local dynamic environmental path planning. The follower robot follows the navigator robot by maintaining a certain ideal distance and relative angle.

[0060] like Figure 2 The location of the navigation robot shown is The position of the following robot is The expected relative distance between the following robot and the leading robot is The expected relative angle is The spatial relationship between the navigating robot and the following robot is as follows:

[0061] ;

[0062] Define the difference between the ideal relative distance and the actual relative distance as By introducing a feedback proportional coefficient By controlling the feedback proportional coefficient, the distance and angle between the following robot and the navigating robot are made to approach the ideal distance and angle. The specific form is as follows:

[0063] ;

[0064] Step S4: Cooperative optimization of the speed of multi-vehicle isomorphic robot formation.

[0065] During the platooning transportation phase, the Model Predictive Control (MPC) method is used to ensure that the robots can maintain a stable platooning formation for transporting goods. First, the state transition equations for robot platooning are constructed, then the overall constraints for platoon stability control are set, and finally, the cost function is set to solve for the driving speed to obtain the speed during stable driving.

[0066] The overall state transition equations for the entire robot formation are as follows:

[0067] ;

[0068] in, It's a robot. The state at any given moment, It's a robot. The state at any given moment, The state matrix, For the input matrix, , as follows:

[0069] ;

[0070] The overall constraints for queue stability control include acceleration constraints, velocity constraints, and formation constraints, as follows:

[0071] ;

[0072] Among them, the maximum speed in the speed constraint For steps Velocity calculated in the spatiotemporal characteristic function of cargo handling ;

[0073] With the optimization objectives of low time cost, stable driving, and maintaining formation, the cost function is defined as follows:

[0074] ;

[0075] in, Indicates the actual arrival time of the robot. For the target arrival time, For the acceleration of the robot's movement, As a weighting factor, the higher the requirement for stable transportation of goods, The larger.

[0076] Example 2:

[0077] This invention also provides a self-organizing formation and collaborative scheduling system for multi-vehicle isomorphic robot clusters, comprising:

[0078] The first processing module is used to construct the representation function of the spatiotemporal characteristics of cargo handling;

[0079] The second processing module is used to determine the target scheduling object of the cluster robot task matching optimal decision strategy based on deep reinforcement learning.

[0080] The third processing module is used to select the cluster handling robot. The algorithm performs path planning in the self-organizing assembly phase and combines it with a dynamic window algorithm to achieve local obstacle avoidance. It also performs self-organizing formation path planning using a cluster-based leader-follower method.

[0081] The fourth processing module is used to adjust the speed and acceleration of each robot in real time using the Model Predictive Control (MPC) method, so as to achieve collaborative optimization of the speed of multi-robot isomorphic robot formation driving.

[0082] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for self-organizing formation and collaborative scheduling of multi-vehicle isomorphic robot clusters, characterized in that, include: Construct a representation function for the spatiotemporal characteristics of cargo handling; Specifically: Based on the robot's travel speed and the requirements of cargo attributes for movement speed, a cargo handling speed characteristic function is constructed based on cargo movement speed and robot travel speed under cargo transportation attribute constraints. The specific form is as follows: in, The maximum speed permitted for transporting goods; This is the maximum speed allowed for the robot to move. This is the speed ratio coefficient; Using the distance traveled during transportation as a known quantity, a time characteristic function is constructed. The specific form of the time function during the robot's cargo transportation process is as follows: in, It is the distance between the robot and the destination of the goods. This is the robot's maximum speed. Based on the speed and spatial characteristics of cargo movement, the representation function of the spatiotemporal characteristics of cargo handling. A deep reinforcement learning-based task matching strategy for cluster robots, optimized based on location and task objectives, is constructed to determine the target scheduling objects for the cluster of transport robots. Specifically: The construction of task decision-making strategies for swarm robots based on deep reinforcement learning includes: state space construction, action space construction, and reward function construction. state space Covering task information Robot status Environmental information state space The format is as follows: Among them, for the task Task information Including cargo weight Cargo length Cargo width Cargo location information Task Information The format is as follows: Among them, for robots Robot status Including robot location Robot speed v ; The requirements for the robot's travel speed are as follows: Among them, environmental information This refers to map information of the storage area, represented as a set in the form of a two-dimensional matrix. Represents coordinates in the environment map The state of being, This indicates that the location is passable. This indicates that there is an obstacle at this location; Action space D For each robot, select whether it participates in the task; The reward function R is used to guide the model in optimizing its objective. For the first The time it takes for a robot to travel to the location of the goods. For the first The energy consumption of a robot traveling to the location of goods. For the first The cost of a robot traveling to the location of goods. , and The reward function, with the corresponding weighting coefficients, is as follows: According to the designed state space S Action space D and reward function R A task matching model for the robot was established using the DQN deep reinforcement learning algorithm, and the selected robots for cluster handling were determined as follows: ; The selected cluster handling robot adopts The algorithm performs path planning in the self-organizing assembly phase and combines it with a dynamic window algorithm to achieve local obstacle avoidance. It also performs self-organizing formation path planning using a cluster-based leader-follower method. The Model Predictive Control (MPC) method is used to adjust the speed and acceleration of each robot in real time, achieving collaborative speed optimization for multi-robot isomorphic platooning. Specifically: The overall state transition equations for the entire robot formation are as follows: in, It's a robot. The state at any given moment, It's a robot. The state at any given moment, For the input matrix, The state matrix, , as follows: The overall constraints for queue stability control include acceleration constraints, velocity constraints, and formation constraints, as follows: in, The maximum speed in the speed constraint is the speed calculated in the spatiotemporal characteristic function of cargo handling. ; With the optimization objectives of low time cost, stable driving, and maintaining formation, the cost function is defined as follows: in, Indicates the actual arrival time of the robot. For the target arrival time, For the acceleration of the robot's movement, These are the weighting coefficients.

2. The self-organizing formation and collaborative scheduling method for multi-vehicle isomorphic robot clusters as described in claim 1, characterized in that, The self-organizing assembly stage path planning is as follows: the A* algorithm is used to generate an initial path based on a static environment map for each robot from its starting position to the common target position, and the dynamic window method (DWA) is combined to adjust the robot's movement in real time, thus solving the problem of the robot relying solely on the A* algorithm for dynamic obstacles. The self-organizing swarm path planning process is as follows: After the robots reach the cargo location, they form a swarm to transport the cargo to the destination. A swarm-based leader-follower method is used for path planning, with the robot closest to the cargo's center of gravity selected as the leader robot. A map of the robot swarm's operating environment is pre-built on the server, using information about other robots outside the swarm as dynamic obstacle information. When the leader robot receives the destination location information, it sends a path planning request to the server. The server performs global path planning, generating the optimal path based on the leader robot's request and the Floyd algorithm, and sends it to the leader robot. The leader robot, based on the optimal path and its kinematic model, generates a path deviation from the optimal path based on environmental information, and then performs real-time deviation correction on the actual path, performing local planning to determine the local route and desired speed. The following robots calculate their own travel route and desired speed based on the leader robot's planned local route and desired speed.

3. A multi-vehicle isomorphic robot swarm self-organizing formation operation collaborative scheduling system that implements the multi-vehicle isomorphic robot swarm self-organizing formation operation collaborative scheduling method of claim 1, characterized in that, include: The first processing module is used to construct the representation function of the spatiotemporal characteristics of cargo handling; The second processing module is used to determine the target scheduling object of the cluster robot task matching optimal decision strategy based on deep reinforcement learning. The third processing module is used to select the cluster handling robot. The algorithm performs path planning in the self-organizing assembly phase and combines it with a dynamic window algorithm to achieve local obstacle avoidance. It also performs self-organizing formation path planning using a cluster-based leader-follower method. The fourth processing module is used to adjust the speed and acceleration of each robot in real time using the Model Predictive Control (MPC) method, so as to achieve collaborative optimization of the speed of multi-robot isomorphic robot formation driving.

Citation Information

Patent Citations

  • Multi-agent reinforcement learning-based multi-AGV task scheduling method

    CN118333254A