A multi-robot cooperation method and device, electronic equipment and storage medium
Through the virtual center of mass module and consensus algorithm module in the MRVCC model, the accuracy and stability problems of multi-robot formations are solved, efficient formation control and obstacle avoidance capabilities are achieved, and the coordination and flexibility of the multi-robot system are improved.
Patent Information
- Application Number
- CN202411417815.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-10-11
AI Technical Summary
Existing multi-robot collaborative formation methods have limitations in accuracy, flexibility and stability, especially in that the navigation structure is greatly affected by single-point failures, the hierarchical control structure is highly complex and has information transmission delay problems, and the hybrid control structure is unstable during strategy switching.
The MRVCC model is adopted, including the virtual center of mass module, the consensus algorithm module and the hierarchical strategy control module. By building a virtual center of mass node shared by multiple machines, collaborative control is carried out based on formation information and radar data to achieve the planning, coordination and obstacle avoidance of the robot formation.
It improves the accuracy and flexibility of multi-robot formations, reduces information interaction delays, enhances the stability and coordination of formations, and enables flexible control of robots in different states.
Smart Images

Figure CN119396139B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot control technology, and in particular to a multi-robot collaboration method, device, electronic equipment and storage medium. Background Art
[0002] As one of the important research directions in the field of robotics, multi-robot collaborative formation control is committed to solving the problem of collaborative motion between multiple robots, including core tasks such as navigation, formation formation and maintenance, and obstacle avoidance. Research in this field is not only related to the intelligence and collaboration of robot systems, but also directly affects the actual performance of multi-robot systems in practical applications.
[0003] Multi-robot collaborative formations can be mainly divided into follower-navigation structure, hierarchical structure, and hybrid control structure. The follower-navigation structure selects a robot as the navigator to make independent decisions, and the remaining robots serve as followers. The followers can quickly respond to the navigator's instructions; the hierarchical control structure divides the multi-robot system into task planning layer, path planning layer, and motion control layer, achieving a certain degree of combination of independent control and centralized control; the hybrid control structure dynamically selects the appropriate control strategy according to the specific situation: centralized control is used in static environments, and distributed control is used in dynamic environments or when encountering obstacles.
[0004] However, the navigation structure is greatly affected by the single-point failure of the navigator during movement. If only the navigator makes independent decisions, the robustness and flexibility of the formation coordination will be reduced; the hierarchical control structure system is highly complex and requires the design and implementation of multi-level control algorithms and communication mechanisms, and there are problems with information transmission and delay; the hybrid control structure is unstable or inconsistent during the switching process between the two control strategies; the above three methods have certain limitations in terms of accuracy, formation flexibility and stability, and the applicable scenarios are relatively single. Summary of the Invention
[0005] In view of this, it is necessary to provide a multi-robot collaboration method, device, electronic device and storage medium to solve the technical problem that the accuracy, flexibility and stability of multi-robot collaborative formations are limited.
[0006] In order to solve the above problems, the present invention provides a multi-robot collaboration method for collaboratively controlling multiple robots based on a constructed MRVCC model; the MRVCC model includes a virtual centroid module, a consensus algorithm module, and a hierarchical strategy control module;
[0007] The multi-robot collaboration method comprises:
[0008] Constructing a virtual center of mass node shared by multiple robots based on the virtual center of mass module, planning the behavior of multiple robots based on the virtual center of mass node, generating a multi-robot formation, and obtaining formation information;
[0009] Coordinating the behaviors of the multiple robots based on the formation information and the consensus algorithm module, and obtaining a minimum formation cumulative error;
[0010] Based on the acquired radar data and odometer information, the status of the robots in the multi-robot formation is determined. Based on the status of the robots, the hierarchical strategy control module is used to execute multiple strategy controls. Based on the multiple strategy controls and the minimum formation cumulative error, the multiple robots are collaboratively controlled to complete the collaborative formation and obstacle avoidance goals of the multiple robots, wherein the multiple strategy controls include safety collaboration, obstacle avoidance decision-making, and formation stability detection.
[0011] In a possible implementation, planning the behavior of multiple robots based on the virtual center of mass node includes:
[0012] Constructing a SAC model, obtaining a policy network and a value network of the robot, training the policy network and the value network of the robot based on the SAC model, obtaining an average reward value of the robot based on the value network of the robot, determining an optimal robot policy network based on the average reward value, and determining a center of mass motion model of the robot based on the optimal robot policy network;
[0013] After planning the formation coordinate positions of the multi-robot formation based on the virtual center of mass node, the robots are navigated based on the formation coordinate positions using the center of mass motion model of the robots to meet the formation requirements and obtain the position coordinates of the robot nodes in the multi-robot formation;
[0014] Determine a target formation position, determine an optimal path based on the target formation position, the position coordinates, and the virtual center of mass node, and plan the behavior of the multi-robots based on the optimal path and the center of mass motion model of the multi-robots so that the multi-robot formation reaches the target formation position.
[0015] In one possible implementation, the SAC model includes an Actor network and a Critic network; and training the robot's policy network and value network based on the SAC model includes:
[0016] Learning the mapping relationship between different states and actions of the robot based on the Actor network;
[0017] The current state of the robot and the actions taken based on the current state are scored based on the Critic network.
[0018] In a possible implementation, the coordinating the behaviors of the plurality of robots based on the formation information and the consensus algorithm module comprises:
[0019] determining a robot node and a master node in the plurality of robot formation, wherein the master node is a virtual centroid node;
[0020] obtaining a position of the robot node and the virtual centroid node based on the formation information, and calculating a relative position vector of the robot node and the virtual centroid node based on the position of the robot node and the virtual centroid node;
[0021] obtaining an expected coordinate of the robot node, and calculating an expected relative position vector of the robot node and the virtual centroid node based on the expected coordinate and the position of the virtual centroid node;
[0022] determining an error of the robot node and the virtual centroid node based on the relative position vector and the expected relative position vector, and determining a formation cumulative error based on the error of the robot node and the virtual centroid node;
[0023] based on the formation cumulative error, coordinating the behaviors of the robot node through the consensus algorithm module, and controlling the communication of the robot node through the virtual centroid node, so that the formation cumulative error of the robot formation is minimized, wherein the behaviors of the robot node are controlled through the centroid motion model.
[0024] In a possible implementation, the controlling the communication of the robot node through the virtual centroid node comprises:
[0025] Step 1, obtaining a formation position parameter of the robot node through the virtual centroid node, and completing the plurality of robot formation through the centroid motion model based on the formation position parameter;
[0026] Step 2, after the plurality of robot formation is formed, refreshing the action control information of the robot node through the virtual centroid node, and sending a communication topic to the robot node;
[0027] Step 3, obtaining topic information of the communication topic, sending the topic information to a motion controller of the robot node, obtaining information of the robot node after the robot is controlled to complete its own motion through the motion controller, and feeding back the information of the robot node to the virtual centroid node to obtain a coordinate and an expected coordinate of the robot node, and determining a formation cumulative error based on the coordinate and the expected coordinate of the robot node;
[0028] Step 4: After the virtual center of mass node obtains the information of the robot node, steps 2 to 4 are repeated.
[0029] In a possible implementation, the calculation formula for the formation cumulative error is:
[0030] ,
[0031] ,
[0032] in, is the cumulative error of the formation, The robot node and the virtual center of mass node are The error in time, for The relative position vector between the robot node and the virtual center of mass node at this moment, for The expected relative position vector of the robot node and the virtual center of mass node at time .
[0033] In one possible implementation, the states of the robots in the multi-robot formation are determined based on the acquired radar data and odometer information, and the hierarchical strategy control module is used to perform multiple strategy controls based on the states of the robots, including:
[0034] Presetting a safe distance between the robot and the obstacle, determining the distance between the robot and the obstacle based on the radar data, comparing the distance between the robot and the obstacle with the safe distance, and when the distance between the robot and the obstacle is greater than the safe distance, using the hierarchical strategy control module to perform safety coordination, wherein the movement of the robot node is guided by the virtual center of mass node;
[0035] When the distance between the robot and the obstacle is less than or equal to the safety distance, the hierarchical strategy control module is used to execute obstacle avoidance decision, wherein the movement of the robot node is controlled by the center of mass motion model of the robot node;
[0036] Determine the target formation position, preset the robot's residence time threshold and the distance threshold between the robot and the target formation position, determine the distance between the robot and the target formation position and the robot's residence time based on the odometer information, and when the distance between the robot and the target formation position is greater than the distance threshold and the residence time is greater than the residence time threshold, use the hierarchical strategy control module to perform formation stability detection, wherein the target formation position information is obtained, and based on the target formation position information, the robot node movement is controlled through the center of mass motion model of the robot node to make the robot state reach a stable state.
[0037] In another aspect, the present application also provides a multi-robot coordination device for coordinating control of a plurality of robots based on a constructed MRVCC model; the MRVCC model comprises a virtual centroid module, a consensus algorithm module, and a hierarchical strategy control module.
[0038] The multi-robot coordination device comprises:
[0039] a virtual centroid unit for constructing a virtual centroid node shared by the plurality of robots based on the virtual centroid module, planning behaviors of the plurality of robots based on the virtual centroid node, generating a multi-robot formation, and obtaining formation information;
[0040] a consensus algorithm unit for coordinating the behaviors of the plurality of robots based on the formation information and the consensus algorithm module, and obtaining a minimum formation cumulative error;
[0041] a hierarchical strategy control unit for determining states of the robots in the multi-robot formation based on the obtained radar data and odometry information, executing a plurality of strategy controls based on the states of the robots using the hierarchical strategy control module, and coordinating control of the plurality of robots based on the plurality of strategy controls and the minimum formation cumulative error, to achieve the goals of coordinated formation and obstacle avoidance of the plurality of robots, wherein the plurality of strategy controls comprise safety coordination, obstacle avoidance decision, and formation stability detection.
[0042] In another aspect, the present application also provides an electronic device comprising a processor and a memory.
[0043] The memory has stored thereon a computer readable program executable by the processor.
[0044] The processor, when executing the computer readable program, implements the steps in the multi-robot coordination method as described above.
[0045] In another aspect, the present application also provides a computer readable storage medium having stored thereon one or more programs executable by one or more processors to implement the steps in the multi-robot coordination method as described above.
[0046] The beneficial effects of the present application are: a virtual centroid node shared by multiple robots is constructed, the robots in the multi-robot formation obtain their formation positions through the shared virtual centroid, direct communication between the robots is replaced, communication delay caused by information interaction between the robots is effectively reduced, response speed of the multi-robot formation is accelerated, precise formation is realized, accuracy of the multi-robot formation is improved, a consensus algorithm based on the multi-robot formation is introduced, each robot in the multi-robot formation is regarded as an independent decision unit, each robot has independent perception, decision and execution capabilities; meanwhile, the robots promote formation and maintenance of the formation according to shared information in the virtual centroid, maximum cooperation of the multi-robot is improved, a layered strategy control module executes safety cooperation, obstacle avoidance decision and formation stability detection, so that the robots in the multi-robot formation can adopt corresponding control strategies in different states, to ensure flexibility and stability of the multi-robot formation. BRIEF DESCRIPTION OF DRAWINGS
[0047] Figure 1 A structure schematic diagram of an embodiment of the MRVCC model provided by the present application is shown in the figure.
[0048] Figure 2 A flowchart of an embodiment of the multi-robot cooperation method provided by the present application is shown in the figure.
[0049] Figure 3 A topological structure schematic diagram of the virtual centroid and the robot node of the multi-robot cooperation method provided by the present application is shown in the figure.
[0050] Figure 4 A schematic diagram of the consensus algorithm based on the centroid control of the multi-robot cooperation method provided by the present application is shown in the figure.
[0051] Figure 5 A structure schematic diagram of an embodiment of the multi-robot cooperation device provided by the present application is shown in the figure.
[0052] Figure 6 A structure schematic diagram of an embodiment of the electronic device provided by the present application is shown in the figure. DETAILED DESCRIPTION
[0053] The preferred embodiments of the present application are specifically described below in combination with the drawings, wherein the drawings form a part of the present application, and are used to explain the principles of the embodiments of the present application, and are not used to limit the scope of the present application.
[0054] The present application discloses a multi-robot cooperation method, device, electronic equipment and storage medium, which can be used in a computer. The method, device or computer readable storage medium involved in the present application can be integrated together with the above-mentioned device, or can be relatively independent.
[0055] A specific embodiment of the present invention discloses a multi-robot collaboration method, which can be executed by a computer, specifically by one or more processors of the computer. Figure 1 As shown, the MRVCC model 100 includes a virtual centroid module 110, a consensus algorithm module 120, and a hierarchical strategy control module 130, and performs collaborative control of multiple robots based on the constructed MRVCC model; Figure 2 As shown, the multi-robot collaboration method includes:
[0056] S201, constructing a virtual center of mass node shared by multiple robots based on the virtual center of mass module 110, planning the behavior of multiple robots based on the virtual center of mass node, generating a multi-robot formation, and obtaining formation information;
[0057] S202, coordinating the behaviors of multiple robots based on the formation information and the consensus algorithm module 120, and obtaining the minimum formation cumulative error;
[0058] S203. Determine the status of the robots in the multi-robot formation based on the acquired radar data and odometer information. Based on the status of the robots, use the hierarchical strategy control module 130 to execute multiple strategy controls. Based on the multiple strategy controls and the minimum formation cumulative error, coordinate control of the multiple robots to complete the coordinated formation and obstacle avoidance goals of the multiple robots. The multiple strategy controls include safety coordination, obstacle avoidance decision-making, and formation stability detection.
[0059] Among them, the MRVCC model is a multi-robot formation model based on virtual center of mass and consensus algorithm. The MRVCC model includes a virtual center of mass module, a consensus algorithm module and a hierarchical strategy control module. The virtual center of mass module transforms the centralized control of multi-robot collaborative formation into virtual center of mass control. The virtual center of mass completes the behavior planning of robots in the effective area, realizing a high degree of coordination in the formation maintenance process. The consensus algorithm module regards each robot in the multi-machine system as an independent node with environmental perception, behavior decision-making, and motion control. Each robot obtains formation information and motion control information at the virtual center of mass according to the consensus algorithm, solving the problem of asynchronous robot action execution caused by interaction delay. The consensus algorithm retains the flexibility of the robot's autonomous decision-making while coordinating multiple machines, and makes independent decisions according to its own situation during the formation formation process, ensuring the flexibility of real-time obstacle avoidance and the efficiency of the task during the formation process.
[0060] Compared with the prior art, the multi-robot collaboration method provided in this embodiment constructs a virtual center of mass node shared by multiple robots based on a virtual center of mass module. The behavior of multiple robots is planned based on the virtual center of mass node, a multi-robot formation is generated, and formation information is obtained. The centralized control of the multi-robot collaborative formation is converted to virtual center of mass control, achieving high coordination in the formation maintenance process. The behavior of multiple robots is coordinated based on the formation information and the consensus algorithm module, and the minimum formation cumulative error is obtained, so that the multi-robot formation maintains a stable state. While coordinating multiple robots, the flexibility of autonomous decision-making is retained. During the formation formation process, the robots make independent decisions based on their own conditions, ensuring the flexibility of real-time obstacle avoidance and the efficiency of the task during the formation formation process. Based on the robot status, a hierarchical strategy control module is used to execute multiple strategy controls. Based on the multiple strategy controls and the minimum formation cumulative error, the multiple robots are collaboratively controlled to achieve the collaborative formation and obstacle avoidance goals of the multiple robots. The robot's status is determined by its perception results, and different strategy controls are adopted based on its own state perception to ensure the flexibility and stability of the multi-robot formation.
[0061] In some embodiments, in step S201, a virtual center of mass node shared by multiple machines is constructed based on a virtual center of mass module, and a topic function of the ROS (Robot Operation System) system is used to construct a virtual center of mass node shared by multiple machines, which is similar to a shared space (Shared Memory). The introduction of a virtual center of mass node will replace the physical navigator in the collaborative motion of multiple robots to control the action of the formation. In a multi-robot system, if the virtual center of mass fails, the faulty node information can be directly copied to a new virtual node. Compared with the physical node, which requires reinitializing the robot and reloading the information and communication records of all formation nodes, the time can be greatly shortened and the efficiency can be improved. The behavior of multiple robots is planned based on the virtual center of mass node, a multi-robot formation is generated, and formation information is obtained. First, in order to ensure the efficiency of virtual center of mass decision-making, SAC (Soft Actor) is used in an environment with obstacles. Critic) trains multiple robot policy networks and value networks at the same time, selects the policy network with the best performance as the center of mass motion model, controls the robot's behavior through the center of mass motion model, builds a SAC model, obtains the robot's policy network, trains the robot's policy network and value network based on the SAC model, obtains the robot's average return value based on the robot's value network, determines the optimal robot policy network based on the average return value, determines the robot's center of mass motion model based on the optimal robot policy network, and the process of training the robot's policy network and value network based on the SAC model can be divided into an Actor network and a Critic network. The mapping relationship between different states and actions of the robot is learned based on the Actor network, and the current state of the robot is determined based on the Critic network. and actions to take based on the current state Scoring is performed to evaluate the quality of the current Actor network; the interaction process of the robot in an environment with obstacles conforms to the Markov decision process, which can be represented by the tuple Indicates that, for The robot's status information at all times includes at least position information, lidar sensor data information, and speed information. for The actions taken at any time based on the current state of the robot, Adopting actions for robots The reward value obtained after the action is evaluated by the reward value. To take action After the robot Status information at all times. The robot obtains its own status information through radar data and odometer ,in, is 360-dimensional radar sensor data, is the distance between the robot’s current position and the target position, is the current direction, is the linear velocity of the robot, is the angular velocity, the speed of the robot , the calculation formula for constraining the robot's linear velocity and angular velocity is:
[0062] ,
[0063] in, is the maximum linear speed, is the maximum angular velocity;
[0064] Its reward function The calculation formula is:
[0065] ,
[0066] in, is the difference between the current distance between the robot and the target position and the distance between the robot and the target position at the previous moment. When , the robot moves towards the target position. When , the robot is far away from the target position, The robot's lidar sensor detects the distance between the current position and the obstacle. When the minimum distance exceeds the safety distance, When the robot has a high probability of collision, a certain penalty is given, that is, a negative reward value. When the distance between the robot's current position and the target position is Within a certain range, it is determined that it has reached the target position and a positive reward value is given.
[0067] Secondly, after planning the formation coordinate position of the multi-robot formation based on the virtual center of mass node, the robot is navigated based on the formation coordinate position through the robot's center of mass motion model to complete the formation requirements and obtain the position coordinates of the robot nodes in the multi-robot formation. In addition to containing the same status information data as the robot node, the virtual center of mass node information also includes the current formation parameter information. In the formation formation stage, each robot obtains its own position coordinates in the current formation in the virtual center of mass node. When the robot Robot_i takes a certain coordinate as its target position, the virtual center of mass node extracts the current coordinate from the selection space, ensuring the exclusivity of the multi-robot system in the formation position, and each robot is navigated through the center of mass motion model obtained by its own training to complete the corresponding formation requirements.
[0068] When all robots in the current environment reach their own formation positions and meet the target formation requirements, the virtual center of mass switches to the formation holding state. Taking the regular quadrilateral and triangle formations as examples, a virtual center of mass node can be introduced at the origin. Each robot obtains formation parameters at the center of mass to achieve a precise target formation. For a schematic diagram of the topological structure of the virtual center of mass and robot nodes, please refer to Figure 3 In this process, each robot randomly initializes its own position and autonomously navigates to the formation coordinate position planned by the center of mass node; in the formation formation stage, the center of mass coordinates of the multi-robots are In a triangle formation, the three robots at the ends of the triangle , , the calculation formula of the parameter position coordinates of the three robots is:
[0069] ,
[0070] ,
[0071] ,
[0072] In order to further reduce the communication burden within the multi-robot system, a publisher node is first initialized as the virtual center of mass, and the initial position coordinates of the virtual center of mass are defined as the origin (0,0). Among them, / Center is the virtual center of mass, which can interact with any robot in the current multi-robot collaborative formation through topic communication. The communication speed of this method is faster than the point-to-point communication between robots, which reduces the communication delay of information transmission between robots and improves the feedback speed of robots, ensuring the execution efficiency of multi-robot collaborative tasks. Therefore, the introduction of the virtual center of mass is the basis for the current multi-robot system to maintain formation in real time.
[0073] Determine the target formation position, determine the optimal path based on the target formation position, the position coordinates of the robots and the virtual center of mass node, and plan the behavior of the multi-robots based on the optimal path and the center of mass motion model of the multi-robots so that the multi-robot formation reaches the target formation position.
[0074] In some embodiments, in step S202, the behaviors of multiple robots are coordinated based on the formation information and the consensus algorithm module, and the minimum formation cumulative error is obtained to keep the multi-robot formation in a stable state. The consensus algorithm is applied to the multi-robot collaborative formation to achieve the necessary formation maintenance of the multi-robot system, that is, to coordinate the robots in the current multi-machine system to maintain a relatively stable collaborative state. The specific process is to keep the attributes or states of the robots stable through local information exchange of the virtual center of mass, such as relative position relationship, speed, direction, etc., to promote the formation and maintenance of the formation. In this process, each robot is regarded as a node in the topological structure, and each node is connected to the virtual center of mass node by an edge. The virtual center of mass node has the same form of state information as the robot, including position coordinates, operation control data, etc., but because it is a virtual node, it does not need to avoid obstacles, so its external sensor data is not recorded. The virtual center of mass node will calculate the optimal path to reach the target position in the current environment and the robot motion control instructions. For a schematic diagram of the consensus algorithm based on center of mass control, please refer to Figure 4 ,like Figure 4 As shown in the figure, the solid box represents the initial position of the robot, the dotted box represents the target position, the straight line represents the moving trajectory, the dot represents the virtual center of mass, and the small flag represents the target position. The existing target navigation algorithm can only realize the independent route planning of each robot based on the target position, and does not consider the collaborative planning problem in the process of multi-robot collaborative formation. Collaborative planning still depends on the physical leader node. After the introduction of virtual center of mass control, each robot can ignore its own decision-making when there is no need to avoid obstacles, and only follow the behavioral instructions sent by the virtual center of mass node to complete the action control. This ensures the collaborative formation control of multi-robots without obstacle avoidance requirements during movement. At the same time, the center of mass consensus algorithm also reduces the computing consumption brought by point-to-point multi-machine communication.
[0075] Determine the robot nodes and master nodes in the multi-robot formation, where the master node is the virtual center of mass node. The virtual center of mass node / Center is used as the master node through the consensus algorithm. It mainly plays a role in the formation maintenance stage. The positions of the robot nodes and the virtual center of mass node are obtained based on the formation position information. The relative position vector of the robot node and the virtual center of mass node is calculated based on the positions of the robot nodes and the virtual center of mass node. The calculation formula is:
[0076] ,
[0077] in, for The relative position vector between the virtual center of mass and the robot node Robot_i at this moment; for The position coordinates of the virtual center of mass node at the moment, for The position coordinates of the robot node at the moment;
[0078] Get the expected coordinates of the robot node, and calculate the expected relative position vector between the robot node and the virtual center of mass node based on the expected coordinates and the position of the virtual center of mass node. The calculation formula is:
[0079] ,
[0080] in, for The expected relative position vector between the robot node and the virtual center of mass node at the moment, for The expected coordinates of the robot node at the moment;
[0081] The error between the robot node and the virtual center of mass node is determined based on the relative position vector and the expected relative position vector. The cumulative error of the formation is determined based on the error between the robot node and the virtual center of mass node. The error calculation formula is:
[0082] ,
[0083] in, The robot node and the virtual center of mass node are Error in time;
[0084] The calculation formula for the formation's cumulative error is:
[0085] ,
[0086] in, is the cumulative error of the formation;
[0087] Based on the formation cumulative error, the behavior of the robot nodes is coordinated through the consensus algorithm module, and the communication of the robot nodes is controlled through the virtual center of mass node to minimize the formation cumulative error of the robot formation. Among them, the behavior of the robot nodes is controlled by the robot's center of mass motion model. When the formation cumulative error If it converges to 0, it means that the robot always maintains the required formation during the movement and achieves the desired formation maintenance task.
[0088] The process of controlling the communication of robot nodes through the virtual center of mass node is:
[0089] Step 1: Obtain the formation position parameters of the robot nodes through the virtual center of mass node. Each robot obtains its own formation position parameters at the virtual center of mass node. , and reaches the corresponding formation position through its own strategy network, and its own strategy network is a center of mass motion model;
[0090] Step 2: After the multi-robot formation is formed, the robot motion control information is continuously refreshed through the virtual center of mass node / Center, and the communication topic is sent to the robot node;
[0091] Step 3: Robot_i receives the topic information and sends it to its own motion controller. After the motion controller controls the robot to complete its own motion, it feeds back its own information to / Center. Its own information includes The coordinates of the robot at that moment , / Center refresh The expected coordinates of the robot at this moment ,determine the formation cumulative error based on the coordinates of the robot nodes and the expected coordinates;
[0092] Step 4: After the center node / Center receives the new information of Robot_i, it repeats steps 2 to 4.
[0093] In some embodiments, in step S203, the virtual center of mass module provides fast communication and feedback conditions for the collaborative maintenance of each robot formation, and the consensus algorithm module provides an execution algorithm for the behavior control of each robot, but the two need to cooperate to jointly realize the formation collaborative task. At the same time, each robot also needs to independently decide on the local route planning problem of obstacle avoidance when there are obstacles. Therefore, in order to cooperate with the simultaneous execution of the virtual center of mass, consensus algorithm and autonomous obstacle avoidance, different strategy control methods are designed for different states of each robot in the multi-machine system; the state of the robot in the multi-robot formation is determined based on the acquired radar data and odometer information, and based on the state of the robot, a hierarchical strategy control module is used to execute multiple strategy controls, including safety collaboration, obstacle avoidance decision-making and formation stability detection. First, the state of the robot in the multi-robot formation is judged, and the safe distance between the robot and the obstacle is preset. , based on the radar data, determine the distance between the robot and the obstacle, and compare the distance between the robot and the obstacle with the safety distance For comparison, when the distance between the robot and the obstacle is greater than the safe distance, the robot is in a safe state, and the calculation formula for the judgment condition of its safe state is:
[0094] ,
[0095] When the distance between the robot and the obstacle is less than or equal to the safety distance When , the robot is in the obstacle avoidance state, and the calculation formula for the obstacle avoidance state is:
[0096] ,
[0097] Determine the target formation position, preset the robot's residence time threshold and the distance threshold between the robot and the target formation position, and determine the distance between the robot and the target formation position and the robot's residence time based on the odometer information. When the distance between the robot and the target formation position is greater than the distance threshold and the residence time is greater than the residence time threshold, the robot is in an unstable state. The calculation formula for the judgment condition of the unstable state is:
[0098] ,
[0099] When the robot is in a safe state, the hierarchical strategy control module is used to perform safe collaboration. According to the update rules of the consensus algorithm, the cumulative error of the multi-robot formation is calculated. The cumulative error of the formation is minimized based on the consensus algorithm. The virtual center of mass node is used to guide the robot to move in the direction of the virtual center of mass. Its action decision adopts the same action as the virtual center of mass to ensure the consistency of the collaborative motion, that is, ;
[0100] When the robot is in the obstacle avoidance state, the hierarchical strategy control module is used to execute the obstacle avoidance decision. In the obstacle avoidance state, if the current robot continues to move in the direction of the current virtual center of mass, it will be dangerous. Therefore, the robot's action decision uses its own trained network , guiding the robots to make decisions according to their own trained networks. That is, the robot nodes’ movements are controlled through their center of mass motion models to determine the movement direction of multiple robots and avoid obstacles. This obstacle avoidance decision ensures that the robots can flexibly adjust their movement paths when facing obstacles.
[0101] When the robot is in an unstable state, the hierarchical strategy control module is used to perform formation stability detection. In an unstable state, the robot usually fails to reach its target formation position within a certain period of time due to current decision-making errors. The multi-robot collaborative formation falls into an unstable state and the current robot needs to subscribe to its target formation position information again. In the current state, based on the target formation position information, the robot node movement is controlled by the center of mass motion model of the robot node. The robot's action decision uses its own trained network , making independent decisions based on the formation's target position to complete the formation requirements and restore the formation to a stable state.
[0102] The three types of strategy control jointly realize the multi-robot collaborative formation process, ensuring that the robots can complete the formation maintenance task under normal circumstances, allowing the robots to change the formation shape and safely avoid the current obstacle when encountering an obstacle, and can restore to a stable formation state in the event of a long period of no communication or decision-making errors. This comprehensive guarantee mechanism enables the multi-robot collaborative formation system to reliably perform tasks in various complex scenarios, and comprehensively realizes the formation formation and maintenance, formation obstacle avoidance, and formation stability detection tasks of the multi-robot collaborative formation.
[0103] In order to better implement the multi-robot collaboration method in the embodiment of the present invention, based on the multi-robot collaboration method, correspondingly, Figure 5 As shown, an embodiment of the present invention further provides a multi-robot collaborative device for collaboratively controlling multiple robots using an MRVCC model constructed; the MRVCC model includes a virtual centroid module, a consensus algorithm module, and a hierarchical strategy control module; the multi-robot collaborative device 500 includes:
[0104] A virtual center of mass unit 501 is used to construct a virtual center of mass node shared by multiple robots based on the virtual center of mass module, plan the behavior of multiple robots based on the virtual center of mass node, generate a multi-robot formation, and obtain formation information;
[0105] The consensus algorithm unit 502 is used to coordinate the behaviors of multiple robots based on the formation information and the consensus algorithm module, and obtain the minimum formation cumulative error;
[0106] The hierarchical strategy control unit 503 is used to determine the status of the robots in the multi-robot formation based on the acquired radar data and odometer information. Based on the status of the robots, a hierarchical strategy control module is used to execute multiple strategy controls. Based on the multiple strategy controls and the minimum formation cumulative error, the multiple robots are collaboratively controlled to complete the collaborative formation and obstacle avoidance goals of the multiple robots. Among them, the multiple strategy controls include safety collaboration, obstacle avoidance decision-making and formation stability detection.
[0107] like Figure 6 The present invention also provides an electronic device 600, which can be a computing device such as a mobile terminal, a desktop computer, a notebook, a PDA, or a server. The electronic device 600 includes a processor 601, a memory 602, and a display 603. Figure 6 Only some of the components of the electronic device 600 are shown, but it should be understood that implementation of all of the shown components is not required, and more or fewer components may be implemented instead.
[0108] The memory 602 can be an internal storage unit of the electronic device 600, such as a hard disk or a memory of the electronic device 600 in some embodiments. The memory 602 can also be an external storage device of the electronic device 600, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, and the like equipped on the electronic device 600 in other embodiments. Further, the memory 602 can include both an internal storage unit and an external storage device of the electronic device 600. The memory 602 is used to store application software and various data installed on the electronic device 600, such as program codes installed on the electronic device 600. The memory 602 can also be used to temporarily store data that has been output or will be output. In an embodiment, the memory 602 stores a multi-robot collaboration program that can be executed by the processor 601 to implement the multi-robot collaboration method of the embodiments of the present application.
[0109] The processor 601 can be a Central Processing Unit (CPU), a microprocessor, or other data processing chip in some embodiments, and is used to run program codes or process data stored in the memory 602, such as the multi-robot collaboration method.
[0110] The display 603 can be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch, and the like in some embodiments. The display 603 is used to display identification information of the multi-robot collaboration program and to display a visual user interface. The components 601-603 of the electronic device 600 communicate with each other through a system bus.
[0111] In some embodiments, the processor 601 implements each step of the multi-robot collaboration method as described in the above embodiments when executing the multi-robot collaboration program in the memory 602. Since the multi-robot collaboration method has been described in detail above, no further description is given here.
[0112] Correspondingly, the embodiments of the present application also provide a computer readable storage medium for storing computer readable programs or instructions, which, when executed by a processor, can implement the steps or functions of the multi-robot collaboration method provided by the above method embodiments.
[0113] In summary, the multi-robot collaboration method, device, electronic device and storage medium provided by the present invention are used to collaboratively control multiple robots based on the constructed MRVCC model; the MRVCC model includes a virtual center of mass module, a consensus algorithm module and a hierarchical strategy control module; the multi-robot collaboration method includes: constructing a virtual center of mass node shared by multiple machines based on the virtual center of mass module, planning the behavior of multiple robots based on the virtual center of mass node, generating a multi-robot formation, and obtaining formation information; coordinating the behavior of multiple robots based on the formation information and the consensus algorithm module, and obtaining the minimum formation cumulative error; determining the status of the robots in the multi-robot formation based on the acquired radar data and odometer information, and based on the status of the robots, using the hierarchical strategy control module to execute multiple strategy controls, and collaboratively controlling multiple robots based on multiple strategy controls and the minimum formation cumulative error to complete the collaborative formation and obstacle avoidance goals of multiple robots.
[0114] Those skilled in the art will appreciate that all or part of the process steps of the above-described embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, such as a magnetic disk, an optical disk, a read-only memory, or a random access memory.
[0115] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.
Claims
1. A multi-robot collaboration method, characterized in that: Used to collaboratively control multiple robots based on the constructed MRVCC model; the MRVCC model includes a virtual center of mass module, a consensus algorithm module, and a hierarchical strategy control module; The multi-robot collaboration method comprises: A virtual center of mass node shared by multiple robots is constructed based on the virtual center of mass module, and behaviors of multiple robots are planned based on the virtual center of mass node to generate a multi-robot formation and obtain formation information, wherein the behaviors of multiple robots are planned based on the virtual center of mass node, including: Constructing a SAC model, obtaining a policy network and a value network of the robot, training the policy network and the value network of the robot based on the SAC model, obtaining an average reward value of the robot based on the value network of the robot, determining an optimal robot policy network based on the average reward value, and determining a center of mass motion model of the robot based on the optimal robot policy network; After planning the formation coordinate positions of the multi-robot formation based on the virtual center of mass node, the robots are navigated based on the formation coordinate positions using the center of mass motion model of the robots to meet the formation requirements and obtain the position coordinates of the robot nodes in the multi-robot formation; Determining a target formation position, determining an optimal path based on the target formation position, the position coordinates, and the virtual center of mass node, and planning the behavior of the multi-robots based on the optimal path and the center of mass motion model of the multi-robots so that the multi-robot formation reaches the target formation position; Coordinating the behaviors of the multiple robots based on the formation information and the consensus algorithm module, and obtaining a minimum formation cumulative error; Based on the acquired radar data and odometer information, the status of the robots in the multi-robot formation is determined. Based on the status of the robots, the hierarchical strategy control module is used to execute multiple strategy controls. Based on the multiple strategy controls and the minimum formation cumulative error, the multiple robots are collaboratively controlled to complete the collaborative formation and obstacle avoidance goals of the multiple robots, wherein the multiple strategy controls include safety collaboration, obstacle avoidance decision-making, and formation stability detection.
2. The multi-robot collaboration method according to claim 1, characterized in that: The SAC model includes an Actor network and a Critic network; the strategy network and value network of the robot are trained based on the SAC model, including: Learning the mapping relationship between different states and actions of the robot based on the Actor network; The current state of the robot and the action taken in the current state are scored based on the Critic network.
3. The multi-robot collaboration method according to claim 1, characterized in that: The coordinating the behaviors of the multiple robots based on the formation information and the consensus algorithm module includes: Determining robot nodes and a master node in the multi-robot formation, wherein the master node is a virtual centroid node; Obtaining positions of a robot node and a virtual center of mass node based on the formation information, and calculating a relative position vector between the robot node and the virtual center of mass node based on the positions of the robot node and the virtual center of mass node; Obtaining the expected coordinates of the robot node, and calculating the expected relative position vector between the robot node and the virtual center of mass node based on the expected coordinates and the position of the virtual center of mass node; Determine an error between the robot node and the virtual center of mass node based on the relative position vector and the expected relative position vector, and determine a formation cumulative error based on the error between the robot node and the virtual center of mass node; Based on the formation cumulative error, the behavior of the robot nodes is coordinated through the consensus algorithm module, and the communication of the robot nodes is controlled through the virtual center of mass node to minimize the formation cumulative error of the robot formation, wherein the behavior of the robot nodes is controlled by the center of mass motion model.
4. The multi-robot collaboration method according to claim 3, characterized in that: The controlling the communication of the robot node through the virtual center of mass node comprises: Step 1: Obtain the formation position parameters of the robot nodes through the virtual center of mass node, and complete the multi-robot formation through the center of mass motion model based on the formation position parameters; Step 2: After the multi-robot formation is formed, the motion control information of the robot nodes is refreshed through the virtual center of mass node, and the communication topic is sent to the robot nodes; Step 3: Acquire topic information of the communication topic, send the topic information to the motion controller of the robot node, obtain information of the robot node after the robot completes its own movement through the motion controller, and feed the information of the robot node back to the virtual center of mass node to obtain the coordinates and expected coordinates of the robot node, and determine the formation cumulative error based on the coordinates of the robot node and the expected coordinates; Step 4: After the virtual center of mass node obtains the information of the robot node, steps 2 to 4 are repeated.
5. The multi-robot collaboration method according to claim 4, characterized in that: The calculation formula of the formation cumulative error is: , , in, is the cumulative error of the formation, The robot node and the virtual center of mass node are The error in time, for The relative position vector between the robot node and the virtual center of mass node at this moment, for The expected relative position vector of the robot node and the virtual center of mass node at time .
6. The multi-robot collaboration method according to claim 3, characterized in that: The states of the robots in the multi-robot formation are determined based on the acquired radar data and odometer information, and the hierarchical strategy control module is used to execute multiple strategy controls based on the states of the robots, including: Presetting a safe distance between the robot and the obstacle, determining the distance between the robot and the obstacle based on the radar data, comparing the distance between the robot and the obstacle with the safe distance, and when the distance between the robot and the obstacle is greater than the safe distance, using the hierarchical strategy control module to perform safety coordination, wherein the movement of the robot node is guided by the virtual center of mass node; When the distance between the robot and the obstacle is less than or equal to the safety distance, the hierarchical strategy control module is used to execute obstacle avoidance decision, wherein the movement of the robot node is controlled by the center of mass motion model of the robot node; Determine the target formation position, preset the robot's residence time threshold and the distance threshold between the robot and the target formation position, determine the distance between the robot and the target formation position and the robot's residence time based on the odometer information, and when the distance between the robot and the target formation position is greater than the distance threshold and the residence time is greater than the residence time threshold, use the hierarchical strategy control module to perform formation stability detection, wherein the target formation position information is obtained, and based on the target formation position information, the robot node movement is controlled through the center of mass motion model of the robot node to make the robot state reach a stable state.
7. A multi-robot collaborative device, characterized in that: Used to collaboratively control multiple robots based on the constructed MRVCC model; the MRVCC model includes a virtual center of mass module, a consensus algorithm module, and a hierarchical strategy control module; The multi-robot collaborative device comprises: A virtual center of mass unit is configured to construct a virtual center of mass node shared by multiple robots based on the virtual center of mass module, plan the behavior of multiple robots based on the virtual center of mass node, generate a multi-robot formation, and obtain formation information, wherein planning the behavior of multiple robots based on the virtual center of mass node includes: Constructing a SAC model, obtaining a policy network and a value network of the robot, training the policy network and the value network of the robot based on the SAC model, obtaining an average reward value of the robot based on the value network of the robot, determining an optimal robot policy network based on the average reward value, and determining a center of mass motion model of the robot based on the optimal robot policy network; After planning the formation coordinate positions of the multi-robot formation based on the virtual center of mass node, the robots are navigated based on the formation coordinate positions using the center of mass motion model of the robots to meet the formation requirements and obtain the position coordinates of the robot nodes in the multi-robot formation; Determining a target formation position, determining an optimal path based on the target formation position, the position coordinates, and the virtual center of mass node, and planning the behavior of the multi-robots based on the optimal path and the center of mass motion model of the multi-robots so that the multi-robot formation reaches the target formation position; A consensus algorithm unit, configured to coordinate the behaviors of the multiple robots based on the formation information and the consensus algorithm module, and obtain a minimum formation cumulative error; A hierarchical strategy control unit is used to determine the status of robots in the multi-robot formation based on the acquired radar data and odometer information, and based on the status of the robots, adopt the hierarchical strategy control module to execute multiple strategy controls, and collaboratively control the multiple robots based on the multiple strategy controls and the minimum formation cumulative error to complete the collaborative formation and obstacle avoidance goals of the multiple robots, wherein the multiple strategy controls include safety collaboration, obstacle avoidance decision-making and formation stability detection.
8. An electronic device, characterized in that: including memory and processor; The memory stores a computer-readable program executable by the processor; When the processor executes the computer-readable program, the steps of the multi-robot collaboration method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the multi-robot collaboration method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Four-rotor unmanned aerial vehicle formation cooperative maneuvering control method
CN114911265A