Unified obstacle avoidance method, control terminal and control system for distributed heterogeneous robots

By inputting the dynamic and static information of heterogeneous robots into the reinforcement learning model and transforming them into control information of the same dimensions, the unified control problem between heterogeneous robots is solved, the system design is simplified, maintenance costs are reduced, and the coordination and efficiency of multi-robot operations are improved.

CN120255507APending Publication Date: 2025-07-04HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510373243.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing obstacle avoidance methods are difficult to achieve unified control between heterogeneous robots, resulting in large training overhead, complex system design and high maintenance costs, and traditional methods are prone to deadlock states.

Method used

By obtaining the dynamic and static information of heterogeneous robots, inputting them into pre-trained reinforcement learning models, and converting them into control information of the same dimensions, achieving unified obstacle avoidance control for each robot.

Benefits of technology

It realizes unified processing of different types of robots, reduces system design and maintenance costs, improves the coordination and efficiency of multi-robot operations, and avoids collision accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255507A_ABST
    Figure CN120255507A_ABST
Patent Text Reader

Abstract

The invention provides a unified obstacle avoidance method of a distributed heterogeneous robot, a control terminal and a control system, and relates to the technical field of path planning. The method comprises the following steps: acquiring dynamic information and static information corresponding to each robot in heterogeneous robots; inputting the dynamic information and the static information corresponding to each robot into a pre-trained reinforcement learning model to obtain control information of each robot; wherein the control information of each robot is in the same dimension; and performing unified obstacle avoidance control on the robots based on the control information of the robots. According to the method, dynamic information and static information of various common robots can be uniformly input into the model and converted into control information of the same dimension, so that the robots are correspondingly controlled. By means of the method, unified processing of different types of robots can be achieved, separate training is avoided, the system design and maintenance cost is reduced, and the collaboration of multi-robot operation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of path planning, and in particular, to a unified obstacle avoidance method, a control terminal and a control system for distributed heterogeneous robots. Background Art

[0002] At present, robots have been widely used in fields such as manufacturing, service, military, and agriculture. In some fields, especially in logistics and warehouses, compared with single robots, multi-robot operations can significantly improve the operation efficiency.

[0003] Currently, robots are mainly classified into three types according to the motion type: omnidirectional robots, differential drive robots, and Ackermann-type robots. However, the current obstacle avoidance schemes often control for a single robot type, and it is difficult to achieve unified control among heterogeneous robots. The training cost for separately training different robots is high, the system design is complex, and the maintenance cost is high; directly using the horizontal and vertical speeds or speed changes as the control inputs for different robots ignores the physical characteristics among different robots, resulting in a low success rate.

[0004] In recent years, reinforcement learning methods, especially the reinforcement learning algorithm based on policy gradient (Proximal Policy Optimization, PPO), have achieved success in continuous control tasks due to their stable policy updates and high sample utilization rates. In the current obstacle avoidance strategies, although traditional methods based on safety barrier calculation do not require training, the obstacle avoidance speeds given in the same situation are the same, and due to incomplete consideration, deadlock states are likely to occur; while existing reinforcement learning methods mostly use speed changes as the model output and are mostly applicable to omnidirectional robots or some differential drive robots. Summary of the Invention

[0005] Embodiments of the present invention provide a unified obstacle avoidance method, a control terminal and a control system for distributed heterogeneous robots to solve the problem that the current obstacle avoidance methods are difficult to achieve unified control among heterogeneous robots.

[0006] In a first aspect, embodiments of the present invention provide a unified obstacle avoidance method for distributed heterogeneous robots, including:

[0007] Obtain the dynamic information and static information corresponding to each robot in the heterogeneous robots;

[0008] Input the dynamic information and static information corresponding to each robot into a pre-trained reinforcement learning model to obtain the control information of each robot; wherein, the control information of each robot is in the same dimension;

[0009] Based on the control information of each robot, perform unified obstacle avoidance control on each robot.

[0010] In a possible implementation, the static information corresponding to each robot includes the type, width, maximum speed, and wheelbase of the robot; input the dynamic information and static information corresponding to each robot into a pre-trained reinforcement learning model to obtain the control information for each robot, including:

[0011] Determine the local environment vector corresponding to each robot according to the dynamic information corresponding to each robot;

[0012] Determine the state vector corresponding to each robot according to the dynamic information corresponding to each robot, the type, width, maximum speed, and wheelbase of the robot;

[0013] Input the state vector and local environment vector corresponding to each robot into the pre-trained reinforcement learning model to obtain the control information for each robot.

[0014] In a possible implementation, the dynamic information corresponding to each robot includes the current position, target position, speed information, and environment information corresponding to each robot;

[0015] Input the dynamic information and static information corresponding to each robot into the pre-trained reinforcement learning model to obtain the control information for each robot, including:

[0016] Determine the relative angular error and target distance according to the current position and target position corresponding to each robot;

[0017] Generate a state vector according to the relative angular error, target distance, speed information, and static information corresponding to each robot;

[0018] Generate a local environment vector according to the environment information corresponding to each robot;

[0019] Input the state vector and local environment vector corresponding to each robot into the pre-trained reinforcement learning model to obtain the control information for each robot.

[0020] In a possible implementation, the pre-trained reinforcement learning model is obtained through the following method:

[0021] Obtain the historical dynamic information and historical static information corresponding to each robot in the heterogeneous robots;

[0022] Use the historical dynamic information and historical static information corresponding to each robot as inputs and input them into the policy network and value network for training to obtain the pre-trained reinforcement learning model.

[0023] In a possible implementation, using the historical dynamic information and historical static information corresponding to each robot as inputs and inputting them into the policy network and value network for training to obtain the pre-trained reinforcement learning model, including:

[0024] Take the historical dynamic information and historical static information corresponding to each robot as inputs, and input them into the policy network and the value network to obtain output results;

[0025] Based on the output results, calculate the optimization function value, and calculate the loss function value based on the optimization function value;

[0026] According to the calculated loss function value, adjust the parameters of the policy network and the value network, and return to the step of taking the historical dynamic information and historical static information corresponding to each robot as inputs, inputting them into the policy network and the value network to obtain output results, until the parameters of the policy network and the value network converge, and obtain a pre-trained reinforcement learning model.

[0027] In a possible implementation, the output of the policy network is the recommended speed normalization result, and the recommended speed normalization result is used to determine the dynamic information and static information of each robot at the next moment; the output of the value network is the policy value; based on the output results, calculate the optimization function value, and calculate the loss function value based on the optimization function value, including:

[0028] Based on the policy value and a pre-determined reward function, calculate the optimization function value, and calculate the loss function value based on the optimization function value; wherein, the pre-determined reward function includes the reward function of each robot and the cooperation reward function determined based on the average value of the reward functions of each robot.

[0029] In a possible implementation, determining the relative angular error and the target distance according to the current position and the target position corresponding to each robot includes:

[0030] Determine the global target vector according to the current position and the target position corresponding to each robot;

[0031] Determine the desired heading angle and the target distance according to the global target vector;

[0032] Determine the relative angular error according to the desired heading angle and the current position.

[0033] In a possible implementation, after performing unified obstacle avoidance control on each robot based on the control information of each robot, it further includes:

[0034] Determine the dynamic information of each robot at the next moment based on the speed information of each robot at the next moment;

[0035] And input the dynamic information and static information of each robot at the next moment into the pre-trained reinforcement learning model to obtain the control information of each robot at the next moment.

[0036] Second aspect, an embodiment of the present invention provides a control terminal, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method in the first aspect or any possible implementation manner of the first aspect above is implemented.

[0037] Third aspect, an embodiment of the present invention provides a control system, including the control terminal provided in the second aspect and heterogeneous robots.

[0038] An embodiment of the present invention provides a unified obstacle avoidance method for distributed heterogeneous robots. Since traditional obstacle avoidance schemes often control a single type of robot and it is difficult to achieve unified control among heterogeneous robots, and traditional control methods are based on the dynamic information of each robot for control, which easily leads to deadlock problems. To solve this problem, in this embodiment, the dynamic information and static information of various common types of robots are uniformly input into the model to convert them into control information of the same dimension for corresponding control of each robot. Through this method, unified processing of different types of robots can be achieved, avoiding separate training, reducing the system design and maintenance costs, and improving the coordination of multi-robot operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flowchart of the implementation of the unified obstacle avoidance method for distributed heterogeneous robots provided by an embodiment of the present invention;

[0040] Figure 2 is a schematic diagram of the global coordinate system and local coordinate system of the unified obstacle avoidance method for distributed heterogeneous robots provided by an embodiment of the present invention;

[0041] Figure 3 is a pre-trained reinforcement learning model architecture diagram;

[0042] Figure 4 is a schematic diagram of the structure of the unified obstacle avoidance device for distributed heterogeneous robots provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The embodiments of the present invention will be described in detail below with reference to the drawings.

[0044] Figure 1 is a flowchart of the implementation of the unified obstacle avoidance method for distributed heterogeneous robots provided by an embodiment of the present invention. As Figure 1 shown, the method may include:

[0045] Step 110: Obtain the dynamic information and static information corresponding to each robot in the heterogeneous robots.

[0046] In this embodiment, the heterogeneous robot can refer to a heterogeneous multi-robot group, which includes at least two or more different types of robots. For example, it can include any two or three of the three types of omnidirectional robots, differential drive robots, and Ackermann-type robots.

[0047] In this embodiment, the dynamic information corresponding to each robot may include the current position, target position, speed information, and environmental information; the static information may include the type, width, maximum speed, and wheelbase of the robot.

[0048] Among them, for any one robot, its current position and target position can be obtained through a positioning system such as laser positioning or visual positioning, etc. At time t, its current position can be represented as (x t , y t , θ t ) in the global coordinate system, and the target position is (x goal , y goal ).

[0049] The speed information is different for different types of robots. Exemplarily, the speed information v t = [v t , 0], where the first dimension of these two parameters is the speed magnitude information, and the second dimension is the steering information; that is, they are the speed in the direction the robot is facing and the lateral speed respectively. Since the omnidirectional robot can move in any direction, the lateral speed obtained from the environment is 0.

[0050] For a differential drive robot, since it is usually represented by wheels driven independently on the left and right, its speed information can be: v t = [v l , v r ; where v l is the left wheel speed; v r is the right wheel speed; for unified control, according to the width information width of the differential drive robot, its speed information can be converted into the form of linear speed and angular speed, that is, v t = [v t , ω t ; where the calculation formula is:

[0051]

[0052] An Ackermann-type robot usually has four wheels, with the front wheels for steering and the rear wheels for driving. The rear wheel linear speed and front wheel steering angle obtained from the environment can directly become the speed information, that is, v t = [v t , steer t .

[0053] Each robot can obtain the surrounding environment information through lidar to get the environmental information lidart. Among them, the environmental information is used to perceive environmental conditions such as obstacles. Multiple frames of data are combined into local environmental information, from which position, speed, and acceleration information are extracted.

[0054] For static information, the width is the information that the robot itself has. It is used to consider the robot size when calculating the differential drive robot speed and path planning to prevent collisions.

[0055] The maximum speed information acquisition methods for different types of robots are different. The maximum speed of an omnidirectional robot is v max =[v max ,0]; The maximum speed of a differential drive robot is determined based on the left wheel speed, right wheel speed, and the robot width, expressed as v max =[v max ,ω max ; The maximum speed of an Ackermann-type robot is determined based on the rear wheel linear speed and the front wheel steering angle, expressed as v max =[v max ,steer max . The maximum speed information is used to limit the actual running speed of the robot to ensure safety and compliance with physical characteristics.

[0056] The wheelbase is the unique information of the Ackermann-type robot and is used for kinematic model calculation. In non-Ackermann-type robots, this value is 0 and is used to determine the motion trajectory when the robot steers.

[0057] Step 120: Input the dynamic information and static information corresponding to each robot into a pre-trained reinforcement learning model to obtain the control information of each robot; among them, the control information of each robot is in the same dimension.

[0058] In this embodiment, in the running scenario of heterogeneous robots, to achieve the collaborative operation and efficient obstacle avoidance of various robots, precise and unified control is required. To achieve unified control, the pre-processed dynamic information and static information corresponding to each robot can be input into a pre-trained reinforcement learning model. In this model, by performing dimensionality conversion on the control methods of different types of robots, the control information of different types of robots is in the same dimension.

[0059] Control information of the same dimension makes it possible for different types of robots to perform cooperative control under a unified framework, eliminating the problem of increased control difficulty caused by differences in robot types. For example, in the scenario of a logistics warehouse, omnidirectional robots are responsible for flexibly transporting small goods, differential drive robots perform precise sorting of goods, and Ackermann robots undertake the transportation tasks of large goods. Under the unified control information dimension, they can work together in an orderly manner, avoid colliding with each other, and improve the overall operation efficiency. This unified control method greatly simplifies the design and management of the system, reduces the maintenance cost, and lays a solid foundation for the wide application of distributed heterogeneous robots in multiple fields.

[0060] Step 130: Based on the control information of each robot, perform unified obstacle avoidance control on each robot.

[0061] In a complex multi-robot cooperation scenario, due to the different motion characteristics and physical structures of heterogeneous robots, there are many challenges in achieving unified obstacle avoidance control. However, the embodiments of the present invention ensure the normal operation of each robot and avoid collision with other robots through the control information of each robot.

[0062] Exemplarily, when it is detected that an obstacle appears on the traveling path of a certain robot, the system will quickly analyze based on the control information of the robot. For an omnidirectional robot, the system may adjust its speed and motion direction according to its control information to enable it to flexibly bypass the obstacle; for a differential drive robot, the system will accurately calculate the rotational speed difference between the left and right wheels and control its steering to avoid the obstacle; for an Ackermann robot, the system will comprehensively consider the linear speed of the rear wheels and the steering angle of the front wheels to ensure that it safely bypasses the obstacle while maintaining the stability of the overall motion.

[0063] Unified obstacle avoidance control not only needs to consider the relationship between a single robot and the obstacle, but also coordinate the motions of each robot to avoid interference between them. Exemplarily, the system will formulate a reasonable motion plan according to the control information of each robot to ensure that each robot can efficiently complete its own tasks on the premise of safety. For example, in a logistics warehouse, many different types of robots are simultaneously engaged in goods handling and storage work. Through unified obstacle avoidance control, they can shuttle between the shelves in an orderly manner, avoid collision accidents, and greatly improve the operation efficiency of the warehouse.

[0064] In summary, in the embodiments of the present invention, it is considered that in the actual operating environment, different types of robots such as omnidirectional robots, differential drive robots, and Ackermann robots each undertake specific tasks. Omnidirectional robots can flexibly shuttle in narrow spaces with their unique omnidirectional movement ability; differential drive robots rely on the speed difference between the left and right wheels to achieve precise steering and are suitable for fine operation tasks; Ackermann robots play a key role in handling large objects with their large load capacity and stable steering mechanism. However, when they work together, if there is a lack of effective unified control, collisions are very likely to occur, seriously affecting the operation efficiency and safety.

[0065] To solve this problem, in this embodiment, by converting the control methods of the three common types of robots into control information in the same dimension, omnidirectional, differential drive, and Ackermann robots can work together under a unified policy network, simplifying the system architecture and reducing the maintenance cost; the model training process takes into account the physical limitations of the robots, making the results output by the model conform to the physical limitations and reducing the gap between simulation and actual operation; the reward function with reference speed can better learn the control methods of other algorithms, accelerating the convergence speed of the model and reducing the training time.

[0066] In an optional embodiment, the static information corresponding to each robot includes the type, width, maximum speed, and wheelbase of the robot; in step 120, the dynamic information and static information corresponding to each robot are input into a pre-trained reinforcement learning model to obtain the control information of each robot, which may include:

[0067] According to the dynamic information corresponding to each robot, the local environment vector corresponding to each robot is determined.

[0068] According to the dynamic information corresponding to each robot, the type, width, maximum speed, and wheelbase of the robot, the state vector corresponding to each robot is determined.

[0069] The state vector and local environment vector corresponding to each robot are input into a pre-trained reinforcement learning model to obtain the control information of each robot.

[0070] In this embodiment, since the model adopted is a reinforcement learning model, it is necessary to first convert each data into vector data.

[0071] Considering the characteristics of the data, two types of local environment vectors and state vectors can be determined based on the dynamic information and static information. Among them, the local environment vector can obtain the position information, speed information, and acceleration information of the local environment through the feature extraction of continuous multi-frame data; the state vector includes the dynamic information of each robot and static physical parameters used to characterize static physical properties.

[0072] In this embodiment, the type of the robot can be represented by one-hot encoding. When one_hot = [1, 0, 0], it can represent an omnidirectional robot; when one_hot = [0, 1, 0], it can represent a differential drive robot; when one_hot = [0, 0, 1], it can represent an Ackermann type robot.

[0073] Among them, when the robot is an omnidirectional robot or an Ackermann type robot, its corresponding width can be set to 0; when the robot is not an Ackermann type robot, the wheelbase can be set to 0.

[0074] In an alternative embodiment, the dynamic information corresponding to each robot includes the current position, target position, speed information, and environmental information corresponding to each robot.

[0075] Inputting the dynamic information and static information corresponding to each robot into a pre-trained reinforcement learning model in step 120 to obtain the control information for each robot may include:

[0076] Determine the relative angular error and target distance based on the current position and target position corresponding to each robot.

[0077] Generate a state vector based on the relative angular error, target distance, speed information, and static information corresponding to each robot.

[0078] Generate a local environment vector based on the environmental information corresponding to each robot.

[0079] Input the state vector and local environment vector corresponding to each robot into a pre-trained reinforcement learning model to obtain the control information for each robot.

[0080] In this embodiment, the state vector can be expressed as:

[0081] states t =[d t ,α t ,v t ,v max ,width,one_hot,L]

[0082] Wherein, d t is the target distance; α t is the relative angular error; v t is the speed information; v max is the maximum speed; width is the width; one_hot is used to represent the type of the robot; L is the wheelbase.

[0083] In the formula, the target distance and the relative angular error can be determined based on the current position and the target position.

[0084] The local environment vector can be expressed as:

[0085] lidars t =[lidar t-2 ,lidar t-1 ,lidar t

[0086] where lidar t-2 is the environmental information at time t-2; lidar t-1 is the environmental information at time t-1; lidar t is the environmental information at time t.

[0087] In an alternative embodiment, determining the relative angular error and the target distance according to the current position and the target position corresponding to each robot may include:

[0088] Determining a global target vector according to the current position and the target position corresponding to each robot.

[0089] Determining the desired heading angle and the target distance according to the global target vector.

[0090] Determining the relative angular error according to the desired heading angle and the current position.

[0091] Figure 2 is a schematic diagram of the global coordinate system and the local coordinate system of the unified obstacle avoidance method for distributed heterogeneous robots provided by the embodiments of the present invention. The following will be described in conjunction with Figure 2 to illustrate this embodiment.

[0092] In this embodiment, the current position of each robot is the rectangle in the figure; the target position is the circle in the figure; the current position can be expressed as (x t ,y t ,θ t ); the target position can be expressed as (x goal ,y goal ).

[0093] The global target vector is calculated by the following formula:

[0094] Δx=x goal -x t ,Δy=y goal -y t

[0095] The desired heading angle is determined by the following formula:

[0096] θ d =arctan2(Δy,Δx)

[0097] The target distance is determined by the following formula: ​

[0098]

[0099] The relative angular error is determined by the following formula:

[0100] α t = normalize(θ d - θ t )

[0101] where the normalize function means normalizing the angle to [-π, π].

[0102] In an optional embodiment, the pre-trained reinforcement learning model is obtained in the following manner:

[0103] Obtain the historical dynamic information and historical static information corresponding to each robot in the heterogeneous robot.

[0104] Take the historical dynamic information and historical static information corresponding to each robot as inputs and input them into the policy network and the value network for training to obtain the pre-trained reinforcement learning model.

[0105] Figure 3 is the architecture diagram of the pre-trained reinforcement learning model. The following will be combined with Figure 3 to illustrate this embodiment:

[0106] In this embodiment, the PPO algorithm is selected for model training. In the PPO algorithm, there are usually two networks that are the same except for the output dimension. One is used as the policy network for outputting action information, and the other is used as the value network for evaluating the quality of the policy network.

[0107] In Figure 3 the shown model includes two parts: the CNN module and the fully connected layer. The lidars data of the surrounding environment information undergoes CNN convolution to extract environmental obstacle features. The fully connected layer fuses the CNN features and the dynamic and static information of the robot itself for training, outputs the mean of continuous actions mean, and the final result after normalization is the result sampled from the Gaussian distribution constructed from mean, and the output has a logarithmic standard deviation vector logstd for subsequent training.

[0108] In this embodiment, first, the historical dynamic information and historical static information corresponding to each robot in the heterogeneous robot are processed to obtain the historical local environment vector and historical state vector corresponding to each robot.

[0109] Then, the historical local environment vector and historical state vector corresponding to each robot are input into the reinforcement learning model architecture; among them, the reinforcement learning model architecture includes a policy network and a value network.

[0110] The policy network is used to output action information, and the value network is used to evaluate the advantages and disadvantages of the policy network based on the action information output by the policy network. During the evaluation process, the policy network and the value network itself are adjusted until the preset conditions are met, and a pre-trained reinforcement learning model is obtained.

[0111] In an alternative embodiment, the historical dynamic information and historical static information corresponding to each robot are used as inputs and input into the policy network and the value network for training to obtain a pre-trained reinforcement learning model, including:

[0112] Using the historical dynamic information and historical static information corresponding to each robot as inputs and inputting them into the policy network and the value network to obtain output results;

[0113] Based on the output results, calculate the optimization function value and calculate the loss function value based on the optimization function value;

[0114] According to the calculated loss function value, adjust the parameters of the policy network and the value network, and return to the step of using the historical dynamic information and historical static information corresponding to each robot as inputs and inputting them into the policy network and the value network to obtain output results, until the parameters of the policy network and the value network converge, and a pre-trained reinforcement learning model is obtained.

[0115] In this embodiment, after using the historical dynamic information and historical static information corresponding to each robot as inputs, the policy network and the value network output corresponding output results. In this network architecture, the generalized advantage estimate value can be used to calculate the advantage function value according to the output structure, where the advantage function can be expressed as:

[0116]

[0117] where T i is the maximum elapsed time before termination of the i-th robot; γλ is parameter information; is the evaluation result of the i-th robot.

[0118] Update the parameters of the policy network according to the PPO policy loss function

[0119] where the policy loss function can be expressed as:

[0120]

[0121] where is the ratio of the new and old policy probabilities; ∈ is the clipping hyperparameter.

[0122] Update the parameters φ of the value network according to the PPO value loss function, where the value loss function is:

[0123]

[0124] Among them, N is all the data of the corresponding batch, that is, all the data of all robots at all times; value i is the value obtained at the i-th moment; R i is the reward obtained at the i-th moment.

[0125] Update the policy according to the loss function, and perform training through repeated iteration until the parameters converge, so as to obtain a pre-trained reinforcement learning model.

[0126] In an alternative embodiment, the output of the policy network is the recommended speed normalization result, and the recommended speed normalization result is used to determine the dynamic information and static information of each robot at the next moment; the output of the value network is the policy value; based on the output results, the optimization function value is calculated, and the loss function value is calculated based on the optimization function value, including:

[0127] Calculate the optimization function value based on the policy value and the pre-determined reward function, and calculate the loss function value based on the optimization function value; among them, the pre-determined reward function includes the reward function of each robot and the cooperation reward function determined based on the average value of the reward functions of each robot.

[0128] In this embodiment, the output of the policy network being the recommended speed normalization result can be expressed as The value of each element therein is [-1, 1], and the dynamic information and static information of each robot at the next moment can be determined in the following manner:

[0129] According to the maximum speed of each robot and the recommended speed normalization result, calculate the actual acting speed of each robot, that is, the local speed information:

[0130]

[0131] According to the type characteristics of each robot, determine the global speed based on the local speed information, and then obtain the speed of each robot at the next moment. After obtaining the speed of each robot at the next moment, the global state can also be updated based on the current position of each robot and the speed of each robot at the next moment.

[0132] For an omnidirectional robot, its local speed can be expressed as:

[0133] v t+1 =[v xlocal ,v ylocal

[0134] Among them, v xlocal represents the local lateral speed after decomposition of the omnidirectional robot; v​ylocal Represents the local longitudinal velocity after the omnidirectional robot is decomposed.

[0135] The global velocity of the omnidirectional robot can be expressed as:

[0136]

[0137] Its velocity at the next moment is:

[0138] Correspondingly, its global state is:

[0139]

[0140] For a differential drive robot, its local velocity can be expressed as:

[0141] v t+1 =[v t+1 ,ω t+1

[0142] The global velocity of the omnidirectional robot, that is, the actual linear velocities of the left and right wheels, can be expressed as:

[0143]

[0144] The actual linear velocities of its left and right wheels are the velocities at the next moment.

[0145] Correspondingly, its global state is:

[0146]

[0147] For an Ackermann-type robot, its local velocity can be expressed as:

[0148] v t+1 =[v t+1 ,steer t+1

[0149] Its global state is:

[0150]

[0151] To enable the model to converge quickly, this embodiment proposes a recommended velocity for the velocity matching reward in the reward function. The recommended velocity can be the velocity given by any other algorithm in any sense under the current situation. In this embodiment, the recommended velocity is the velocity to directly reach the target position based on the robot's own position, orientation, and target position without any obstacles.

[0152] For the omnidirectional robot, the recommended control is directly obtained according to the target azimuth and the maximum velocity:

[0153] ​​

[0154] For a differentially driven robot, when the target distance is less than a certain distance and the target is behind the robot, i.e., d t < d th and |α t | > α crit , a backward strategy v ref = -v max is adopted. Otherwise, move forward v ref = v max . The angular velocity information cannot exceed a certain range, i.e., w ref = clip(α t , -ω max , ω max ). That is, the recommended speed is

[0155]

[0156] For an Ackermann-type robot, the linear velocity strategy is the same, and the steering angle is In this case, the robot reaches the target position in an arc, and the recommended speed is

[0157]

[0158] In this embodiment, in order to better learn the control methods of other algorithms, accelerate the convergence speed of the model, and reduce the training time, the pre-determined reward function includes the reward functions of each robot and the cooperative reward function determined based on the mean value of the reward functions of each robot. Among them, the pre-determined reward function can be expressed as:

[0159]

[0160] Among them, represents the self-reward of each robot; η represents the proportion of the cooperative reward among each robot; R coordination represents the cooperative reward among each robot.

[0161] Optionally, the self-reward of each robot can be expressed as:

[0162]

[0163] Among them, α0 is the initial weight, λ is the decay factor, α(t) decreases with the number of training epochs, and βγδ are the weights of different penalties.

[0164] is the speed penalty, which is expressed as:

[0165] R speed = -||v t+1 -v ref || 2

[0166] is the distance reward, which is expressed as:

[0167]

[0168] where the distance reward represents the reduction of the distance between the robot and the target.

[0169] A large negative reward is given when a collision occurs, which is the collision penalty. When a collision occurs it is 0 at other times.

[0170] In this embodiment, to reduce turning operations and increase path smoothness, a penalty is imposed on the turning speed, which is the smoothness penalty and is expressed as:

[0171]

[0172] To encourage the cooperative operation of multiple robots, the average reward of all robots is taken as an additional reward and added to each robot according to a certain proportion. The cooperation reward R coordination can be expressed as:

[0173]

[0174] Based on the given reward function, calculate the reward result corresponding to each policy network each time; and calculate the optimization function value according to the reward result and the policy value value to obtain the final pre-trained reinforcement learning model.

[0175] In an alternative embodiment, after performing unified obstacle avoidance control on each robot based on the control information of each robot, it further includes:

[0176] Based on the speed information of each robot at the next moment, determine the dynamic information of each robot at the next moment;

[0177] And input the dynamic information and static information of each robot at the next moment into the pre-trained reinforcement learning model to obtain the control information of each robot at the next moment.

[0178] In this embodiment, since the dynamic information of each robot at the next moment has been calculated as an intermediate quantity in the process of obtaining the model output, therefore, in practical applications, the known dynamic information of each robot at the next moment can be directly used as the intermediate quantity, and the static information is input into the pre-trained reinforcement learning model to obtain the control information of each robot at the next moment, reducing the computational complexity.

[0179] Of course, in order to ensure the control accuracy each time, it is also possible to re-collect the dynamic information of each robot after performing unified obstacle avoidance control on each robot based on the control information of each robot, so as to perform new control on each robot and ensure that each robot can successfully complete the task.

[0180] In summary, the embodiment of the present invention converts the control methods of three common types of robots into control information of the same dimension, realizes the collaborative work of omnidirectional, differential drive, and Ackermann-type robots under a unified policy network, simplifies the system architecture and reduces the maintenance cost; the model training process considers the physical limitations of the robots, so that the results output by the model conform to the physical limitations, reducing the gap between simulation and actual operation; the reward function with reference speed can better learn the control methods of other algorithms, speeds up the convergence speed of the model, and reduces the training time.

[0181] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0182] The following is an apparatus embodiment of the present invention. For the details not described in detail therein, reference may be made to the corresponding method embodiments above.

[0183] Figure 4 FIG. is a schematic structural diagram of a unified obstacle avoidance apparatus for distributed heterogeneous robots provided by an embodiment of the present invention. For the sake of convenience of description, only the parts related to the embodiment of the present invention are shown and are described in detail as follows:

[0184] As Figure 4 shown, the unified obstacle avoidance apparatus 4 for distributed heterogeneous robots includes:

[0185] An acquisition module 41, configured to acquire the dynamic information and static information corresponding to each robot in the heterogeneous robots;

[0186] A processing module 42, configured to input the dynamic information and static information corresponding to each robot into a pre-trained reinforcement learning model to obtain the control information of each robot; wherein, the control information of each robot is in the same dimension;

[0187] A control module 43, configured to perform unified obstacle avoidance control on each robot based on the control information of each robot.

[0188] In a possible implementation, the static information corresponding to each robot includes the type, width, maximum speed, and wheelbase of the robot; the processing module 42 is specifically configured to:

[0189] Determine the local environment vector corresponding to each robot according to the dynamic information corresponding to each robot;

[0190] Determine the state vector corresponding to each robot according to the dynamic information corresponding to each robot, the type, width, maximum speed, and wheelbase of the robot;

[0191] Input the state vector and local environment vector corresponding to each robot into a pre-trained reinforcement learning model to obtain the control information of each robot.

[0192] In a possible implementation, the dynamic information corresponding to each robot includes the current position, target position, speed information, and environment information corresponding to each robot;

[0193] The processing module 42 is specifically configured to:

[0194] Determine the relative angular error and target distance according to the current position and target position corresponding to each robot;

[0195] Generate a state vector according to the relative angular error, target distance, speed information, and static information corresponding to each robot;

[0196] Generate a local environment vector according to the environment information corresponding to each robot;

[0197] Input the state vector and local environment vector corresponding to each robot into a pre-trained reinforcement learning model to obtain the control information of each robot.

[0198] In a possible implementation, the processing module 42 is specifically configured to:

[0199] Obtain the historical dynamic information and historical static information corresponding to each robot in the heterogeneous robots;

[0200] Use the historical dynamic information and historical static information corresponding to each robot as inputs and input them into a policy network and a value network for training to obtain a pre-trained reinforcement learning model.

[0201] In a possible implementation, the processing module 42 is specifically configured to:

[0202] Use the historical dynamic information and historical static information corresponding to each robot as inputs and input them into a policy network and a value network to obtain an output result;

[0203] Based on the output result, calculate the optimized function value, and calculate the loss function value based on the optimized function value;

[0204] According to the calculated loss function value, adjust the parameters of the policy network and the value network, and return the step of taking the historical dynamic information and historical static information corresponding to each robot as the input, inputting them into the policy network and the value network, and obtaining the output result, until the parameters of the policy network and the value network converge, and obtain the pre-trained reinforcement learning model.

[0205] In a possible implementation manner, the output of the policy network is the recommended speed normalization result, and the recommended speed normalization result is used to determine the dynamic information and static information of each robot at the next moment; the output of the value network is the policy value. The processing module 42 is specifically used for:

[0206] Based on the policy value and the pre-determined reward function, calculate the optimized function value, and calculate the loss function value based on the optimized function value; wherein, the pre-determined reward function includes the reward function of each robot and the cooperation reward function determined based on the average value of the reward functions of each robot.

[0207] In a possible implementation manner, the processing module 42 is specifically used for:

[0208] Determine the global target vector according to the current position and the target position corresponding to each robot;

[0209] Determine the expected heading angle and the target distance according to the global target vector;

[0210] Determine the relative angle error according to the expected heading angle and the current position.

[0211] In a possible implementation manner, after performing unified obstacle avoidance control on each robot based on the control information of each robot,

[0212] The acquisition module 41 is further used to determine the dynamic information of each robot at the next moment based on the speed information of each robot at the next moment;

[0213] The processing module 42 is further used to input the dynamic information and static information of each robot at the next moment into the pre-trained reinforcement learning model, and obtain the control information of each robot at the next moment.

[0214] An embodiment of the present invention further provides a control terminal, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method in the above method embodiment is implemented. Exemplarily, the control terminal may be a desktop computer, a notebook, a palm computer, etc., which is not limited herein.

[0215] An embodiment of the present invention provides a control system, which includes a control terminal and heterogeneous robots.

[0216] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. Without special instructions and logical conflicts, the terms and / or descriptions between different embodiments are consistent and can be mutually referred to. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.

[0217] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A unified obstacle avoidance method for distributed heterogeneous robots, characterized in that, Including: Obtaining the dynamic information and static information corresponding to each robot in the heterogeneous robots; Inputting the dynamic information and static information corresponding to each robot into a pre-trained reinforcement learning model to obtain the control information of each robot; wherein, the control information of each robot is in the same dimension; Based on the control information of each robot, performing unified obstacle avoidance control on each robot.

2. The unified obstacle avoidance method for distributed heterogeneous robots according to claim 1, characterized in that The static information corresponding to each robot includes the type, width, maximum speed, and wheelbase of the robot; the inputting the dynamic information and static information corresponding to each robot into a pre-trained reinforcement learning model to obtain the control information of each robot includes: Determining the local environment vector corresponding to each robot according to the dynamic information corresponding to each robot; Determining the state vector corresponding to each robot according to the dynamic information, type, width, maximum speed, and wheelbase of the robot corresponding to each robot; Inputting the state vector and local environment vector corresponding to each robot into a pre-trained reinforcement learning model to obtain the control information of each robot.

3. The unified obstacle avoidance method for distributed heterogeneous robots according to claim 1, characterized in that, The dynamic information corresponding to each robot includes the current position, target position, speed information, and environment information corresponding to each robot; The inputting the dynamic information and static information corresponding to each robot into a pre-trained reinforcement learning model to obtain the control information of each robot includes: Determining the relative angular error and target distance according to the current position and target position corresponding to each robot; Generating a state vector according to the relative angular error, target distance, speed information, and static information corresponding to each robot; Generating a local environment vector according to the environment information corresponding to each robot; Inputting the state vector and local environment vector corresponding to each robot into a pre-trained reinforcement learning model to obtain the control information of each robot.

4. The unified obstacle avoidance method for distributed heterogeneous robots according to claim 1, characterized in that The pre-trained reinforcement learning model is obtained through the following method: Obtaining the historical dynamic information and historical static information corresponding to each robot in the heterogeneous robots; Taking the historical dynamic information and historical static information corresponding to each robot as inputs and inputting them into a policy network and a value network for training to obtain a pre-trained reinforcement learning model.

5. The unified obstacle avoidance method for distributed heterogeneous robots according to claim 4, wherein, The taking the historical dynamic information and historical static information corresponding to each robot as inputs and inputting them into a policy network and a value network for training to obtain a pre-trained reinforcement learning model includes: Taking the historical dynamic information and historical static information corresponding to each robot as inputs and inputting them into a policy network and a value network to obtain an output result; Based on the output result, calculating an optimization function value and calculating a loss function value based on the optimization function value; According to the calculated loss function value, adjusting the parameters of the policy network and the value network, and returning to the step of taking the historical dynamic information and historical static information corresponding to each robot as inputs and inputting them into a policy network and a value network to obtain an output result until the parameters of the policy network and the value network converge, and then obtaining a pre-trained reinforcement learning model.

6. The unified obstacle avoidance method for distributed heterogeneous robots according to claim 5, characterized in that The output of the policy network is a recommended speed normalization result, and the recommended speed normalization result is used to determine the dynamic information and static information of each robot at the next moment; The output of the value network is the policy value; Based on the output result, calculating an optimization function value and calculating a loss function value based on the optimization function value, including: Calculating an optimization function value based on the policy value and a pre-determined reward function, and calculating a loss function value based on the optimization function value; wherein, the pre-determined reward function includes the reward functions of each robot and a cooperation reward function determined based on the average value of the reward functions of each robot.

7. The unified obstacle avoidance method for distributed heterogeneous robots according to claim 3, characterized in that The determining the relative angular error and the target distance according to the current position and the target position corresponding to each robot includes: Determining a global target vector according to the current position and the target position corresponding to each robot; Determining the desired heading angle and the target distance according to the global target vector; Determining the relative angular error according to the desired heading angle and the current position.

8. The unified obstacle avoidance method for distributed heterogeneous robots according to claim 1, characterized in that After performing unified obstacle avoidance control on each robot based on the control information of each robot, it further includes: Determining the dynamic information of each robot at the next moment based on the speed information of each robot at the next moment; And inputting the dynamic information and the static information of each robot at the next moment into a pre-trained reinforcement learning model to obtain the control information of each robot at the next moment.

9. A control terminal, characterized in that, It includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the method according to any one of claims 1 to 8.

10. A control system, characterized in that, It includes a control terminal according to claim 9, and heterogeneous robots.