Intelligent task supervision method for multi-nonholonomic constraint mobile robot system

By combining a multi-agent reinforcement learning task supervisor and a model predictive control redundant planner, the problems of poor dynamic performance and high computational and storage burden in multi-nonholonomic constraint mobile robot systems are solved, and more efficient behavior priority switching and task execution are achieved.

CN116245286BActive Publication Date: 2025-12-30FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310255761.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-16
Publication Date
2025-12-30
Estimated Expiration
2043-03-16

AI Technical Summary

Technical Problem

Existing zero-space behavior control methods have poor dynamic performance in multi-nonholonomic constraint mobile robot systems, and existing task supervisors rely on manually designed rules, resulting in high computational and storage resource requirements and excessive computational and storage burden.

Method used

By combining a multi-agent reinforcement learning task supervisor and a model predictive control redundant planner, we model the behavior priority switching through cooperative Markov game, learn the optimal joint behavior priority strategy, and design optimization problems under security constraints, thereby reducing the burden of online computation and storage.

Benefits of technology

This improves the dynamic performance of task execution in multi-nonholonomic constraint mobile robot systems, avoids manual rule design, reduces online computing and storage burden, and enhances the reliability and practicality of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116245286B_ABST
    Figure CN116245286B_ABST
Patent Text Reader

Abstract

The application provides an intelligent task supervision method for a multi-nonholonomic constraint mobile robot system, comprising the following steps: step S1, introducing a nonholonomic constraint in a behavior design framework of a null space behavior control method, deducing a behavior design paradigm NSBC-NCs of the null space behavior control with the nonholonomic constraint, and completing basic behavior design of the multi-nonholonomic constraint mobile robot system on the basis of the paradigm, and simultaneously combining the designed basic behaviors into a compound behavior of the multi-nonholonomic constraint mobile robot through a null space projection technology in different priority orders; and through learning an optimal joint behavior priority strategy, guiding the multi-nonholonomic constraint mobile robot system to switch the behavior priority in actual use, and further improving the dynamic performance of a task execution process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent robot technology, and in particular to an intelligent task monitoring method for multi-nonholonomic constraint mobile robot systems. Background Technology

[0002] In recent years, nonholonomically constrained mobile robots have attracted widespread attention due to the prevalence of nonholonomic constraints in real-world systems. Through swarm intelligence techniques, multiple nonholonomically constrained mobile robots can achieve better task performance than individual robots. Because task objectives are often complex, nonholonomically constrained mobile robots must simultaneously perform multiple conflicting tasks. As is well known, multi-task conflict resolution is an open problem in the field of multi-agent systems. Behavioral control provides a practical approach by modeling and fusing behaviors. In particular, a method called Null-Space-based Behavioral Control (NSBC) not only implements the highest-priority behaviors but also executes some lower-priority behaviors in its null space. Generally, determined behavior priorities are needed to complete the null-space projection, and the priority of each individual behavior is pre-determined by a module called a task supervisor.

[0003] Despite this, early work used fixed and pre-set behavior priorities, leading to poor dynamic performance of zero-space behavior control methods. To overcome the shortcomings of fixed priorities, researchers have proposed several task supervisors, such as the Finite State Automaton Mission Supervisor (FSAMS), the Fuzzy Mission Supervisor (FMS), and the Model Predictive Control Mission Supervisor (MPCMS). The FSAMS switches behavior priorities hidden in states by manually designing state transition rules. While the FSAMS is easy to implement, it requires numerical logic rules, which can sometimes be difficult to design. The Fuzzy Logic Mission Supervisor uses fuzzy logic tables instead of numerical logic rules, thus greatly reducing the difficulty of rule design. Similar to the Fuzzy Logic Mission Supervisor, event-triggered techniques offer a potential solution by setting thresholds. Unfortunately, all of the above methods rely on human intelligence to design different types of rules. Therefore, the Model Predictive Control Mission Supervisor describes behavior priority switching as an optimal mode switching problem. By solving for the optimal behavior priority in real time, the Model Predictive Control Mission Supervisor avoids manual rule design. Because the model predictive control task supervisor must resolve the optimal behavior priority for each sampling period, the online computation and storage burden is high, especially when the number of behaviors is relatively large, this problem is further exacerbated. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide an intelligent task supervision method for multi-nonholonomic constraint mobile robot systems. By learning the optimal joint behavior priority strategy, the method guides the multi-nonholonomic constraint mobile robot system to switch behavior priorities in practical use, thereby improving their dynamic performance in the task execution process.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: an intelligent task supervision method for multi-nonholonomic constraint mobile robot systems, comprising the following steps:

[0006] Step S1: Within the behavior design framework of the zero-space behavior control method, nonholonomic constraints are introduced, and the behavior design paradigm NSBC-NCs with nonholonomic constraints zero-space behavior control is derived. Based on this paradigm, the basic behavior design of the multi-nonholonomic constraint mobile robot system is completed. At the same time, through zero-space projection technology, the designed basic behaviors are combined into composite behaviors of the multi-nonholonomic constraint mobile robot in different priority orders.

[0007] Step S2 innovatively models the behavior switching problem of zero-space behavior control as a cooperative Markov game problem, sets the reference velocity command of composite behavior as the action set of reinforcement learning algorithm, selects the position, velocity and potential field value of multi-nonholonomic constraint robot as the state set of reinforcement learning algorithm, and designs a reasonable reward function to construct the multi-agent reinforcement learning task supervisor MARLMS.

[0008] Step S3 involves constructing a cost function for the optimization problem using tracking performance and inertia penalty as indicators, setting a dynamic model and obstacle avoidance as constraints for the optimization problem, thereby designing the Model Predictive Redundancy Planner (MPCRR). In a preferred embodiment, step S1 specifically involves:

[0009] Step S11: Kinematic Modeling of Nonholonomically Constrained Mobile Robot

[0010] Consider a row of N (N>2) nonholonomically constrained mobile robots, where each agent has 2 auxiliary wheels and 2 drive wheels; i = 1,...,N; the linear velocity v of the i-th nonholonomically constrained mobile robot is... i and angular velocity ω i They are respectively represented as

[0011]

[0012]

[0013] in, and These are the speeds of the left and right drive wheels, respectively. It is the distance between the left and right drive wheels. Represents the set of real numbers;

[0014] Define the position and orientation of the i-th nonholonomically constrained mobile robot as follows: and The kinematic equations of the i-th nonholonomically constrained mobile robot are then modeled as follows:

[0015]

[0016] in, These are, respectively, position and velocity in a generalized sense. It is a nonholonomic constraint matrix;

[0017] Assumption 1: The multi-nonholonomic constraint mobile robot system operates in a static scene, where all obstacles are static but not fixed, and some obstacles are unknown;

[0018] Assumption 2: The multi-nonholonomic constraint mobile robot system is a small-scale multi-agent system, for example (N≤20);

[0019] Step S12: Derivation of the null space behavior control paradigm with nonholonomic constraints

[0020] Assume that each nonholonomically constrained mobile robot has M basic behaviors, where the j-th basic behavior of the i-th nonholonomically constrained mobile robot can be represented by a task variable. For j = 1, ..., M, the mathematical model is as follows:

[0021]

[0022] Among them, h i,j (·): For task functions;

[0023] Then, task variables The differential form is expressed as

[0024]

[0025] in, The Jacobian matrix representing the task;

[0026] Finally, the reference velocity command for the j-th basic behavior of the i-th nonholonomically constrained mobile robot is expressed as:

[0027]

[0028] in, It is J i,j The right pseudo-inverse matrix, It is the expected task. It is a bonus to the task. It is the error in the task;

[0029] Step S13: Design of basic behaviors

[0030] Without loss of generality, the formation maintenance, formation reconfiguration, and obstacle avoidance behaviors are designed as follows:

[0031] Formation Behavior FM:

[0032] Formation-keeping behavior aims to drive a multi-nonholonomic constrained mobile robot system to form and maintain a desired formation. The corresponding task function, desired task, and task Jacobian matrix are expressed as follows:

[0033]

[0034]

[0035]

[0036] in, It is the position of the i-th nonholonomic constraint mobile robot. This is the location of the first nonholonomically constrained mobile robot. This is the expected formation task function for the first nonholonomically constrained mobile robot. It represents the relative formation positions of the first nonholonomic constraint mobile robot and the i-th nonholonomic constraint mobile robot. It is the expected direction of formation behavior; formation reconfiguration behavior FR:

[0037] The formation reconfiguration behavior aims to drive a multi-nonholonomic constrained mobile robot system to reconfigure a desired formation. The corresponding task function, desired task, and task Jacobian matrix are expressed as follows:

[0038]

[0039]

[0040]

[0041] in, It is the expected reconstruction task function of the first nonholonomically constrained mobile robot;

[0042] Obstacle Avoidance Behavior (OA):

[0043] Obstacle avoidance behavior aims to drive a multi-nonholonomic constraint mobile robot system to avoid obstacles near its path. The corresponding task function, expected task, and task Jacobian matrix are expressed as follows:

[0044]

[0045]

[0046]

[0047] in, d is the minimum distance between the i-th nonholonomically constrained mobile robot and the obstacle. OA For a safe distance, It is the relative position with the minimum distance. is the expected direction of obstacle avoidance behavior, and + and - represent the left and right sides of the obstacle in the i-th nonholonomic constraint mobile robot, respectively;

[0048] Step S14: Design of Composite Behaviors

[0049] A composite task is a combination of multiple basic behaviors arranged in a certain priority order; setting Let j be the task function of the i-th nonholonomically constrained mobile robot, where j m ∈NM N M ={1,...,M}, m j M represents the dimension of the task space, and M represents the number of tasks; define a time-dependent priority function g. i (j m ,t):N M ×[0,∞]→N M Simultaneously, define a task hierarchy with the following rules:

[0050] 1) A having g i (j α Priority task j α Cannot interfere with g i (j β Priority task j β If g i (j α )≥g i (j β ),

[0051] 2) The mapping relationship from speed to task speed is given by the task's Jacobian matrix. express;

[0052] 3) Has the lowest priority task m M The dimension may be greater than Therefore, it is necessary to ensure that dimension m n Greater than the total dimension of all tasks;

[0053] 4)g i (j m The value of ) is assigned by the task supervisor based on the task requirements and sensor information;

[0054] By assigning a certain priority to the basic tasks, the velocity of the composite task at time t is expressed as:

[0055]

[0056]

[0057]

[0058] in, It is a behavior priority. It is the augmented Jacobian matrix of the null spatial projection.

[0059] In a preferred embodiment, step S2 specifically comprises:

[0060] The problem of switching behavior priorities is described as a cooperative Markov game, where all non-holonomically constrained mobile robots contribute a team's reward; the joint set of states and joint sets of behaviors are defined as S = {s...} t} and B = {b t},in It is a joint position. It is a flag indicating the priority of joint actions. It is the formation marker position. It is the combined potential field value. The calculation is as follows:

[0061]

[0062]

[0063]

[0064] Where Z is the number of obstacle sampling points, and exp(·) represents the exponential function. Indicates in d OA Density of internal obstacles ω represents the distance between the nonholonomically constrained mobile robot and the obstacle sampling point. o ≥1 is a slack variable; a potential field is generated by sampling along the surface of the obstacle at half a safe distance. Furthermore, The reward function is designed as follows:

[0065] r t =r1+r2, (22)

[0066]

[0067]

[0068] in, These represent the identifiers for no formation, reconfigured formation, and desired formation states, respectively; r1 and r2 are the reward signals for achieving the task objective and reducing behavior switching, respectively.

[0069] A multi-nonholonomic constrained mobile robot system interacts with its environment at time step t, and they observe a joint state s. t Based on a Strategy selection joint behavior b t Receive a team reward r t And transition to the next joint state s t+1 ; This refers to a multi-nonholonomic constraint mobile robot system with one The probability of selecting a random joint action b t and with a The probability of selecting the joint action with the largest Q value ζ is an exponent; then, the experience is stored in the experience pool and labeled with a loose value as follows:

[0070]

[0071] T t+1 (φ(s t ),b t )=γ l T t (φ(s t ),b t ), (26)

[0072]

[0073] Among them, κ l It is the moderation factor of the easing value, T t It is the decay temperature, φ(·) is the hash autoencoder function, and γ l It is the discount factor, d γ It is the attenuation rate;

[0074] Since excessively high Q-values ​​always impair accurate learning, a Dueling network structure and an average Q-value are introduced to optimize the network structure; therefore, Q-value updates are based on a relaxation value l. t Calculation as follows

[0075]

[0076] Where, α t ∈(0,1) is the learning rate, and χ~U(0,1) represents a random variable. It is a timing difference error. The offline training of the multi-agent reinforcement learning task supervisor stops after all rounds; finally, the learned joint policy guides the multi-nonholonomically constrained mobile robot to select the optimal joint behavior priority in the real-world scenario; note that once the joint behavior priority is determined, the reference velocity v... r,i and reference trajectory x i,r Calculations are performed according to formulas (3) and (16)-(18).

[0077] In a preferred embodiment, step S3 specifically comprises:

[0078] Note that the multi-agent reinforcement learning task supervisor and model predictive control must use the same kinematic model (3) and the same sampling time Δt; then, the cost function of the i-th nonholonomic constrained mobile robot is designed as follows:

[0079]

[0080] in, Let k be the tracking error from time step k to time step k+κ. γ is a positive definite matrix. C ≥0 is the penalty factor, γ v Let ||·|| be the velocity factor, and ||·|| be the Euclidean norm.

[0081] Therefore, a set of discrete, fixed-step-size, finite-dimensional nonlinear optimization problems are described as follows:

[0082]

[0083] in, and Let be the initial position and initial velocity of the i-th nonholonomic constraint mobile robot, respectively; k+κ|k represents the time step from k to k+κ. The nonholonomic constraint matrix for the model predictive control redundant planner;

[0084] This rolling optimization problem can be solved using the MATMPC nonlinear solver toolbox; if the safety constraints are not violated, then and The inertia penalty term is so small that it can be ignored; otherwise, the model predictive control redundant planner would replan the reference instructions. The inertia penalty term is set to prevent multi-nonholonomic constraint mobile robot systems from remaining in some minimum extreme states.

[0085] Compared with existing technologies, this invention has the following advantages: Taking multiholonomic constraint mobile robot systems as the research object, this invention proposes an intelligent task supervisor design method for multi-nonholonomic constraint mobile robot systems, addressing the task supervisor design problem based on the zero-space behavior control method. First, this method presents for the first time a zero-space behavior control paradigm with nonholonomic constraints. The proposed zero-space behavior control framework with nonholonomic constraints ensures that the reference velocity command after zero-space projection does not violate the nonholonomic constraints. Compared with classical zero-space behavior control methods, the proposed zero-space behavior control with nonholonomic constraints exhibits stronger robustness to some extreme value problems. Second, by modeling behavior priority switching as a Markov game, a multi-agent reinforcement learning task supervisor is designed to learn the optimal joint behavior priority strategy. Compared with existing task supervisors, the multi-agent reinforcement learning task supervisor not only avoids relying on human intelligence to design behavior priority switching rules, but also... Shifting online computation to the offline stage reduces the computational and storage burden on the online system. Finally, a model predictive control redundant planner was designed by constructing a rolling optimization problem with safety constraints. This planner typically executes the learning strategy of the multi-agent reinforcement learning task supervisor and only re-plans the reference instructions when the safety constraints are violated. Since the multi-agent reinforcement learning task supervisor and the model predictive control redundant planner form an integrated decision-making and control framework, it overcomes the challenges caused by the inconsistency between the learning environment and the actual environment. This is the key to enabling multi-agent reinforcement learning algorithms to be applied to practical nonholonomically constrained mobile robots. This structure is called the intelligent task supervisor, which has strong reliability and practicality. Attached Figure Description

[0086] Figure 1 This is a principle block diagram of a design method for an intelligent task supervisor for a multi-nonholonomic constraint mobile robot system according to an embodiment of the present invention.

[0087] Figure 2 This is a schematic diagram of the i-th nonholonomic constraint mobile robot according to an embodiment of the present invention;

[0088] Figure 3 This is a pseudocode diagram of a multi-agent reinforcement learning task supervisor according to an embodiment of the present invention;

[0089] Figure 4 This is a schematic diagram of obstacle sampling according to an embodiment of the present invention, (a) obstacle sampling point, (b) potential field value;

[0090] Figure 5 This is a schematic diagram of the model predictive control redundant planner according to an embodiment of the present invention;

[0091] Figure 6 This is a diagram showing the selection of simulation parameter values ​​in an embodiment of the present invention;

[0092] Figure 7 The potential field diagrams of the environment in this embodiment of the invention are shown in (a) the learning environment and (b) the time working environment.

[0093] Figure 8 The following is a comparison of the task performance of the traditional zero-space behavior control method and the behavior control method with nonholonomic constraints in this embodiment of the invention: (a) trajectory, (b) direction.

[0094] Figure 9 The following is a comparison of the task performance of the finite state machine task supervisor and the proposed intelligent task supervisor in this invention: (a) trajectory, (b) direction, (c) distance between the nonholonomically constrained mobile robot and obstacles, and (d) behavior priority.

[0095] Figure 10 A comparison chart of the online iteration time of the model predictive control task supervisor and the proposed intelligent task supervisor in this invention embodiment;

[0096] Figure 11 The following is a comparison chart of the task performance of the intelligent task supervisor with and without the model predictive control redundant planner according to an embodiment of the present invention: (a) behavior priority, (b) online iteration time, (c) trajectory, and (d) distance between the nonholonomic constraint mobile robot and obstacles.

[0097] Figure 12 This is a schematic diagram of the experimental configuration of an embodiment of the present invention, (a) AgileX Limo, (b) experimental apparatus;

[0098] Figure 13 These are snapshots of experiments on multiple nonholonomically constrained mobile robots with the proposed intelligent task supervisor according to embodiments of the present invention: (a) 0 seconds, (b) 39 seconds, (c) 53 seconds, and (d) 79 seconds.

[0099] Figure 14 The following is a diagram showing the task performance results of a traditional behavior control method with an intelligent task supervisor according to an embodiment of the present invention: (a) trajectory, (b) direction, (c) distance between the nonholonomically constrained mobile robot and obstacles, and (d) behavior priority.

[0100] Figure 15 The following is a task performance result diagram of the nonholonomic constraint behavior control method with finite state machine task supervisor according to an embodiment of the present invention: (a) trajectory, (b) direction, (c) distance between the nonholonomic constraint mobile robot and obstacles, and (d) behavior priority.

[0101] Figure 16The following is a diagram showing the task performance results of the nonholonomic constraint behavior control method with intelligent task supervisor according to an embodiment of the present invention: (a) trajectory, (b) direction, (c) distance between the nonholonomic constraint mobile robot and obstacles, and (d) behavior priority. Detailed Implementation

[0102] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0103] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0104] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0105] Step 1: Behavioral Design with Nonholonomic Constraints and Zero-Space Behavioral Control

[0106] Kinematic modeling of nonholonomically constrained mobile robots

[0107] Consider a row of N (N>2) nonholonomically constrained mobile robots, where each agent has 2 auxiliary wheels and 2 drive wheels. A schematic diagram of the i-th nonholonomically constrained mobile robot is shown below. Figure 2 As shown, i = 1,...,N. The linear velocity v of the i-th nonholonomically constrained mobile robot. i and angular velocity ω i They can be represented as

[0108]

[0109]

[0110] in, and These are the speeds of the left and right drive wheels, respectively. It is the distance between the left and right drive wheels. It represents the set of real numbers.

[0111] Define the position and orientation of the i-th nonholonomically constrained mobile robot as follows: and Then the kinematic equations of the i-th nonholonomically constrained mobile robot can be modeled as follows:

[0112]

[0113] in, These are, respectively, position and velocity in a generalized sense. It is a nonholonomic constraint matrix.

[0114] Assumption 1: The multi-nonholonomic constraint mobile robot system operates in a static scene where all obstacles are static but not fixed, and some obstacles are unknown.

[0115] Assumption 2: The multi-nonholonomic constraint mobile robot system is a small-scale multi-agent system, for example (N≤20).

[0116] Derivation of the Null Space Behavior Control Paradigm with Nonholonomic Constraints

[0117] Assume that each nonholonomically constrained mobile robot has M basic behaviors, where the j-th basic behavior of the i-th nonholonomically constrained mobile robot can be represented by a task variable. For j = 1, ..., M, the mathematical model is as follows:

[0118]

[0119] Among them, h i,j (·): This is the task function.

[0120] Then, task variables The differential form can be expressed as

[0121]

[0122] in, The Jacobian matrix represents the task.

[0123] Finally, the reference velocity command for the j-th basic behavior of the i-th nonholonomically constrained mobile robot can be expressed as:

[0124]

[0125] in, It is J i,j The right pseudo-inverse matrix, It is the expected task. It is a bonus to the task. It is the error in the task.

[0126] Design of basic behaviors

[0127] Without loss of generality, the formation maintenance, formation reconfiguration, and obstacle avoidance behaviors are designed as follows:

[0128] Formation Maintenance (FM):

[0129] Formation-keeping behavior aims to drive a multi-nonholonomic constrained mobile robot system to form and maintain a desired formation. The corresponding task function, desired task, and task Jacobian matrix are expressed as follows:

[0130]

[0131]

[0132]

[0133] in, It is the position of the i-th nonholonomic constraint mobile robot. This is the location of the first nonholonomically constrained mobile robot. This is the expected formation task function for the first nonholonomically constrained mobile robot. It represents the relative formation positions of the first nonholonomic constraint mobile robot and the i-th nonholonomic constraint mobile robot. This is the expected direction of the formation behavior. Formation Reconstruction (FR):

[0134] The formation reconfiguration behavior aims to drive a multi-nonholonomic constrained mobile robot system to reconfigure a desired formation. The corresponding task function, desired task, and task Jacobian matrix are expressed as follows:

[0135]

[0136]

[0137]

[0138] in, It is the expected reconstruction task function of the first nonholonomically constrained mobile robot.

[0139] Obstacle Avoidance (OA):

[0140] Obstacle avoidance behavior aims to drive a multi-nonholonomic constraint mobile robot system to avoid obstacles near its path. The corresponding task function, expected task, and task Jacobian matrix are expressed as follows:

[0141]

[0142]

[0143]

[0144] in, d is the minimum distance between the i-th nonholonomically constrained mobile robot and the obstacle. OA For a safe distance, It is the relative position with the minimum distance. is the expected direction of obstacle avoidance behavior, and + and - represent the left and right sides of the obstacle on the i-th nonholonomic constraint mobile robot, respectively.

[0145] Design of composite behaviors

[0146] A composite task is a combination of multiple basic behaviors arranged in a certain priority order. (Settings) Let j be the task function of the i-th nonholonomically constrained mobile robot, where j m ∈N M N M ={1,...,M}, m j Let M represent the dimension of the task space and M represent the number of tasks. Define a time-dependent priority function g. i (j m ,t):N M ×[0,∞]→N M At the same time, define a task hierarchy with the following rules:

[0147] 1) A having g i (j α Priority task j α Cannot interfere with g i (j β Priority task j β If g i (j α )≥g i (j β ),

[0148] 2) The mapping relationship from speed to task speed is given by the task's Jacobian matrix. express.

[0149] 3) Has the lowest priority task m M The dimension may be greater than Therefore, it is necessary to ensure that dimension m n It is greater than the total dimension of all tasks.

[0150] 4)g i (j m The value of ) is assigned by the task supervisor based on the task requirements and sensor information.

[0151] By assigning a certain priority to the basic tasks, the velocity of the composite task at time t can be expressed as:

[0152]

[0153]

[0154]

[0155] in, It is a behavior priority. It is the augmented Jacobian matrix of the null spatial projection.

[0156] Step 2: Design of a Multi-Agent Reinforcement Learning Task Supervisor

[0157] The problem of switching action priorities can be described as a cooperative Markov game, where all non-integrity constrained mobile robots contribute to a team's reward. The pseudocode diagram of a multi-agent reinforcement learning task supervisor is shown below. Figure 3 As shown. Specifically, the set of states and the set of behaviors of the union are defined as S = {s...} t} and B = {b t},in It is a joint position. It is a flag indicating the priority of joint actions. It is the formation marker position. It is the combined potential field value. The following can be calculated:

[0158]

[0159]

[0160]

[0161] Where Z is the number of obstacle sampling points, and exp(·) represents the exponential function. Indicates in d OA Density of internal obstacles ω represents the distance between the nonholonomically constrained mobile robot and the obstacle sampling point. o ≥1 represents a slack variable. A potential field can be generated by sampling along the surface of the obstacle at half a safe distance, such as... Figure 4 As shown. Furthermore, The reward function is designed as follows:

[0162] r t =r1+r2, (22)

[0163]

[0164]

[0165] in, These represent the identifiers for no formation, reconfigured formation, and desired formation states, respectively. r1 and r2 are the reward signals for achieving the task objective and reducing behavior switching, respectively.

[0166] A multi-nonholonomic constrained mobile robot system interacts with its environment at time step t, and they observe a joint state s. t Based on a Strategy selection joint behavior b t Receive a team reward r t And transition to the next joint state s t+1 . This refers to a multi-nonholonomic constraint mobile robot system with one The probability of selecting a random joint action b t and with a The probability of selecting the joint action with the largest Q value ζ is an exponent. Then, the experience is stored in the experience pool and labeled with a loose value as follows:

[0167]

[0168] T t+1 (φ(s t ),b t )=γ l T t (φ(s t ),b t ), (26)

[0169]

[0170] Among them, κ l It is the moderation factor of the leniency value, T t It is the decay temperature, φ(·) is the hash autoencoder function, and γ l It is the discount factor, d γ It is the attenuation rate.

[0171] Since excessively high Q-values ​​always impair accurate learning, a Dueling network structure and an average Q-value are introduced to optimize the network structure. Therefore, Q-value updates are based on a relaxation value l. t Calculation as follows

[0172]

[0173] Where, α t ∈(0,1) is the learning rate, and χ~U(0,1) represents a random variable. It is a timing difference error.

[0174] The offline training of the multi-agent reinforcement learning task supervisor stops after all rounds. Finally, the learned joint policy guides the multi-nonholonomically constrained mobile robot to select the optimal joint behavior priority in the real-world scenario. Note that once the joint behavior priority is determined, the reference velocity v... r,i and reference trajectory x i,r It can be calculated according to formulas (3) and (16)-(18).

[0175] Step 3: Design of the Model Predictive Control Redundant Planner

[0176] The Model Predictive Control Redundant Planner aims to improve the reliability of supervisors for multi-agent reinforcement learning tasks. Its schematic diagram is shown below. Figure 5 As shown. Note that the multi-agent reinforcement learning task supervisor and model predictive control must use the same kinematic model (3) and the same sampling time Δt. Then, the cost function of the i-th nonholonomic constrained mobile robot is designed as follows:

[0177]

[0178] in, Let k be the tracking error from time step k to time step k+κ. γ is a positive definite matrix. C ≥0 is the penalty factor, γ v Let be the velocity factor, and ||·|| be the Euclidean norm.

[0179] Therefore, a set of discrete, fixed-step-size, finite-dimensional nonlinear optimization problems can be described as follows:

[0180]

[0181]

[0182]

[0183]

[0184]

[0185] in, and Let be the initial position and initial velocity of the i-th nonholonomically constrained mobile robot, respectively. Let k+κ|k represent the time step from k to k+κ. This is the nonholonomic constraint matrix for the model predictive control redundant planner.

[0186] This rolling optimization problem can be solved using the MATMPC nonlinear solver toolbox. If the safety constraints are not violated, then... and The inertia penalty term is so small that it can be ignored. Otherwise, the model predictive control redundant planner would replan the reference instructions. The inertia penalty term is set to prevent the multi-nonholonomic constraint mobile robot system from remaining in some minimum extreme state.

[0187] To provide a detailed introduction to this invention, a simulation and an experimental example are given below to demonstrate the effectiveness and superiority of the proposed intelligent task monitor design method for multi-nonholonomic constraint mobile robot systems.

[0188] 1. Simulation Comparison and Analysis

[0189] The numerical simulation considers three multi-nonholonomic mobile robot systems moving to a target location while avoiding obstacles along the path. It is assumed that the multi-nonholonomic mobile robot systems are capable of environmental detection. All parameters used in the simulation are as follows: Figure 6 As shown, obstacle 6 is an unknown obstacle, and the position of obstacle 1 is... The location of obstacle 2 is The location of obstacle 3 is The location of obstacle 4 is The location of obstacle 5 is The location of obstacle 6 is Potential field diagrams of learning and actual work environments are as follows Figure 7 As shown, and the simulation results are as follows. Figure 8-11 As shown. Figure 8 A comparison was made between traditional zero-space behavior control methods and zero-space behavior control methods with nonholonomic constraints. The traditional method caused the nonholonomically constrained mobile robot system to enter a minimum extreme state, while the method with nonholonomic constraints did not exhibit this phenomenon. This is because the traditional method uses a mass kinematics model, while the method with nonholonomic constraints uses a nonholonomic constraint kinematics model. Therefore, the method with nonholonomic constraints is more robust to extreme value problems and is more suitable for nonholonomically constrained mobile robot systems. Figure 9-10 The task performance of the proposed intelligent task supervisor, finite state machine task supervisor, and model predictive control task supervisor was compared. Compared with the finite state machine task supervisor, the intelligent task supervisor has a lower frequency of behavior priority switching and does not violate safety constraints. Compared with the model predictive control task supervisor, the intelligent task supervisor has less online iteration time and better real-time performance. Figure 11A comparison was made between the intelligent task supervisor with and without a model predictive control redundant planner, where obstacle 6 was reconfigured. Intelligent task supervisors without a model predictive control redundant planner exhibit a few violations of safety constraints, while those with a model predictive control redundant planner strictly adhere to these constraints. Therefore, the model predictive control redundant planner enhances the reliability and practicality of intelligent task supervisors. All comparisons demonstrate the effectiveness and superiority of the proposed intelligent task supervisor method.

[0190] 2. Experimental Comparison and Analysis

[0191] The physical experiment also considered three multi-nonholonomic constraint mobile robot systems moving to the target position while avoiding obstacles in the path. A schematic diagram of the experimental setup is shown below. Figure 12 As shown, the AgileX Limos (ALs) robot is set to differential mode, making it a typical nonholonomic constrained mobile robot system. During task execution, the ALs robot uses LiDAR to detect obstacles. The experimental comparison results are as follows... Figure 13-16 As shown. Figure 13 A snapshot is shown of a multi-nonholonomic constraint mobile robot performing a task, equipped with an intelligent task supervisor. Figure 14-16 The graph shows a comparison of task performance in a real-world experimental environment, comparing the traditional behavior control method with an intelligent task supervisor, the nonholonomic constraint behavior control method with a finite state machine task supervisor, and the nonholonomic constraint behavior control method with an intelligent task supervisor. The traditional behavior control method with an intelligent task supervisor caused the ALs robot to enter a minimum extreme state, failing to reach the target position. The nonholonomic constraint behavior control method with a finite state machine task supervisor exhibited unsatisfactory task switching performance, leading to violations of safety constraints. The nonholonomic constraint behavior control method with an intelligent task supervisor not only enabled the ALs robot to reach the target position but also avoided any violations of safety constraints. All comparisons demonstrate the reliability and practicality of the proposed intelligent task supervisor method.

Claims

1. An intelligent task supervision method for a multi-nonholonomic constraint mobile robot system, characterized by: The method comprises the following steps: Step S1, in the behavior design framework of the null-space behavior control method, nonholonomic constraints are introduced, a behavior design paradigm NSBC-NCs of the null-space behavior control with nonholonomic constraints is derived, and basic behaviors of the multi-nonholonomic constraint mobile robot system are designed on the basis of the paradigm, and the designed basic behaviors are combined into compound behaviors of the multi-nonholonomic constraint mobile robot in different priority orders through the null-space projection technology; Step S2, the behavior switching problem of the null-space behavior control is modeled as a cooperative Markov game problem, the reference speed command of the compound behavior is set as an action set of a reinforcement learning algorithm, the position, the speed and the potential field value of the multi-nonholonomic constraint robot are selected as a state set of the reinforcement learning algorithm, a reward function is designed, and thus a multi-agent reinforcement learning task supervisor MARLMS is constructed; Step S3, a cost function of an optimization problem is constructed with the tracking performance and the inertia penalty as indexes, a dynamic model and obstacle avoidance are set as constraints of the optimization problem, and thus a model predictive redundant planner MPCRR is designed; The step S3 is specifically: Note that the multi-agent reinforcement learning task supervisor and model predictive control must use the same kinematic model and the same sampling time ; then, the cost function of the first nonholonomic mobile robot is designed as follows: (29) wherein is a tracking error from time step to time step, is a positive definite matrix, is a penalty factor, is a speed factor, is the Euclidean norm; Therefore, a set of discrete fixed-step finite-dimensional nonlinear optimization problems are described as (30) wherein, and are the initial position and initial velocity of the first nonholonomic mobile robot, respectively; denotes the time step from to , and is the nonholonomic constraint matrix of the model predictive control redundancy planner. The rolling optimization problem is solved by the MATMPC nonlinear solver toolbox; if the safety constraints are not violated, then and because the inertial penalty term is small to be ignored; otherwise, the model predictive control redundancy planner re-plans the reference command; the inertial penalty term is set to avoid the multi-incomplete constraint mobile robot system to keep in some minimum extreme states.

2. The intelligent task monitoring method for multi-nonholonomic constraint mobile robot system according to claim 1, wherein, The step S1 is specifically: Step S11: kinematic modeling of the nonholonomic constraint mobile robot Consider a row A nonholonomic constraint mobile robot, in which each agent has 2 auxiliary wheels and 2 drive wheels; ;No. The linear velocity of a nonholonomically constrained mobile robot and angular velocity They are respectively represented as (1) (2) wherein, and are the velocities of the left and right drive wheels, respectively, is the distance between the left and right drive wheels, denotes the set of real numbers; Definition of the position and orientation of the first nonholonomic mobile robot are and respectively, then the kinematic equation of the first nonholonomic mobile robot is modeled as (3) wherein, , are generalized position and velocity, respectively, is a nonholonomic constraint matrix; Hypothesis 1: the multi-nonholonomic constraint mobile robot system works in a static scene, in which all obstacles are static but not fixed, and part of the obstacles are unknown; Hypothesis 2: the multi-nonholonomic constraint mobile robot system is a small-scale multi-agent system; Step S12: derivation of the null-space behavior control paradigm with nonholonomic constraints Assume that each nonholonomic mobile robot has basic behaviors, where the basic behavior of the nonholonomic mobile robot can use a task variable , , and be mathematically modeled as follows (4) wherein, is a task function; Then, the task variable is expressed as a differential form (5) wherein, J represents the Jacobian matrix of the task; Finally, the reference velocity command of the first basic behavior of the nonholonomic mobile robot is expressed as ​ (6) wherein is the right pseudo-inverse matrix of is the desired task, is the gain of the task, is the error of the task; Step S13: design of basic behaviors Without loss of generality, the formation keeping, formation reconstruction and obstacle avoidance behaviors are designed as follows: Formation keeping behavior FM: The formation keeping behavior aims to drive the multi-nonholonomic constraint mobile robot system to form and maintain a desired formation, and the corresponding task function, the desired task and the task Jacobian matrix are respectively represented as: (7) (8) (9) wherein, is the position of the th non-holonomic constrained mobile robot, is the position of the 1st non-holonomic constrained mobile robot, is the desired formation task function of the 1st non-holonomic constrained mobile robot, is the formation relative position between the 1st non-holonomic constrained mobile robot and the th non-holonomic constrained mobile robot, is the direction of the formation behavior expectation; Formation reconstruction behavior FR: The formation reconstruction behavior aims to drive the multi-nonholonomic constraint mobile robot system to reconstruct a desired formation, and the corresponding task function, the desired task and the task Jacobian matrix are respectively represented as: (10) (11) (12) wherein, is the desired reconstruction task function of the first nonholonomic mobile robot; Obstacle avoidance behavior OA: The obstacle avoidance behavior aims to drive the multi-nonholonomic constraint mobile robot system to avoid obstacles near the path, and the corresponding task function, the desired task and the task Jacobian matrix are respectively represented as: (13) (14) (15) wherein, is the minimum distance of the nonholonomic mobile robot to the obstacle, is the safety distance, , is the relative position of the minimum distance, is the desired direction of the obstacle avoidance behavior, and denote that the obstacle is on the left and right side of the nonholonomic mobile robot, respectively. Step S14: design of compound behaviors A composite task is a combination of multiple primitive behaviors in a certain priority order; set For the first Nonholonomic mobile robot task function, where , , Indicates the dimension of the task space, Indicates the number of tasks; define the time-dependent priority function ; At the same time, define a task hierarchy with the following rules: 1) a task having priority cannot interfere with a task having priority if , , ; 2) the mapping from velocity to task velocity is represented by the Jacobian matrix of the task ; 3) the task with the lowest priority The dimensionality of the task can be greater than Thus, make sure the dimensionality is greater than the total dimensionality of all tasks; 4) The values are assigned by the task supervisor according to the task's requirements and sensor information; By assigning a given priority to the basic tasks, The speed of the complex task at the moment is expressed as (16) (17) (18) wherein is the behavior priority, is the augmented Jacobian matrix of the null-space projection.

3. The intelligent task monitoring method for multi-nonholonomic constraint mobile robot system of claim 1, wherein, The step S2 is specifically: The switching problem of behavior priority is described as a cooperative Markov game, where multiple nonholonomic mobile robots contribute to a team's reward; the joint state and joint action set are defined as and where , is the joint position, is the joint behavior priority flag, is the formation flag, is the joint potential field value, is calculated as follows: (19) (20) (21) where, is the number of obstacle sampling points, denotes the exponential function, denotes the density of obstacles within denotes the distance between the nonholonomic mobile robot and the obstacle sampling points, is the slack variable; along the surface of the obstacle, sample the potential field with half of the safety distance, in addition, , and the reward function is designed as follows​ (22) (23) (24) wherein, are reward signals respectively set to achieve the task goal and reduce behavior switching.​​ Multi-nonholonomic mobile robot system interacts with the environment at a time step, they observe the joint state , based on a policy, they select a joint action , get a team reward , and transition to the next joint state ; means that the multi-nonholonomic mobile robot system selects a random joint action with a probability of , and selects the joint action with the largest value with a probability of , is an index; then, the episode is stored in the experience pool and labeled with a loose value as follows (25) (26) (27) wherein, is a modest factor of leniency, is a decay temperature, is a hash autoencoding function, is a discount factor, is a decay rate; The Dueling network structure and average values are introduced to optimize the network structure; therefore, value updates are based on loose values are computed as follows (28) wherein, is a learning rate, denotes a random variable, is a time-difference error, ; The offline training of the multi-agent reinforcement learning task supervisor stops at the end of all episodes; finally, the learned joint policy guides the multi-nonholonomic constraint mobile robot to select the optimal joint behavior priority in the actual scene; note that when the joint behavior priority is determined, the reference velocity and the reference trajectory are calculated according to formulas (3) and (16)-(18).