Local planning method of mobile robots based on MPC dynamic game in crowd environment

By establishing an MPC dynamic game model and identifying pedestrian weight parameters, the problem of freezing and obstacle avoidance of local path planning of mobile robots in dynamic crowd environments is solved, and the effect of rapid reaching the target point and energy saving and avoidance is achieved.

CN116185005BActive Publication Date: 2025-08-19SHANGHAI JIAOTONG UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211674633.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2025-08-19
Estimated Expiration
2042-12-26

AI Technical Summary

Technical Problem

In the dynamic crowd environment, the local path planning of mobile robots has problems with freezing and insufficient obstacle avoidance capabilities, especially in high-density people, which is difficult to achieve safe and comfortable path planning.

Method used

Establish an MPC dynamic game model, and use linear approximation and inverse optimal control algorithms to identify pedestrian decision weight parameters, and combine robot positioning and pedestrian detection algorithms to realize local planning control in dynamic crowd environments.

Benefits of technology

In a dynamic crowd environment, mobile robots can quickly reach target points, avoid pedestrians and save energy, and the decision-making process is in line with the trade-offs of real pedestrians.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116185005B_ABST
    Figure CN116185005B_ABST
Patent Text Reader

Abstract

The present invention provides a method for local planning of a mobile robot in a crowd environment based on MPC dynamic game theory, comprising: establishing a dynamic game model under the MPC framework with the robot and surrounding pedestrians as participants; linearizing and approximating the dynamic game model to obtain an approximate dynamic game model, and solving the optimality conditions of the approximate model using the maximum principle; designing an inverse optimal control algorithm based on the optimality conditions, and using the inverse optimal control algorithm to identify weight parameters guiding pedestrian decision-making from real pedestrian trajectories; using the identified approximate dynamic game model as a local planner, combined with robot positioning and pedestrian detection algorithms, to achieve local planning and control of the robot in a dynamic crowd environment. This method enables the mobile robot to achieve excellent local planning and control effects in a dynamic crowd environment, such as quickly reaching the target point, avoiding pedestrians, and saving energy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robotics technology, and in particular to a local planning method for a mobile robot based on MPC dynamic game in a crowd environment. Background Art

[0002] With the development of mobile robotics technology, more and more robots are becoming integrated into people's daily lives, providing services that benefit humanity. Examples include medical service robots (smart wheelchairs), food delivery robots, and outdoor patrol robots. These applications often involve large numbers of dynamic pedestrians, which poses significant challenges to the robots' local path planning. For example, unpredictable pedestrian movements can lead to obstacle avoidance failures. Therefore, improving mobile robots' adaptability to dynamic crowd environments and enhancing their local obstacle avoidance capabilities are key technologies in the field of mobile robotics.

[0003] Researchers at home and abroad have also done a lot of research on the local planning problem of robots in dynamic crowd environments, which can be roughly divided into three categories:

[0004] One is reactive local planning. In early studies, researchers designed reactive local planning algorithms that treat pedestrians as general dynamic obstacles. For example, the velocity obstacle algorithm designed by Paolo et al. [1] (Paolo Fiorini and ZviShiller. Motion planning in dynamic environments using velocity obstacles. The International Journal of Robotics Research, 17(7): 760–772, 1998.) can achieve a relatively good real-time obstacle avoidance effect for dynamic obstacles. This local planning method can ensure safety because it takes into account all potential obstacles. However, since it avoids dynamic obstacles detected in real time, it only focuses on short-term next step planning and lacks long-term planning.

[0005] The second is predictive local planning. Predictive local planning first predicts the future trajectory of pedestrians and then makes reasonable decisions based on the predicted future trajectory, thereby successfully planning a safe and comfortable path in the long term. In low-density crowd environments, this planning method performs well. However, when the crowd density increases, the robot often stagnates until the crowd disperses, which is the common robot freezing problem. For details, please refer to the work of Peter et al. [2] (Peter Trautman and Andreas Krause. Unfreezing the robot: Navigation in dense, interacting crowds. In Intelligent Robots and Systems (IROS), 2010 IEEE / RSJ International Conference on, pages 797–803. IEEE, 2010.).

[0006] The third approach is cooperative local planning. The freezing problem arises because previous local planning approaches treat all individuals individually, ignoring interactions between humans and between humans and robots. Therefore, the current research trend is towards cooperative local planning. To improve obstacle avoidance performance, some researchers have mathematically modeled human-human and human-robot interactions, such as common social force models and dynamic game models. Chinese invention patent publication number CN111752276A uses social force models to enable robots to generate local obstacle avoidance behaviors that align with pedestrian expectations. Other researchers have drawn on learning concepts to learn interactions from real pedestrian trajectories and incorporate them into planning, such as reinforcement learning algorithms. Chinese invention patent publication number CN115185281A employs reinforcement learning, taking into account the correlation characteristics between robots and pedestrians, and between pedestrians, to predict pedestrians' future trajectories and then plan for obstacle avoidance. Direct modeling approaches, such as the former, involve weight design, which often requires extensive experience and is cumbersome and laborious to debug. Direct learning approaches, such as the latter, lack interpretability and often have limited generalization capabilities.

[0007] To address this issue, direct modeling can be combined with direct learning. First, based on the theories of optimal control and dynamic game theory, assuming that the robot and pedestrians consistently make optimal planning decisions in their environment in a rational manner, a dynamic game model is constructed between the robot itself and the locally observed pedestrians. This allows the robot to determine the optimal control variables under Nash equilibrium conditions. Simultaneously, utilizing inverse optimal control theory, the weight parameters of the cost function in the dynamic game model are identified offline from real pedestrian trajectories, ensuring that the weight design in the established model is more consistent with human-defined norms.

[0008] Based on previous work, this paper innovatively proposes a local planning algorithm for mobile robots based on MPC dynamic game in crowded environments. This algorithm can be applied to path planning tasks for manned intelligent wheelchairs and catering delivery robots in work scenarios with a large number of pedestrians, such as stations and restaurants. Summary of the Invention

[0009] In view of the defects in the prior art, the purpose of the present invention is to provide a local planning method for mobile robots based on MPC dynamic game in a crowd environment.

[0010] According to the present invention, a local planning method for a mobile robot based on MPC dynamic game in a crowd environment includes:

[0011] Model building steps: Build a dynamic game model under the MPC framework with the robot and surrounding pedestrians as participants;

[0012] Approximate model steps: linearize and approximate the dynamic game model to obtain an approximate dynamic game model, and use the maximum value principle to solve the optimality conditions of the approximate model;

[0013] Identification model steps: Design an inverse optimal control algorithm based on the optimality condition, and use the inverse optimal control algorithm to identify the weight parameters that guide pedestrian decision-making from the actual pedestrian trajectory;

[0014] Local planning step: The identified approximate dynamic game model is used as a local planner, combined with the robot positioning and pedestrian detection algorithms to achieve local planning and control of the robot in a dynamic crowd environment.

[0015] Preferably, the model building step includes:

[0016] Steps for designing state equations: Establish the particle kinematic equations reflecting the motion states of pedestrians and robots respectively;

[0017] Design the cost function: Design the corresponding target item, interaction item, and control item based on the pedestrian's motion decision, and combine their weight coefficients to form the cost function.

[0018] Design dynamic game steps: According to the state equation and cost function, design the optimal control problem of the intelligent agent, combine the optimal control problems of all intelligent agents to obtain a dynamic game model, and use the MPC framework to obtain the MPC dynamic game model, where the intelligent agent is a pedestrian or a mobile robot.

[0019] Preferably, for a mobile robot with nonholonomic constraints, a feedback linearization method is used to convert the unicycle kinematic equation of the robot into a particle kinematic equation of a point outside the robot.

[0020] Preferably, the step of designing a dynamic game includes:

[0021] For each intelligent agent, its decision-making process is constructed as an optimal control problem, and its state equation and cost function are designed;

[0022] At each sampling moment, based on the pedestrian position obtained by real-time detection and the mobile robot position obtained by real-time positioning, the finite-time open-loop optimal control problem of each intelligent agent, consisting of a real-time updated cost function and state equation, is solved online and simultaneously, and the first element of the obtained optimal control sequence is applied to each intelligent agent; at the next sampling moment, the solution operation is repeated, and the new measurement value is used as the initial condition for calculating the optimal control of the multi-agent at the next moment, and the dynamic game problem is refreshed and re-solved.

[0023] Preferably, the approximate model step includes:

[0024] Define the approximate model step: In the process of solving the optimal control problem for each intelligent agent, the cost function is Taylor expanded at the initial conditions to approximate it into a linear quadratic form, taking into account the nonlinear characteristics of the interaction terms in the cost function. This approximates the MPC-LQ dynamic game problem.

[0025] Steps to define the optimality conditions: Under the MPC framework, at each sampling moment, all agents update their own optimal control problems and jointly form a dynamic game problem; solve the optimal control problem of each agent to obtain the optimal solution of each agent, which satisfies the condition that each agent cannot unilaterally change its own decision to make the cost lower when other agents are optimal, that is, the Nash equilibrium definition; by simultaneously solving the optimal control problems of all agents, the optimality conditions in the sense of Nash equilibrium of the dynamic game problem are obtained; for the approximate MPC-LQ dynamic game problem, the Pontryagin maximum principle is used to obtain the optimality conditions in the form of a matrix equation; by solving this matrix equation, the optimal control and optimal trajectory of all agents within the MPC prediction period can be obtained.

[0026] Preferably, the model identification step includes:

[0027] Data collection steps: Use the trajectory collection system to collect multiple sets of real pedestrian trajectories for identifying weight parameters;

[0028] Parameter identification step: The pedestrian decision-making process is a process of weighing various factors. The importance of each factor is reflected in the weight parameters of the cost function θ = [θ1, θ2] TTo make the dynamic game model more accurate, the inverse optimal control algorithm is used to identify the weight coefficient that best describes the pedestrian trade-off process from the collected real pedestrian trajectory dataset. The inverse optimal control algorithm obtains the optimal weight by minimizing the bi-norm of the difference between the real pedestrian trajectory and the optimal trajectory. The designed optimization problem is as follows:

[0029]

[0030] st F k (θ)·S k =C k (θ)

[0031] ≥0

[0032] in: It represents the real position of the i-th agent at time k+1 when collecting the real pedestrian trajectory for the n-th time; represents the optimal position of the i-th agent at time k+1 obtained by solving the approximate MPC-LQ dynamic model at sampling time k when the weight coefficient is θ; M represents the total number of agents; L represents the total number of repeated sampling experiments; F k (θ)·S k =C k (θ) is the optimality condition for the matrix equation obtained by solving the approximate model at sampling time k; S k By x t|k and λ t|k ,,k≤t≤N+k-1, respectively representing the optimal trajectory of M agents in the MPC prediction cycle ( is one of them) and Lagrange multiplier (which can be further transformed into optimal control), N is the number of prediction steps of the MPC model; F k (θ) and C k (θ) consists of the positions of the robot and surrounding pedestrians at sampling time k and the target point related terms.

[0033] Preferably, the local planning step includes:

[0034] Information acquisition step: The goal is to obtain the information required by the dynamic game model in real time at the sampling time, including the positions of the robot and surrounding pedestrians at the sampling time, as well as the target point. The robot's position at the sampling time is obtained by the localization algorithm, and the target point is obtained by the global planning algorithm; the pedestrian's position at the sampling time is obtained by the pedestrian detection algorithm, and the target point is inferred from environmental information;

[0035] Game solving steps: Substitute the positions of the robot and pedestrian at the current sampling moment and the corresponding target point into the identified approximate dynamic game model. Through the optimality conditions under the Nash equilibrium, solve the optimal control under the Nash equilibrium at the next moment, realize local avoidance interaction, and repeat the cycle until the robot reaches the final target point.

[0036] Preferably, the pedestrian detection algorithm adopts a 3D pedestrian detection algorithm, which returns real-time three-dimensional coordinate information and uses coordinate transformation to further transform it into a map coordinate system referenced by the robot positioning.

[0037] Preferably, the Nash equilibrium means that when other pedestrians are other intelligent agents in the dynamic game and the optimal control under the Nash equilibrium condition is adopted, the control of the robot obtained by solving the above dynamic game model is the optimal control that minimizes its cost function.

[0038] Preferably, the model identification step is performed in an offline manner, and the local planning step is performed in an online manner.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] 1. The present invention enables a mobile robot to achieve excellent local planning and control effects such as quickly reaching a target point, avoiding pedestrians, and saving energy in a dynamic crowd environment;

[0041] 2. The present invention enables the process of weighing various factors when a mobile robot makes decisions in a dynamic crowd environment to be more consistent with the weighing process of real pedestrians. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:

[0043] Figure 1 This is a flow chart of the local planning method of a mobile robot based on MPC dynamic game in a crowd environment provided by the present invention. DETAILED DESCRIPTION

[0044] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.

[0045] The present invention provides a local planning method for mobile robots in crowd environments based on MPC dynamic game, which is applicable to the local planning and control problem of mobile robots in dynamic crowd environments. Specifically, it includes the following steps:

[0046] Model building steps: Assuming that the robot and pedestrians always make optimal planning decisions in a rational manner in the environment, design the state equation and cost function, and establish a dynamic game model under the MPC framework with the robot and surrounding pedestrians as participants.

[0047] Approximate model steps: Linearize the dynamic game model to obtain an approximate dynamic game model, and use the maximum principle to solve the optimality conditions of the approximate model. Because the nonlinear characteristics of the cost function of the original dynamic game model make it difficult to solve, linear approximation is used.

[0048] Model identification step: Based on the optimality conditions, an inverse optimal control algorithm is designed. The inverse optimal control algorithm is used to identify the weight parameters that guide pedestrian decision-making from the actual pedestrian trajectories. In this step, the actual pedestrian trajectories are collected using the trajectory acquisition system. The weight parameters that guide pedestrian decision-making are identified based on the actual trajectories, resulting in an approximate dynamic game model that more accurately depicts the real pedestrian trade-off process.

[0049] Local planning step: The identified approximate dynamic game model is used as a local planner, combined with the robot positioning and pedestrian detection algorithms to achieve local planning and control of the robot in a dynamic crowd environment.

[0050] The steps to build the model include:

[0051] 1. Design the state equation: Establish the particle kinematic equations for the pedestrian and robot, respectively. For the pedestrian in the game, the particle kinematic equation is used. For the robot in the game, if it has nonholonomic constraints, feedback linearization can be used to transform the robot's unicycle kinematic equation into the particle kinematic equation for a point outside the robot, achieving a unified state equation representation.

[0052] In a specific embodiment, a pedestrian state equation is designed. If the i-th agent is a pedestrian, its state equation is the particle kinematics equation, and the mathematical model is as follows:

[0053]

[0054] in: represents the xy coordinate position of the i-th agent on the two-dimensional plane at time t when the i-th agent is a pedestrian; represents the speed of the i-th agent in the xy direction of the two-dimensional plane at time t when the i-th agent is a pedestrian; Δt∈R + Indicates the sampling interval (symbol R 2 represents a two-dimensional set of real numbers, R + represents the set of positive real numbers).

[0055] In a specific embodiment, a robot state equation is designed. If the i-th agent is a robot (wheelchair), its state equation is the kinematic equation of a unicycle, and the mathematical model is as follows:

[0056]

[0057]

[0058]

[0059] Its state quantity [x t y t θ t ] T is a three-dimensional vector, representing the xy coordinates of the two-dimensional plane and the steering angle at time t; the control amount [v t w t ] T are two-dimensional vectors representing the linear velocity and angular velocity at time t.

[0060] The robot's unicycle model doesn't match the pedestrian's point kinematic model, making it difficult to unify the state equations. To address this, feedback linearization is used to transform the nonlinear unicycle model into a point kinematic model for a point P directly in front of the robot, at a distance ∈(m) from the robot's center of mass, achieving unification of the state equations.

[0061] Feedback linearization is used to derive the particle kinematic equation of point P. The mathematical model is as follows:

[0062]

[0063] in: It indicates the xy coordinate position of point P on the two-dimensional plane at time t when the i-th agent is a robot; represents the speed of point R in the xy direction of the two-dimensional plane at time t when the i-th intelligent agent is a robot; Δt∈R + Indicates the sampling interval (symbol R 2 represents a two-dimensional set of real numbers, R + represents the set of positive real numbers).

[0064] 2. Cost function design steps: Design the corresponding target items, interaction items, and control items based on the pedestrian's motion decision-making, and combine their respective weight coefficients to form a cost function; when designing, fully consider the pedestrian's decision-making characteristics such as goal-orientedness, avoidance, and effort-saving.

[0065] In a specific embodiment, the target item is designed to take into account that during navigation, the intelligent body will reach the target point as quickly as possible. The target item is designed to be in the form of a quadratic form. Indicates the position of the i-th agent at time t and its target point Minimizing this term allows the agent to move toward the goal point.

[0066] In a specific embodiment, the interaction item is designed. Considering that during the navigation process, the intelligent body will actively interact with other intelligent bodies to avoid each other, the interaction item is designed in the form of represents the interaction loss term between the i-th agent and the j-th agent at time t. Minimizing this term can increase the distance between the two agents and achieve the effect of mutual avoidance.

[0067] In a specific embodiment, the control item is designed to consider that the intelligent body should reduce energy consumption as much as possible during navigation, and the control item is designed in the form of a quadratic represents the control loss term of the ith agent at time t. Minimizing this term can reduce the energy consumption of the agent.

[0068] 3. Design dynamic game steps: Based on the state equation and cost function, design the optimal control problem for each intelligent agent, combine the optimal control problems of all intelligent agents to obtain a dynamic game model, and use the MPC framework to obtain the MPC dynamic game model. The intelligent agent is a pedestrian or a mobile robot.

[0069] In a specific implementation, a dynamic game is constructed. For each intelligent agent in the environment, its decision-making process is modeled as an optimal control problem, and its state equation and cost function are designed; the interaction term in the cost function represents the interaction between the behaviors of multiple intelligent agents. When making decisions, each intelligent agent not only considers its own information (target item, control item), but also considers the influence of the behavior of other intelligent agents (interaction term), thus constituting a dynamic game problem.

[0070] In a specific implementation, an MPC framework is adopted. At the same time, in order to achieve instant interaction between multiple agents, a model predictive control (MPC) framework is adopted. At each sampling moment, based on the current measurement information obtained (pedestrian positions obtained by real-time detection and mobile robot positions obtained by real-time positioning), the finite-time open-loop optimal control problem of each agent consisting of a real-time updated cost function and state equation is solved online and simultaneously, and the first element of the obtained optimal control sequence is applied to each agent; at the next sampling moment, the above process is repeated, and the new measurement value is used as the initial condition for calculating the optimal control of the multiple agents at the next moment, and the dynamic game problem is refreshed and re-solved.

[0071] Assume that at the kth sampling moment, the total number of agents detected is M, and the MPC prediction cycle is designed to be N steps. Combining the above definitions of the state equation and cost function, the MPC dynamic game model is as follows:

[0072]

[0073]

[0074]

[0075] in: represents the cost function of the i-th agent at the k-th sampling moment; represents the target point of the i-th agent; θ=[θ1,θ2] T represents the weight parameter that reflects the decision-making trade-off process; It represents the observed coordinate position of the i-th agent at the k-th sampling moment, and also serves as the initial value of the finite-time open-loop optimization problem defined at the k-th sampling moment in the MPC rolling optimization step, that is, the position at time k represents the optimal position and control of the i-th agent at time t based on the observation at time k.

[0076] The approximate model steps include:

[0077] ① Define the approximate model step: For the i-th agent, the interaction term in its cost function It has nonlinear characteristics, and in practical applications, there are problems such as difficulty in solving Nash equilibrium and too long a time consumption. In order to ensure the real-time performance of the algorithm, the cost function is adjusted based on the initial conditions. Perform Taylor expansion at the point and approximate it to a linear quadratic problem, simplifying the difficulty of solving it.

[0078] After linearization, the approximate MPC-LQ dynamic game model is as follows:

[0079]

[0080]

[0081]

[0082] in: express x k|k The Hessian matrix of Related items; j=1,…,M,j≠i means x k|k The Hessian matrix of Related items; express x k|k The gradient matrix is Related items.

[0083] ② Define the optimality condition steps: Under the MPC framework, at each sampling moment, all agents update their own optimal control problems and jointly form a dynamic game problem; at the same time, solve the optimal control problem of each agent to obtain the optimal solution of each agent. This optimal solution satisfies the condition that when other agents are optimal, each agent cannot unilaterally change its own decision to make the cost smaller, which is the Nash equilibrium definition; by simultaneously solving the optimal control problems of all agents, the optimality conditions of the dynamic game problem in the sense of Nash equilibrium are obtained; for the approximate MPC-LQ dynamic game problem, the Pontryagin Maximum Principle (PMP) is used to obtain its optimality conditions.

[0084] First, the constrained optimization problem is transformed into an unconstrained optimization problem to obtain the Hamiltonian function as follows:

[0085]

[0086] in: is the corresponding Lagrange multiplier.

[0087] right Taking the derivative and setting it to 0, we get:

[0088]

[0089] Substituting into the state equation we can get:

[0090]

[0091] right Taking the derivative and setting it to 0, we can get:

[0092]

[0093] Putting all the equations for the agents together, we get the optimality condition in the form of a matrix equation as follows:

[0094]

[0095] in:( represents the Kronecker product; I 2MN , I N , I M , I2 represent the unit matrices of size 2, N, M, and 2 respectively; 1 N×1 represents an all-1 vector of size N×1; F 1,k (θ), F2, F3, C 1,k The numerical subscripts in (θ) and C2 only indicate different variables and have no other specific meanings. The subscript k indicates that it is related to the measurement value at time k.

[0096]

[0097]

[0098]

[0099] It represents the optimal trajectory of M agents in the MPC prediction period N steps obtained by solving the approximate game model based on k-time observations.

[0100] It means that based on k-time observations, the approximate game model is solved to obtain the N-step Lagrange multipliers of the MPC prediction period of M agents, which can be further transformed to obtain the optimal control.

[0101] The steps of identifying the model include:

[0102] a. Data collection step: Use the trajectory collection system to collect multiple sets of real pedestrian trajectories for identification of weight parameters;

[0103] b. Parameter identification step: The pedestrian decision-making process is a process of weighing various factors. The importance of each factor is reflected in the weight parameters of the cost function θ = [θ1, θ2] T If the model is manually designed, it may not be accurate. To make the dynamic game model more accurate, the inverse optimal control algorithm is used to identify the weight coefficients that best describe the pedestrian's trade-off process from the collected real pedestrian trajectory data set. The inverse optimal control algorithm assumes that the difference between the optimal trajectory obtained under the optimal weight and the actual pedestrian trajectory should be minimized. Therefore, the optimal weight is obtained by minimizing the bi-norm of the difference between the actual pedestrian trajectory and the optimal trajectory obtained by solving the solution. The designed optimization problem is as follows:

[0104]

[0105] st F k (θ)·S k =C k (θ)

[0106] ≥0

[0107] in: It represents the real position of the i-th agent at time k+1 when collecting the real pedestrian trajectory for the n-th time; represents the optimal position of the i-th agent at time k+1 obtained by solving the approximate MPC-LQ dynamic model at sampling time k when the weight coefficient is θ; M represents the total number of agents; L represents the total number of repeated sampling experiments; F k (θ)·S k =C k(θ) is the optimality condition for the matrix equation obtained by solving the approximate model at sampling time k; S k By x t|k and λ t|k ,,k≤t≤N+k-1, respectively representing the optimal trajectory of M agents in the MPC prediction cycle ( is one of them) and Lagrange multiplier (which can be further transformed into optimal control), N is the number of prediction steps of the MPC model; F k (θ) and C k (θ) consists of the positions of the robot and surrounding pedestrians at sampling time k and the target point related terms.

[0108] In a preferred embodiment, the model identification step is performed offline.

[0109] The local planning steps include:

[0110] 1) Information acquisition step: Real-time acquisition of the information required by the dynamic game model at time k, including the positions of the robot and the surrounding pedestrians at time k. and target points The robot's position at time k is obtained by the positioning algorithm, the target point at time k is obtained by the global planning algorithm, the pedestrian's position at time k is obtained by the pedestrian detection algorithm, and the target point at time k is inferred from environmental information;

[0111] 2) Game solving step: Substitute the positions of the robot and pedestrian at the current sampling moment and the corresponding target point into the identified approximate dynamic game model. Based on the optimality conditions under the Nash equilibrium, solve the optimal control under the Nash equilibrium at the next moment, realize local avoidance interaction, and repeat the cycle until the robot reaches the final target point.

[0112] More specifically, the global planning algorithm used in the information acquisition step is mainly aimed at low-dynamic features (static obstacles). By establishing a global map of the static environment in advance, and then using A* and other methods to perform global path planning on the established two-dimensional grid map, the waypoints from the robot's current position to the target position are planned as the temporary target points of the robot in local planning.

[0113] More specifically, the pedestrian detection algorithm used in the information acquisition step is a 3D pedestrian detection algorithm, which needs to return real-time three-dimensional coordinate information, and then further convert it into the map coordinate system referenced by the robot positioning through coordinate transformation; the environmental information used to infer the target point refers to directional environmental structures such as elevator doors and emergency exits.

[0114] More specifically, the Nash equilibrium means that when other pedestrians are other intelligent agents in the dynamic game and the optimal control under the Nash equilibrium condition is adopted, the control of the robot obtained by solving the above dynamic game model is the optimal control that minimizes its cost function.

[0115] More specifically, the local planning step solves the approximate dynamic game model to find the optimal control for the xy direction velocity of point P. and It needs to be converted into the robot's control quantity linear velocity v through the feedback linearization inverse process t and angular velocity w t The conversion equation is as follows:

[0116]

[0117] In a preferred embodiment, the local planning step is performed online.

[0118] Those skilled in the art are well aware that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various devices, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; the devices, modules, and units for implementing various functions can also be considered as both software modules implementing the method and structures within the hardware component.

[0119] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.

Claims

1. A local planning method for mobile robots based on MPC dynamic game in a crowd environment, characterized by: include: Model building steps: Build a dynamic game model under the MPC framework with the robot and surrounding pedestrians as participants; Approximate model steps: linearize and approximate the dynamic game model to obtain an approximate dynamic game model, and use the maximum value principle to solve the optimality conditions of the approximate model; Identification model steps: Design an inverse optimal control algorithm based on the optimality condition, and use the inverse optimal control algorithm to identify the weight parameters that guide pedestrian decision-making from the actual pedestrian trajectory; Local planning step: Using the identified approximate dynamic game model as a local planner, combined with robot positioning and pedestrian detection algorithms, local planning and control of the robot in a dynamic crowd environment is achieved; The approximate model step includes: Define the approximate model step: In the process of solving the optimal control problem for each intelligent agent, the cost function is Taylor expanded at the initial conditions to approximate it into a linear quadratic form, taking into account the nonlinear characteristics of the interaction terms in the cost function. This approximates the MPC-LQ dynamic game problem. Steps to define the optimality conditions: Under the MPC framework, at each sampling moment, all agents update their own optimal control problems and jointly form a dynamic game problem; solve the optimal control problem of each agent to obtain the optimal solution of each agent, which satisfies the condition that each agent cannot unilaterally change its own decision to make the cost lower when other agents are optimal, that is, the Nash equilibrium definition; by simultaneously solving the optimal control problems of all agents, the optimality conditions in the sense of Nash equilibrium of the dynamic game problem are obtained; for the approximate MPC-LQ dynamic game problem, the Pontryagin maximum principle is used to obtain the optimality conditions in the form of a matrix equation; by solving this matrix equation, the optimal control and optimal trajectory of all agents within the MPC prediction period can be obtained.

2. The mobile robot local planning method based on MPC dynamic game in a crowd environment according to claim 1 is characterized by: The model building step includes: Steps for designing state equations: Establish the particle kinematic equations reflecting the motion states of pedestrians and robots respectively; Design the cost function: Design the corresponding target item, interaction item, and control item according to the pedestrian's motion decision, and combine their weight coefficients to form the cost function; Design dynamic game steps: According to the state equation and cost function, design the optimal control problem of the intelligent agent, combine the optimal control problems of all intelligent agents to obtain a dynamic game model, and use the MPC framework to obtain the MPC dynamic game model, where the intelligent agent is a pedestrian or a mobile robot.

3. The mobile robot local planning method based on MPC dynamic game in a crowd environment according to claim 2, characterized in that: Aiming at the mobile robot with nonholonomic constraints, the feedback linearization method is used to transform the unicycle kinematic equation of the robot into the kinematic equation of a particle outside the robot.

4. The mobile robot local planning method based on MPC dynamic game in a crowd environment according to claim 2 is characterized in that: The steps of designing a dynamic game include: For each intelligent agent, its decision-making process is constructed as an optimal control problem, and its state equation and cost function are designed; At each sampling moment, based on the pedestrian position obtained by real-time detection and the mobile robot position obtained by real-time positioning, the finite-time open-loop optimal control problem of each intelligent agent, consisting of a real-time updated cost function and state equation, is solved online and simultaneously, and the first element of the obtained optimal control sequence is applied to each intelligent agent; at the next sampling moment, the solution operation is repeated, and the new measurement value is used as the initial condition for calculating the optimal control of the multi-agent at the next moment, and the dynamic game problem is refreshed and re-solved.

5. The mobile robot local planning method based on MPC dynamic game in a crowd environment according to claim 1 is characterized in that: The identification model step includes: Data collection steps: Use the trajectory collection system to collect multiple sets of real pedestrian trajectories for identifying weight parameters; Identification parameter steps: The pedestrian decision-making process is a process of weighing various factors. The importance of each factor is reflected in the weight parameters of the cost function. The inverse optimal control algorithm is used to identify the weight coefficient that best describes the pedestrian's trade-off process from the collected real pedestrian trajectory dataset. The inverse optimal control algorithm obtains the optimal weight by minimizing the bi-norm of the difference between the real pedestrian trajectory and the optimal trajectory. The designed optimization problem is as follows: in: Indicates the time when collecting real pedestrian trajectories Repeated collection of An intelligent agent in The actual location at the moment; Indicates that the weight coefficient is By solving the sampling time The approximate MPC-LQ dynamic model obtained An intelligent agent in The optimal position at the moment; represents the total number of agents; Indicates the total number of repeated acquisition experiments; That is to solve the sampling time Optimality conditions for the matrix equality obtained from the approximate model; Depend on and , Composition, respectively The optimal trajectory and Lagrange multiplier of each agent in the MPC prediction cycle, is one of them, the Lagrange multiplier can be further transformed into optimal control, is the number of prediction steps of the MPC model; and At the sampling moment, the robot and the surrounding pedestrians The position and target point related items are composed.

6. The mobile robot local planning method based on MPC dynamic game in a crowd environment according to claim 1 is characterized in that: The local planning step includes: Information acquisition step: Real-time acquisition of the information required by the dynamic game model at the sampling moment, including the positions of the robot and surrounding pedestrians at the sampling moment, as well as the target point. The robot's position at the sampling moment is obtained by the positioning algorithm, the target point is obtained by the global planning algorithm, the pedestrian's position at the sampling moment is obtained by the pedestrian detection algorithm, and the target point is inferred from environmental information. Game solving steps: Substitute the positions of the robot and pedestrian at the current sampling moment and the corresponding target point into the identified approximate dynamic game model. Through the optimality conditions under the Nash equilibrium, solve the optimal control under the Nash equilibrium at the next moment, realize local avoidance interaction, and repeat the cycle until the robot reaches the final target point.

7. The mobile robot local planning method based on MPC dynamic game in a crowd environment according to claim 6, characterized in that: The pedestrian detection algorithm adopts a 3D pedestrian detection algorithm, which returns real-time three-dimensional coordinate information and uses coordinate transformation to further convert it into a map coordinate system referenced by the robot positioning.

8. The mobile robot local planning method based on MPC dynamic game in a crowd environment according to claim 6, characterized in that: The Nash equilibrium means that when other pedestrians are other intelligent agents in the dynamic game and the optimal control under the Nash equilibrium condition is adopted, the control of the robot obtained by solving the above dynamic game model is the optimal control that minimizes its cost function.

9. The mobile robot local planning method based on MPC dynamic game in a crowd environment according to claim 1, characterized in that: The identification model step is performed offline, and the local planning step is performed online.

Citation Information

Patent Citations

  • Local path planning method and device, computer readable storage medium and robot

    CN111752276A

  • Mobile robot obstacle avoidance method and device based on pedestrian prediction

    CN115185281A

  • Collaborative strategy inversion identification method based on multi-agent game playing

    CN112270103A

  • Master-slave game type human-machine cooperative steering control method in ice and snow environment

    CN113553726A