Navigation Method and Device for Mobile Robot
By combining the combined model of memory pool, state prediction module and value estimation module, pedestrian motion information is processed, and the problem of inaccurate navigation prediction caused by ignoring pedestrian interaction in mobile robot navigation is solved, and higher navigation prediction accuracy and navigation action optimization are achieved.
Patent Information
- Application Number
- CN202510060918.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-15
AI Technical Summary
In the prior art, mobile robot navigation methods ignore the interaction between pedestrians, resulting in inaccurate navigation prediction.
By obtaining pedestrian location information, using a combined model of memory pool, state prediction module and value estimation module, we process historical, neighbor and future motion information, accurately predict pedestrian status, and determine target navigation actions based on robot status.
It improves the accuracy of navigation prediction of mobile robots in different scenarios, enhances the ability to predict pedestrian motion trajectories, optimizes navigation movements, and avoids collisions with pedestrians.
Smart Images

Figure CN119472702B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot navigation, and in particular to a navigation method and device for a mobile robot. Background Art
[0002] Mobile robot navigation refers to the technology in which a mobile robot perceives the surrounding environment, autonomously plans a path, and moves along a predetermined or optimized path.
[0003] For a mobile robot, navigating in a crowd is a huge challenge. On the one hand, people in the environment cannot be simply regarded as static or dynamic obstacles, and there may be some complex interaction relationships between people. On the other hand, the movement of pedestrians is random and anisotropic, making it difficult to accurately model based on certain rules.
[0004] It can be seen that the mobile robot navigation method in the related art has the technical problem of inaccurate navigation prediction caused by ignoring the interaction relationship between pedestrians. Summary of the Invention
[0005] The present invention provides a navigation method and device for a mobile robot, so as to solve the defect that the mobile robot navigation method in the prior art has inaccurate navigation prediction caused by ignoring the interaction relationship between pedestrians, and achieve the improvement of the prediction accuracy of the mobile robot in different scenarios.
[0006] The present invention provides a navigation method for a mobile robot, including the following steps.
[0007] Obtain the pedestrian position information collected by the mobile robot, where the pedestrian position information includes: the position information sequence of each pedestrian; input the pedestrian position information into the robot navigation model to obtain the target navigation action output by the robot navigation model. The robot navigation model includes: a memory pool, a state prediction module, and a value estimation module, where: through the memory pool, within the current observation time window, use the position information sequence of the target pedestrian before the target time step as historical motion information; use the position information sequences of other pedestrians before the target time step as neighbor motion information; use the position information sequence of the target pedestrian after the target time step as future motion information; through the state prediction module, based on the historical motion information, the neighbor motion information, and the future motion information, determine the predicted pedestrian state of the target pedestrian; use each pedestrian within the observation time window as the target pedestrian to determine the predicted pedestrian state of each pedestrian; through the value estimation module, obtain the robot state of the mobile robot, and based on the robot state and the predicted pedestrian state of each pedestrian, determine the target navigation action of the mobile robot.
[0008] A navigation method for a mobile robot provided according to the present invention, the memory pool includes: a first memory pool, a second memory pool, and a third memory pool; the first memory pool is used to store the robot state and environmental rewards at each time step during the training of the robot navigation model; the second memory pool is used to store the position information sequences of each pedestrian within a historical observation time window; the third memory pool is used to store the position information sequences of each pedestrian within the current observation time window.
[0009] A navigation method for a mobile robot provided according to the present invention, the state prediction module includes: a recurrent neural network and a decoder; determining the predicted pedestrian state of the target pedestrian based on the historical motion information, the neighbor motion information, and the future motion information includes: encoding the historical motion information, the neighbor motion information, and the future motion information through the recurrent neural network based on the attention mechanism to obtain the social feature vector of the target pedestrian; decoding the social feature vector through the decoder to obtain the predicted pedestrian state of the target pedestrian.
[0010] A navigation method for a mobile robot provided according to the present invention, the value estimation module includes: a relational graph network and a value network; determining the target navigation action of the mobile robot based on the robot state and the predicted pedestrian state of each pedestrian includes: constructing a feature matrix and a relational matrix through the relational graph network based on the robot state and the predicted pedestrian state of each pedestrian; calling a graph convolutional network to encode based on the feature matrix and the relational matrix to obtain an interaction relationship feature; estimating the action value of the interaction relationship feature through the value network based on a preset action value function, and determining the action with the highest value in the action space as the target navigation action.
[0011] A navigation method for a mobile robot provided according to the present invention, before inputting the pedestrian position information into the robot navigation model to obtain the target navigation action output by the robot navigation model, the method further includes: training a preset robot navigation model based on knowledge distillation to obtain a trained robot navigation model.
[0012] A navigation method for a mobile robot provided by the present invention, the training includes: teacher model training and student model training. Training a preset robot navigation model based on knowledge distillation to obtain a trained robot navigation model, including: performing a running simulation of a preset teacher model in a target scenario based on an optimal mutual collision avoidance strategy to obtain target scenario experience data; performing a running simulation of a preset student model in a source scenario to obtain source scenario experience data; initializing a state prediction module and a value estimation module in the preset teacher model based on the target scenario experience data; jointly training the state prediction module and the value estimation module in the preset teacher model after initialization based on reinforcement learning and supervised learning to obtain a trained teacher model; initializing a state prediction module and a value estimation module in the preset student model based on the source scenario experience data through the trained teacher model; jointly training the state prediction module and the value estimation module in the preset student model after initialization based on reinforcement learning and supervised learning to obtain a trained student model; using the trained student model as the trained robot navigation model.
[0013] The present invention also provides a navigation device for a mobile robot, including the following modules: an acquisition module, configured to acquire pedestrian position information collected by the mobile robot, where the pedestrian position information includes: a position information sequence of each pedestrian; a robot navigation module, configured to input the pedestrian position information into a robot navigation model to obtain a target navigation action output by the robot navigation model. The robot navigation model includes: a memory pool, a state prediction module, and a value estimation module, where: through the memory pool, within a current observation time window, using the position information sequence of a target pedestrian before a target time step as historical motion information; using the position information sequences of other pedestrians before the target time step as neighbor motion information; using the position information sequence of the target pedestrian after the target time step as future motion information; through the state prediction module, determining a predicted pedestrian state of the target pedestrian based on the historical motion information, the neighbor motion information, and the future motion information; using each pedestrian within the observation time window as the target pedestrian to determine the predicted pedestrian state of each pedestrian; through the value estimation module, obtaining a robot state of the mobile robot, and determining a target navigation action of the mobile robot based on the robot state and the predicted pedestrian state of each pedestrian.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the navigation method of the mobile robot as described in any one of the above.
[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the navigation method of the mobile robot as described in any one of the above is implemented.
[0016] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the navigation method of the mobile robot as described in any one of the above is implemented.
[0017] The navigation method and device of the mobile robot provided by the present invention collect the position information of pedestrians through the mobile robot, including the position information sequence of each pedestrian; within the current observation time window, store and process the historical motion information, neighbor motion information, and future motion information of the target pedestrian, and can more comprehensively understand the motion pattern of the pedestrian; based on the attention mechanism, can accurately determine the predicted pedestrian state of the target pedestrian; by taking each pedestrian within the observation time window as the target pedestrian respectively, the predicted pedestrian state of each pedestrian can be determined, thereby enhancing the prediction accuracy of the pedestrian motion trajectory; obtain the robot state of the mobile robot, and combine the predicted pedestrian state of each pedestrian to determine the target navigation action of the mobile robot. Thus, the robot can more accurately judge the relative position relationship between itself and the pedestrian, thereby optimizing the navigation action and avoiding collisions with the pedestrian; furthermore, it can process the pedestrian motion information in different scenarios, enabling the robot to maintain stable navigation accuracy in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the description of the embodiments or the prior art will be briefly introduced one by one below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 is a schematic flowchart of the navigation method of the mobile robot provided by the present invention.
[0020] Figure 2 is a schematic diagram of the navigation strategy of the mobile robot provided by the present invention.
[0021] Figure 3 is a schematic structural diagram of the knowledge distillation training framework provided by the present invention.
[0022] Figure 4 is a schematic module structure diagram of the navigation device of the mobile robot provided by the present invention.
[0023] Figure 5 is a schematic entity structure diagram of the electronic device provided by the present invention. Detailed Implementation Modes
[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0025] For a mobile robot, navigating in a crowd is a huge challenge. On the one hand, people in the environment cannot be simply regarded as static or dynamic obstacles, and there may be some complex interaction relationships between people. On the other hand, the movement of pedestrians is random and anisotropic, making it difficult to accurately model based on certain rules. The robot must observe and predict the movement of surrounding humans in this complex and dynamic environment, while maintaining a certain safety distance to avoid collisions.
[0026] In related technologies, some methods help the robot make decisions by making assumptions about human movement behaviors, but these assumptions are usually unrealistic. There are also some methods that avoid modeling pedestrians by directly learning a policy network, but these methods ignore the interaction relationships between pedestrians in various scenarios.
[0027] The models in related technologies are trained in specific scenarios. Although they can achieve good results in the original scenarios, when the scenarios change, there may be a certain degree of deviation between the prediction results of the models and the actual states. Therefore, the performance of the methods will decline. In fact, most of the existing work does not discuss the ability to apply the navigation strategy of the robot to different scenarios, that is, the scenario generalization performance of the model, which is crucial for the actual application of the robot.
[0028] The present invention provides a navigation method for a mobile robot. On the one hand, a pedestrian state predictor based on the SocialVAE (Social Variational AutoEncoder) method is designed to predict the pedestrian state according to the historical position information of pedestrians and their neighbors, and social interactions between pedestrians are considered, which not only improves the prediction accuracy but also reduces the dependence on scenarios. On the other hand, a method for training a robot navigation model based on knowledge distillation is designed to distill the trained pedestrian state prediction module in different scenarios, so that the finally trained navigation model can learn information about different scenarios.
[0029] Optionally, the navigation method of the mobile robot according to the embodiments of the present application can be executed by a server, or can be executed by a terminal device (e.g., a mobile robot), or can also be jointly executed by the server and the terminal device. Taking the execution of the navigation method of the mobile robot in this embodiment by the server as an example.
[0030] Figure 1 is a schematic flow chart of the navigation method of the mobile robot provided by the present invention. As Figure 1 shown, the method includes the following steps.
[0031] Step 101, obtain the pedestrian position information collected by the mobile robot, where the pedestrian position information includes: the position information sequence of each pedestrian.
[0032] In some embodiments, the mobile robot is pre-equipped with multiple hardware devices for collecting pedestrian position information and its own state information.
[0033] For example, a laser sensor, a sonar sensor, an infrared sensor, and a camera. Among them, the laser sensor is used to obtain high-precision distance information and is applicable to indoor and outdoor environments. The sonar sensor is applicable to indoor environments and can measure the distance to obstacles. The infrared sensor is used for obstacle avoidance and close-range detection and is usually used as an auxiliary sensor. The camera is used for visual navigation and positioning and can capture the image information of pedestrians.
[0034] In some embodiments, the mobile robot senses the surrounding environment in real time through sensors, including pedestrians, obstacles, etc. The sensor data is pre-processed, such as denoising, filtering, etc., to improve the accuracy and reliability of the data. Image processing algorithms (such as object detection and tracking algorithms) are used to detect pedestrians in the images captured by the camera, or through data fusion of the laser sensor and the sonar sensor, three-dimensional detection and tracking of pedestrians are realized. According to the results of pedestrian detection and tracking, the position information sequence of each pedestrian is generated, and the position information sequence includes the position coordinates of the pedestrian (such as X, Y, Z coordinates), speed, direction, etc.
[0035] Step 102, input the pedestrian position information into the robot navigation model to obtain the target navigation action output by the robot navigation model.
[0036] Refer to Figure 2 , Figure 2 is a schematic diagram of the mobile robot navigation strategy provided by the present invention, which includes: Environment, Memory, and Action Space. Among them, the Action Space includes: State Predictor and Value Estimator, represents the target navigation action.
[0037] The above-mentioned robot navigation model includes: a memory pool, a state prediction module, and a value estimation module, where, it includes:
[0038] Step 1021, through the memory pool, within the current observation time window, take the position information sequence of the target pedestrian before the target time step as historical motion information; take the position information sequence of other pedestrians before the target time step as neighbor motion information; take the position information sequence of the target pedestrian after the target time step as future motion information.
[0039] According to the navigation method of the mobile robot provided by the present invention, the above-mentioned memory pool includes: a first memory pool, a second memory pool, and a third memory pool;
[0040] The first memory pool is used to store the robot state and environmental reward at each time step during the training process of the robot navigation model;
[0041] The second memory pool is used to store the position information sequence of each pedestrian within the historical observation time window;
[0042] The third memory pool is used to store the position information sequence of each pedestrian within the current observation time window.
[0043] As Figure 2 shown, through the movement and interaction of the mobile robot in the environment, record the state and environmental reward of the mobile robot at each moment into the memory pool. In order to conduct training and testing, three memory pools are designed.
[0044] The first memory pool (Experience Memory) is used to store the state and reward during the training process. The second memory pool (Trajectory Memory) is used to store the historical trajectories of pedestrians (the position information sequence of each pedestrian within the historical observation time window), that is, store the corresponding position information of humans in time series. These two memory pools are only used and updated during training, and are respectively used to provide training data for the value estimation module and the state prediction module. The third memory pool (Horizon Memory) is used to store the position information sequence of pedestrians in the latest observation time window of the current round. Therefore, this memory pool maintains a fixed time length. In order to facilitate the prediction of pedestrian states in the simulation step, at the beginning of each round, only pedestrians will be observed first, and the robot action decision will be made after the time length of this memory pool reaches the length of the observation time window.
[0045] It should be noted that the collected pedestrian position information includes the position information sequences of all n humans within a past continuous time period T0. The position information sequence of the first T1 time steps of this sequence is used as historical position information, and the position information sequence of the subsequent T2 time steps is used as future position information.
[0046] For each target pedestrian i, the historical motion information of the target pedestrian is obtained based on the historical position information of the first T1 time steps of the target pedestrian, including motion information such as position, velocity, and acceleration. The neighbor motion information of the target pedestrian is obtained based on the historical position information of the first T1 time steps of the other n - 1 humans, also including motion information such as position, velocity, and acceleration. The future motion information of the target pedestrian is obtained based on the future position information of the target pedestrian in the subsequent T2 time steps. T1 and T2 are pre-set parameters, where T0 = T1 + T2.
[0047] In the training step, the collected pedestrian position information here comes from the sampling of the second memory pool. The goal is to make a prediction based on the historical position information of the first T1 time steps, and compare the prediction result with the future position information of the subsequent T2 time steps to optimize the parameters of the predictor. In the simulation step, the collected pedestrian position information comes from the third memory pool. At this time, the historical observation data only includes the known position information of T1 time steps, and the goal is to predict the future position information of T2 time steps based on the known position information of T1 time steps.
[0048] In some embodiments, the first memory pool (experience memory pool) stores the states and rewards of the mobile robot during training, and is used to train the value estimation module or the state prediction module. The content stored in the first memory pool includes the state of the mobile robot at each moment and the rewards returned by the environment. After each interaction between the mobile robot and the environment, new experiences (robot states and environmental rewards) are added to the first memory pool.
[0049] The second memory pool (trajectory memory pool) stores the historical trajectories of pedestrians and is used to train the state prediction module or imitation learning. Each entry of the stored content contains the position information sequence of a pedestrian within a certain time period, and this position information may include coordinates, velocity, direction, etc. When a pedestrian moves, new position information is added to the trajectory memory pool.
[0050] The third memory pool (horizontal memory pool) stores the sequence of the positions of pedestrians in the latest observation time window of the current round, which is used for predicting the states of pedestrians in the simulation steps. The stored content includes the continuous sequence of the positions of pedestrians within the current round, and the length of the sequence is equal to the length of the observation time window. At the beginning of each round, only observations are made without making action decisions. When the time length in the memory pool reaches the length of the observation time window, the robot starts to make action decisions based on the information in the memory pool. During the observation process, as new human position information arrives, the content in the memory pool is continuously updated, and past information is discarded to maintain a fixed time length.
[0051] Through the embodiments of the present invention, rich training data is provided by storing the states and rewards of the mobile robot during the training process; by storing the historical trajectories of humans, the mobile robot can learn the behavior patterns of humans; by storing the sequence of the positions of humans in the latest observation time window of the current round, the mobile robot can predict the possible positions of pedestrians in a short time, so as to adjust its navigation path in time and avoid potential collision risks.
[0052] Step 1022, through the state prediction module, based on the historical motion information, neighbor motion information, and future motion information, determine the predicted pedestrian state of the target pedestrian; each pedestrian within the observation time window is used as the target pedestrian respectively to determine the predicted pedestrian state of each pedestrian.
[0053] According to the navigation method of the mobile robot provided by the present invention, the above state prediction module includes: a recurrent neural network and a decoder;
[0054] Based on the historical motion information, neighbor motion information, and future motion information, determining the predicted pedestrian state of the target pedestrian includes:
[0055] Through the recurrent neural network, encode the historical motion information, neighbor motion information, and future motion information based on the attention mechanism to obtain the social feature vector of the target pedestrian;
[0056] Through the decoder, decode the social feature vector to obtain the predicted pedestrian state of the target pedestrian.
[0057] For the state prediction of the mobile robot, assuming that the actions of the mobile robot can be perfectly executed, the next state of the mobile robot can be directly calculated from the current state and actions.
[0058] For the state prediction of pedestrians, the embodiments of the present invention design a pedestrian state predictor based on SocialVAE. This pedestrian state predictor first preprocesses the data in the memory pool or to obtain the historical motion information in the historical observations , neighbor movement information and future movement information .
[0059] Among them, SocialVAE (Social Variational AutoEncoder) is a method for pedestrian trajectory prediction. Its core lies in using the timewise variational autoencoder architecture and combining it with stochastic recurrent neural networks for prediction.
[0060] Based on the attention mechanism and recurrent neural network, encode the above historical movement information, neighbor movement information, and future movement information to extract the movement features of pedestrians, the movement features of surrounding neighbors, and the social feature vectors between pedestrians and surrounding neighbors. The encoded social feature vectors are processed by the decoder to finally obtain the predicted pedestrian state , where represents the predicted time interval, and the predicted pedestrian state includes the position, speed, acceleration, and moving direction of the pedestrian, etc., which specifically depends on the requirements of the application.
[0061] Among them, the social feature vector refers to the feature vector encoded according to the target pedestrian movement information (i.e., historical movement information and future movement information) and the surrounding neighbor movement information (i.e., neighbor movement information). The social feature vector implicitly describes the relative position and speed relationship between the target pedestrian and the surrounding neighbors.
[0062] Through the embodiments of the present invention, through the encoding process of the attention mechanism and RNN, the model can finely extract the movement features of pedestrians and the movement features of neighbors, as well as the social interaction features between them, which helps the model to more accurately understand the movement patterns and intentions of pedestrians.
[0063] Step 1023, through the value estimation module, obtain the robot state of the mobile robot, and based on the robot state and the predicted pedestrian state of each pedestrian, determine the target navigation action of the mobile robot.
[0064] According to the navigation method of the mobile robot provided by the present invention, the above value estimation module includes: a relational graph network and a value network;
[0065] Based on the robot state and the predicted pedestrian state of each pedestrian, determining the target navigation action of the mobile robot includes:
[0066] Through the relational graph network, based on the robot state and the predicted pedestrian state of each pedestrian, construct a feature matrix and a relational matrix;
[0067] Call the graph convolutional network to encode based on the feature matrix and the relationship matrix to obtain interaction relationship features;
[0068] Through the value network, based on a preset action-value function, perform action-value estimation on the interaction relationship features, and determine the action with the highest value in the action space as the target navigation action.
[0069] In the embodiment of the present invention, the value estimation module includes a relational graph network and a value network.
[0070] The relational graph is a directed graph. The nodes in the graph represent the states of the agents. The directed edges between the nodes represent the degree of attention of one agent to another agent. Among them, the agents include robots and pedestrians. In the relational graph network, first encode the robot state and the pedestrian state respectively through different multi-layer perceptrons into state feature vectors of equal length, and combine the state feature vectors together to form a feature matrix. According to the similarity degree between every two feature vectors (which can be measured by the embedded gaussian function), obtain the relationship vectors, and combine all the relationship vectors into a relationship matrix. The relational graph network infers the feature matrix and the relationship matrix based on the current state (multiple predicted pedestrian states and the robot state).
[0071] Among them, the robot state is known information that can be obtained in real time, including position, speed, etc. The action is any action selected from the action space. The above action is the speed. According to the current position and speed, the position at the next moment can be calculated, so as to obtain the robot state at the next moment.
[0072] Then extract the interaction relationship features between the robot, the person, and between people through the graph convolutional network (the interaction relationship features are vectors obtained by further encoding the feature matrix and the relationship matrix, without a clear meaning, but implicitly represent the relative interaction relationship between the agents, that is, the degree of attention between the agents). The value network is a multi-layer perceptron (MLP), which performs action-value estimation according to the interaction relationship features extracted from the relational graph. Using the obtained action-value function , select the current optimal navigation action from the action space (Action Space) as the target navigation action.
[0073] In some embodiments, a relational graph is constructed through a relational graph network (Relational Graph Network, RGN). The nodes in the relational graph represent the states of the agents (robots and pedestrians). Each node contains the current state information of the agent, such as position, speed, etc. The directed edges between the nodes represent the degree of attention of one agent to another agent. This degree of attention can be quantified by calculating parameters such as relative position, speed difference, and predicted trajectory overlap, and used as the weight of the edge.
[0074] Construct a feature matrix and a relationship matrix based on the relationship graph. Among them, the feature matrix is used to store the feature vectors of each node, and these vectors describe the state information of the agent; the relationship matrix stores the interaction relationship features between each pair of nodes, usually a two-dimensional matrix, where each element represents the interaction intensity or attention degree between a pair of nodes.
[0075] Input the feature matrix and the relationship matrix into the graph convolutional network. Through the graph convolutional network (GCN), use the relationship matrix to update the feature matrix to capture the interaction relationships between agents, and output the interaction relationship features.
[0076] Through the value network (multi-layer perceptron), based on the interaction relationship features, through multi-layer non-linear transformations (such as ReLU activation function) and linear transformations, extract high-level feature representations; output the action value function, and the action value function represents the expected benefits or costs obtained by performing different actions in a given state; according to the action value function, select the action with the highest value from the action space as the current optimal navigation action.
[0077] Through the embodiments of the present invention, through the relationship graph network and GCN, the module can accurately capture the interaction relationships between agents; through the value network, accurate action value estimation can be performed based on the interaction relationship features, and the optimal navigation action can be selected for the robot.
[0078] According to the navigation method of the mobile robot provided by the present invention, before inputting the pedestrian position information into the robot navigation model and obtaining the target navigation action output by the robot navigation model, the above method further includes:
[0079] Train a preset robot navigation model based on knowledge distillation to obtain a trained robot navigation model.
[0080] Refer to Figure 3 , Figure 3 is a schematic structural diagram of the knowledge distillation training framework provided by the present invention, which includes input data (Data), a teacher model (Teacher Model) and a student model (Student Model). Among them, the teacher model includes: a first encoder (Encoder1), a second encoder (Encoder2) and a decoder (Decoder); the student model includes: a first encoder (Encoder1), a second encoder (Encoder2) and a decoder (Decoder).
[0081] In an embodiment of the present invention, a preset robot navigation model is trained based on Knowledge Distillation, aiming to optimize a simpler or more compact Student Model by extracting knowledge from one or more complex and superior-performing Teacher Models, so as to obtain a trained robot navigation model.
[0082] Through the embodiment of the present invention, not only can the performance of the model be improved, but also the complexity and computational requirements of the model can be reduced, facilitating deployment in practical applications.
[0083] According to the navigation method of the mobile robot provided by the present invention, the above training includes: Teacher Model training and Student Model training. Training the preset robot navigation model based on Knowledge Distillation to obtain a trained robot navigation model includes:
[0084] Based on the preset Teacher Model, perform running simulation in the target scenario based on the best mutual collision avoidance strategy to obtain target scenario experience data;
[0085] Based on the preset Student Model, perform running simulation in the source scenario to obtain source scenario experience data;
[0086] Based on the target scenario experience data, initialize the state prediction module and value estimation module in the preset Teacher Model;
[0087] Based on reinforcement learning and supervised learning, jointly train the state prediction module and value estimation module in the preset Teacher Model after initialization to obtain a trained Teacher Model;
[0088] Through the trained Teacher Model, initialize the state prediction module and value estimation module in the preset Student Model based on the source scenario experience data;
[0089] Based on reinforcement learning and supervised learning, jointly train the state prediction module and value estimation module in the preset Student Model after initialization to obtain a trained Student Model;
[0090] Use the trained Student Model as the trained robot navigation model.
[0091] Refer to Figure 3 , for the learning of the Teacher Model, in the embodiment of the present invention, first use the best mutual collision avoidance strategy to perform simulation demonstration in the target scenario, and then use the collected experience and data (memory pool ) to perform simulation learning, and initialize the value estimation module and state prediction module. Then, we use reinforcement learning and supervised learning to jointly train the value estimation module and state prediction module. The loss function of the value estimation module is the mean squared error (MSE) between the estimated value and the target value. The loss function of the state prediction module includes the mean squared error between the predicted state and the true state, as well as the KL divergence between the posterior distribution and the prior distribution .
[0092] The loss function of the value estimation module can refer to the following formula (1):
[0093] (1)
[0094] where represents the loss function of the value estimation module of the teacher model, represents the mean squared error function, represents the value estimation module of the teacher model after initialization, represents the estimated value of the teacher model, represents the target value.
[0095] The loss function of the state prediction module can refer to the following formula (2):
[0096] (2)
[0097] where represents the loss function of the state prediction module of the teacher model, represents the mean squared error function, represents the state prediction module of the teacher model after initialization, represents the predicted state of the teacher model (state prediction module), represents the true state, represents the posterior distribution of the teacher model and the KL divergence between the prior distribution of the teacher model
[0098] In some embodiments, the Optimal Reciprocal Collision Avoidance (ORCA) aims to ensure that multiple agents (such as robots and pedestrians) can move safely and efficiently in a shared space and avoid collisions.
[0099] During the simulation demonstration process, the interaction experience between the robot and the environment is collected, including information such as state, action, and reward, and this information is stored in the memory pool . These experiences will be used for subsequent simulation learning and model training.
[0100] Initialize the value estimation module using the experiences in the memory pool. The value estimation module is used to estimate the value or expected return of performing a specific action in a given state. Also initialize the state prediction module using the experiences in the memory pool. The state prediction module is used to predict the pedestrian state.
[0101] Optimize the value estimation module through reinforcement learning algorithms (such as Q-learning, Deep Q-Network, Policy Gradients, etc.). In this process, the mobile robot will try different actions and update its value estimation based on the obtained rewards.
[0102] Use the experiences in the memory pool to supervise the training of the state prediction module. This usually involves minimizing the difference between the predicted state and the true state, and ensuring that the posterior distribution of the model is consistent with the prior distribution.
[0103] For the training of the student module, initialize the student model using the teacher model trained in the previous process, and then train the value estimation module and the state prediction module through reinforcement learning and supervised learning in the source scenario. The value estimation module is still updated in a similar way to formula (1), and the loss function is as shown in the following formula (3). The state prediction module is updated under the guidance of the teacher model, as Figure 3 shown:
[0104] (3)
[0105] where, represents the loss function of the value estimation module of the student model, represents the mean squared error function, represents the value estimation module of the student model after initialization, represents the estimated value of the student model (value estimation module), represents the target value.
[0106] Specifically, obtain the hard error of the student model according to formula (4) :
[0107] (4)
[0108] where, represents the hard error, represents the mean squared error function, represents the predicted state of the student model (state prediction module), represents the true state, represents the posterior distribution of the student model and the KL divergence between the prior distribution of the student model.
[0109] According to the predicted state of the student model (state prediction module) and the predicted state of the teacher model (state prediction module) the mean squared error between them, as well as the prior distribution of the student model and the prior distribution of the teacher model the KL divergence between them, to obtain the soft error , specifically referring to the following formula (5):
[0110] (5)
[0111] where, represents the soft error, represents the mean squared error function, represents the predicted state of the student model (state prediction module), represents the predicted state of the teacher model (state prediction module), represents the prior distribution of the student model and the prior distribution of the teacher model the KL divergence between them.
[0112] The average of the hard error and the soft error is used as the final error for backpropagation, specifically referring to the following formula (6):
[0113] (6)
[0114] where, represents the final error, represents the hard error, represents the soft error.
[0115] In some embodiments, the KL divergence is used to measure the difference between two probability distributions.
[0116] In the embodiments of the present invention, both the teacher model and the student model are based on the robot navigation model. The teacher model represents a larger and more complex neural network, having strong performance and accuracy, but usually requiring high computational resources. The student model represents a smaller and simpler neural network, aiming to obtain performance close to that of the teacher model while having higher operating efficiency and less resource requirements.
[0117] The target scenario is the training environment or dataset where the mobile robot needs to navigate so that the mobile robot exhibits the target performance.
[0118] For example, the target scenario may contain various complex elements such as obstacles, pedestrians, traffic signals, etc., which require the mobile robot to perform accurate perception and judgment. The target scenario may involve different time periods (such as day and night), weather conditions (such as sunny and rainy days), and different geographical locations (such as urban streets and rural paths). The navigation tasks in the target scenario may be somewhat challenging, such as obstacle avoidance, path planning, pedestrian tracking, etc.
[0119] The source scenario is the actual running scenario or dataset of the mobile robot during subsequent operation.
[0120] Through the embodiments of the present invention, while reducing the dependence of the robot navigation model on the scenario, fully excavating the model's information understanding ability for different scenarios, improving the applicability of the model to various scenarios, it can effectively improve the performance of the navigation strategy model in different scenarios and ensure that the mobile robot can better handle the crowd navigation tasks under different scenarios.
[0121] The navigation device of the mobile robot provided by the present invention will be described below. The navigation device of the mobile robot described below can be mutually referred to corresponding to the navigation method of the mobile robot described above.
[0122] Reference Figure 4 , Figure 4 is a schematic diagram of the module structure of the navigation device of the mobile robot provided by the present invention.
[0123] An acquisition module 401, configured to acquire the pedestrian position information collected by the mobile robot, where the pedestrian position information includes: the position information sequence of each pedestrian;
[0124] A robot navigation module 402, configured to input the pedestrian position information into the robot navigation model to obtain the target navigation action output by the robot navigation model. The robot navigation model includes: a memory pool, a state prediction module, and a value estimation module, where, including:
[0125] Through the memory pool, within the current observation time window, the position information sequence of the target pedestrian before the target time step is used as historical motion information; the position information sequences of other pedestrians before the target time step are used as neighbor motion information; the position information sequence of the target pedestrian after the target time step is used as future motion information;
[0126] Through the state prediction module, based on the historical motion information, neighbor motion information, and future motion information, determine the predicted pedestrian state of the target pedestrian; take each pedestrian within the observation time window as the target pedestrian and determine the predicted pedestrian state of each pedestrian;
[0127] Through a value estimation module, obtain the robot state of the mobile robot, and determine the target navigation action of the mobile robot based on the robot state and the predicted pedestrian states of each pedestrian.
[0128] Specifically, the navigation device of the mobile robot provided by the present invention can implement all the method steps implemented by the above-mentioned navigation method embodiment of the mobile robot, and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiment will not be specifically described herein.
[0129] Figure 5 It is a schematic physical structure diagram of the electronic device provided by the present invention. As Figure 5 shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 complete mutual communication through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute the navigation method of the mobile robot. The method includes: obtaining the pedestrian position information collected by the mobile robot, where the pedestrian position information includes: the position information sequence of each pedestrian; inputting the pedestrian position information into the robot navigation model to obtain the target navigation action output by the robot navigation model. The robot navigation model includes: a memory pool, a state prediction module, and a value estimation module, where, including: through the memory pool, within the current observation time window, use the position information sequence of the target pedestrian before the target time step as historical motion information; use the position information sequences of other pedestrians before the target time step as neighbor motion information; use the position information sequence of the target pedestrian after the target time step as future motion information; through the state prediction module, based on the historical motion information, neighbor motion information, and future motion information, determine the predicted pedestrian state of the target pedestrian; use each pedestrian within the observation time window as the target pedestrian to determine the predicted pedestrian state of each pedestrian; through the value estimation module, obtain the robot state of the mobile robot, and determine the target navigation action of the mobile robot based on the robot state and the predicted pedestrian states of each pedestrian.
[0130] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0131] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the navigation method of the mobile robot provided by the above-mentioned various methods. The method includes: obtaining pedestrian position information collected by the mobile robot, where the pedestrian position information includes: a position information sequence of each pedestrian; inputting the pedestrian position information into a robot navigation model to obtain a target navigation action output by the robot navigation model. The robot navigation model includes: a memory pool, a state prediction module, and a value estimation module, where, including: through the memory pool, within the current observation time window, using the position information sequence of the target pedestrian before the target time step as historical motion information; using the position information sequences of other pedestrians before the target time step as neighbor motion information; using the position information sequence of the target pedestrian after the target time step as future motion information; through the state prediction module, based on the historical motion information, neighbor motion information, and future motion information, determining the predicted pedestrian state of the target pedestrian; taking each pedestrian within the observation time window as the target pedestrian and determining the predicted pedestrian state of each pedestrian; through the value estimation module, obtaining the robot state of the mobile robot, and based on the robot state and the predicted pedestrian state of each pedestrian, determining the target navigation action of the mobile robot.
[0132] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the navigation method of the mobile robot provided by the above-mentioned various methods. The method includes: obtaining the pedestrian position information collected by the mobile robot, where the pedestrian position information includes: the position information sequence of each pedestrian; inputting the pedestrian position information into the robot navigation model to obtain the target navigation action output by the robot navigation model. The robot navigation model includes: a memory pool, a state prediction module, and a value estimation module, where, including: through the memory pool, within the current observation time window, taking the position information sequence of the target pedestrian before the target time step as historical motion information; taking the position information sequences of other pedestrians before the target time step as neighbor motion information; taking the position information sequence of the target pedestrian after the target time step as future motion information; through the state prediction module, based on the historical motion information, neighbor motion information, and future motion information, determining the predicted pedestrian state of the target pedestrian; taking each pedestrian within the observation time window as the target pedestrian and determining the predicted pedestrian state of each pedestrian; through the value estimation module, obtaining the robot state of the mobile robot, and based on the robot state and the predicted pedestrian state of each pedestrian, determining the target navigation action of the mobile robot.
[0133] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0134] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A navigation method for a mobile robot, characterized in that: include: Acquire pedestrian position information collected by the mobile robot, wherein the pedestrian position information includes: a position information sequence of each pedestrian; The pedestrian position information is input into a robot navigation model to obtain a target navigation action output by the robot navigation model, wherein the robot navigation model includes: a memory pool, a state prediction module and a value estimation module, wherein: Through the memory pool, within the current observation time window, the position information sequence of the target pedestrian before the target time step is used as the historical motion information; the position information sequence of other pedestrians before the target time step is used as the neighbor motion information; and the position information sequence of the target pedestrian after the target time step is used as the future motion information; Determine the predicted pedestrian state of the target pedestrian based on the historical motion information, the neighbor motion information and the future motion information through the state prediction module; take each pedestrian in the observation time window as the target pedestrian and determine the predicted pedestrian state of each pedestrian; acquiring a robot state of the mobile robot through the value estimation module, and determining a target navigation action of the mobile robot based on the robot state and the predicted pedestrian state of each pedestrian; Before inputting the pedestrian position information into the robot navigation model to obtain the target navigation action output by the robot navigation model, the method further includes: The preset robot navigation model is trained based on knowledge distillation to obtain a trained robot navigation model; The training includes: teacher model training and student model training. The preset robot navigation model is trained based on knowledge distillation to obtain a trained robot navigation model, including: Based on the preset teacher model, a simulation is run in the target scenario based on the optimal mutual collision avoidance strategy to obtain the target scenario experience data; Run simulation in the source scene based on the preset student model to obtain source scene experience data; Initializing the state prediction module and the value estimation module in the preset teacher model based on the target scene experience data; Based on reinforcement learning and supervised learning, the state prediction module and the value estimation module in the preset teacher model after initialization are jointly trained to obtain a trained teacher model; Initializing the state prediction module and the value estimation module in the preset student model based on the source scene experience data through the trained teacher model; Based on reinforcement learning and supervised learning, the state prediction module and the value estimation module in the preset student model after initialization are jointly trained to obtain a trained student model; Using the trained student model as a trained robot navigation model; Among them, the loss function of the value estimation module of the teacher model can refer to the following formula: ; in, represents the loss function of the value estimation module of the teacher model, represents the mean square error function, represents the value estimation module of the teacher model after initialization, represents the estimated value of the value estimation module of the teacher model after the initialization, represents the target value; The loss function of the state prediction module of the teacher model can refer to the following formula: ; in, represents the loss function of the state prediction module of the teacher model, represents the mean square error function, represents the state prediction module of the teacher model after initialization, represents the predicted state of the state prediction module of the teacher model after the initialization, Indicates the real state, represents the posterior distribution of the teacher model and the prior distribution of the teacher model The KL divergence between .
2. The navigation method of a mobile robot according to claim 1, characterized in that: The memory pool includes: a first memory pool, a second memory pool and a third memory pool; The first memory pool is used to store the robot state and environment reward of the robot navigation model at each time step during the training process; The second memory pool is used to store the position information sequence of each pedestrian in the historical observation time window; The third memory pool is used to store the position information sequence of each pedestrian in the current observation time window.
3. The navigation method of a mobile robot according to claim 1, characterized in that: The state prediction module includes: a recurrent neural network and a decoder; Determining the predicted pedestrian state of the target pedestrian based on the historical motion information, the neighbor motion information, and the future motion information includes: By means of the recurrent neural network, the historical motion information, the neighbor motion information and the future motion information are encoded based on an attention mechanism to obtain a social feature vector of the target pedestrian; The social feature vector is decoded by the decoder to obtain the predicted pedestrian state of the target pedestrian.
4. The navigation method of a mobile robot according to claim 1, characterized in that: The value estimation module includes: a relationship graph network and a value network; The determining the target navigation action of the mobile robot based on the robot state and the predicted pedestrian state of each pedestrian includes: Constructing a feature matrix and a relationship matrix based on the robot state and the predicted pedestrian state of each pedestrian through the relationship graph network; Calling a graph convolutional network to encode based on the feature matrix and the relationship matrix to obtain interactive relationship features; Through the value network, the action value of the interactive relationship feature is estimated based on a preset action value function, and the action with the highest value in the action space is determined as the target navigation action.
5. A navigation device for a mobile robot, characterized in that: include: An acquisition module, used to acquire pedestrian position information collected by a mobile robot, wherein the pedestrian position information includes: a position information sequence of each pedestrian; A robot navigation module is used to input the pedestrian position information into a robot navigation model to obtain a target navigation action output by the robot navigation model, wherein the robot navigation model includes: a memory pool, a state prediction module and a value estimation module, including: Through the memory pool, within the current observation time window, the position information sequence of the target pedestrian before the target time step is used as the historical motion information; the position information sequence of other pedestrians before the target time step is used as the neighbor motion information; and the position information sequence of the target pedestrian after the target time step is used as the future motion information; Determine the predicted pedestrian state of the target pedestrian based on the historical motion information, the neighbor motion information and the future motion information through the state prediction module; take each pedestrian in the observation time window as the target pedestrian and determine the predicted pedestrian state of each pedestrian; acquiring a robot state of the mobile robot through the value estimation module, and determining a target navigation action of the mobile robot based on the robot state and the predicted pedestrian state of each pedestrian; Before inputting the pedestrian position information into the robot navigation model to obtain the target navigation action output by the robot navigation model, the navigation device of the mobile robot is further used to: The preset robot navigation model is trained based on knowledge distillation to obtain a trained robot navigation model; The training includes: teacher model training and student model training. The preset robot navigation model is trained based on knowledge distillation to obtain a trained robot navigation model, including: Based on the preset teacher model, a simulation is run in the target scenario based on the optimal mutual collision avoidance strategy to obtain the target scenario experience data; Run simulation in the source scene based on the preset student model to obtain source scene experience data; Initializing the state prediction module and the value estimation module in the preset teacher model based on the target scene experience data; Based on reinforcement learning and supervised learning, the state prediction module and the value estimation module in the preset teacher model after initialization are jointly trained to obtain a trained teacher model; Initializing the state prediction module and the value estimation module in the preset student model based on the source scene experience data through the trained teacher model; Based on reinforcement learning and supervised learning, the state prediction module and the value estimation module in the preset student model after initialization are jointly trained to obtain a trained student model; Using the trained student model as a trained robot navigation model; Among them, the loss function of the value estimation module of the teacher model can refer to the following formula: ; in, represents the loss function of the value estimation module of the teacher model, represents the mean square error function, represents the value estimation module of the teacher model after initialization, represents the estimated value of the value estimation module of the teacher model after the initialization, represents the target value; The loss function of the state prediction module of the teacher model can refer to the following formula: ; in, represents the loss function of the state prediction module of the teacher model, represents the mean square error function, represents the state prediction module of the teacher model after initialization, represents the predicted state of the state prediction module of the teacher model after the initialization, Indicates the real state, represents the posterior distribution of the teacher model and the prior distribution of the teacher model The KL divergence between .
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the navigation method of the mobile robot according to any one of claims 1 to 4 is implemented.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the navigation method of the mobile robot as claimed in any one of claims 1 to 4 is implemented.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the navigation method of the mobile robot as claimed in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Robot navigation method and system, robot and storage medium
CN110955242A
Pedestrian trajectory prediction method based on motion intention extraction
CN116433710A