Robot navigation using an advanced policy model and a trained

Through collaborative use of advanced policy models and low-level policy models, the problem of robot navigation in the existing technology underperformance in real environments is solved, efficient and secure navigation is achieved, and computing resources are saved.

CN119973981APending Publication Date: 2025-05-13GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510040279.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-11-29
Filing Date
2019-11-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing robot navigation systems are difficult to deploy successfully in real environments, mainly due to the high sample complexity of reinforcement learning algorithms, which leads to poor performance on real robots.

Method used

Robot navigation is used in collaborative approaches using advanced policy models and low-level policy models. The advanced strategy model is used to remotely plan and generate advanced actions, while the low-level strategy model is used to generate fine low-level actions, avoid obstacles and achieve efficient movement.

Benefits of technology

Efficient and secure robot navigation is achieved by eliminating the need for map generation, saving computing resources, and avoiding storing large amounts of map data on the robot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119973981A_ABST
    Figure CN119973981A_ABST
Patent Text Reader

Abstract

Mobile robot navigation is trained and / or performed using both a high-level policy model and a low-level policy model. An advanced output generated using the advanced policy model in each iteration indicates a corresponding advanced action of the robot motion when navigating to the navigation target. The low-level output generated in each iteration is based on a corresponding high-level action determined for the iteration and based on observation (s) of the iteration. The low-level policy model is trained to generate low-level outputs that define low-level action (s) that define robotic motion more finely than high-level actions-and generate obstacle-avoiding and / or efficient (e.g., range and / or time efficient) low-level action (s).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with an international filing date of November 27, 2019, Chinese application number 201980078868.4, and invention name “Robot navigation using high-level strategy models and trained low-level strategy models”. Background Art

[0002] Robot navigation is one of the fundamental challenges in robotics. To operate efficiently, various mobile robots need to navigate robustly in dynamic environments. Robot navigation is generally defined as finding a path from a starting location to a target location and executing that path in a robust and safe manner. Typically, robot navigation requires the robot to sense its environment, localize itself relative to the target, reason about obstacles within its immediate vicinity, and infer a long-range path to the target.

[0003] Traditionally, robot navigation systems rely on feature extraction and geometry-based inference to localize the robot and map its environment. When a map of the robot's environment is generated or given, the robot can use the map to find a navigation path using a planning algorithm.

[0004] Recently, training neural network policy models using reinforcement learning (RL) has become an option for robot navigation. Policy models trained using reinforcement learning with the help of robot experience data learn to associate raw observations with actions without the need for mapping or explicit path planning. However, various current policy models trained using RL are difficult to successfully deploy on real robots. The difficulty may be due to, for example, the high sample complexity of RL algorithms. Such high sample complexity means that neural network policy models can usually only be successfully trained in simulated environments. When implemented on a real robot, neural network policy models trained in simulated environments may fail and / or perform poorly. This may be due to, for example, images and / or other observation data captured by real sensors of the real robot being visually different from the simulated observation data used to train the neural network policy model. Summary of the invention

[0005] Embodiments disclosed herein relate to training and / or using both a high-level policy model and a low-level policy model for mobile robot navigation. For example, the high-level policy model and the low-level policy model can be used collaboratively to perform point-to-point navigation, wherein the mobile robot navigates from a current posture to a navigation target in an environment, such as a specific location in the environment, a specific object in the environment, or other navigation targets in the environment. Both the high-level policy model and the low-level policy model can be machine learning models, such as neural network models. In various embodiments, the high-level policy model is a recurrent neural network (RNN) model and / or the low-level policy model is a feedforward neural network model, such as a convolutional neural network (CNN) model.

[0006] The high-level policy model is used to generate a high-level output indicating which of a plurality of discrete high-level actions should be implemented to reach the navigation target based on the target label of the navigation target and based on (one or more) current robot observations (e.g., observation data). As a non-limiting example, the high-level actions may include "forward", "turn right" and "turn left". The low-level policy model is used to generate low-level action outputs based on (one or more) current robot observations (which may optionally be different from the current robot observations used to generate the high-level outputs) and optionally based on high-level actions selected based on the high-level outputs. The low-level action outputs define low-level actions that define the robot motion more finely than the high-level actions. As a non-limiting example, the low-level actions may define the corresponding angular velocity and the corresponding linear velocity of each of one or more wheels of the mobile robot. The low-level action outputs may then be used to control one or more actuators of the mobile robot to achieve the corresponding low-level actions. Continuing with the non-limiting example, control commands may be provided to one or more motors driving the (one or more) wheels so that the (one or more) wheels each achieve their corresponding angular velocity and linear velocity.

[0007] The high-level policy model and the low-level policy model are used in collaboration and are used in each iteration of multiple iterations during navigating the mobile robot to the navigation target - taking into account the new current observations of each iteration. The high-level output generated using the high-level policy model at each iteration indicates the corresponding high-level action of the robot movement when navigating to the navigation target. The high-level policy model is trained to be capable of long-range planning and is trained to generate corresponding high-level actions that attempt to move the mobile robot closer to the navigation target at each iteration. The low-level output generated at each iteration is based on the corresponding high-level action determined for the iteration and based on (one or more) observations for the iteration. The low-level policy model is trained to generate low-level outputs that define (one or more) low-level actions, which define the robot movement more finely than the high-level actions - and generate (one or more) low-level actions that avoid obstacles and / or are efficient (e.g., distance and / or time efficiency). The high-level and low-level policy models used separately but in collaboration enable the high-level policy model to be used to determine high-level actions that are guided by the deployment environment and attempt to move the mobile robot to the navigation target. However, the high-level actions determined using the high-level policy model cannot be used to accurately guide the robot. On the other hand, a low-level policy model can be used to generate low-level actions that can accurately guide the robot and efficiently and safely (e.g., avoid obstacles) achieve high-level actions. As described herein, in various embodiments, the low-level policy model is used to generate control commands for only a subset (e.g., one or more) of high-level actions, and for (one or more) high-level actions that are not in the subset, the low-level actions can be predefined or otherwise determined. For example, in an embodiment including "forward", "turn left", and "turn right" as candidate high-level actions, a low-level policy model can be used to generate a low-level action for the "forward" high-level action, and the corresponding fixed low-level actions are used for the "turn left" and "turn right" high-level actions.

[0008] High-level and low-level strategies can be used collaboratively to achieve efficient mobile robot navigation in an environment without relying on a map of the environment to find a navigation path using a planning algorithm. Therefore, navigation can be performed in an environment without generating a map and without referencing a map. Eliminating map generation can save various robot and computer resources that would otherwise be required to generate a detailed map of the environment. In addition, map-based navigation typically requires the storage of maps on the robot, which requires a large amount of storage space. Eliminating the need to reference a map in navigation can prevent the need to store maps in the limited storage resources of a mobile robot.

[0009] Various embodiments utilize supervised training to train a high-level policy model. For example, some of these various embodiments perform supervised training by using (one or more) real-world observations (e.g., images and / or other observation data) as at least a portion of the input to be processed by the high-level policy model during supervised training; and using a baseline truth navigation path in a real environment to generate a loss as a supervisory signal during supervised training. The baseline truth navigation path can be generated using (one or more) path planning algorithms (e.g., shortest path), can be based on human demonstrations of feasible navigation paths, and / or generated in other ways. Because supervised learning has lower sample complexity, it can achieve more efficient training compared to reinforcement training techniques. Therefore, at least compared to the enhancement technique, a smaller amount of resources (e.g., resources of (one or more) processors, memory resources, etc.) can be utilized in the supervised training techniques described herein. In addition, using real-world observations during training of the high-level policy model can improve the performance of the model on a real-world robot compared to using only simulated observations. For example, this may be due to the observations used in training being real-world observations that are visually similar to observations made on a real robot during model use. As described above, the supervised training method described in this paper can achieve efficient training of high-level policy models while leveraging real-world observations.

[0010] Various embodiments additionally or alternatively utilize reinforcement training to train the low-level policy model, and optionally utilize a robot simulator to perform the reinforcement training. In some of these various embodiments, the low-level policy model is trained by: utilizing (one or more) simulated observations and (one or more) high-level actions from the robot simulator as at least a portion of the input to be processed by the low-level policy model during reinforcement training; and utilizing simulated data from the robot simulator to generate rewards for training the low-level policy model. The rewards are generated based on a reward function (such as a reward function that penalizes robot collisions when reaching a navigation goal and rewards faster speeds and / or shorter distances). For example, the reward function can severely penalize motion that results in collisions, while rewarding collision-free motion based on how fast and / or straight the motion is.

[0011] In some embodiments, the (one or more) simulated observations used in the intensive training of the low-level policy model are simulated one-dimensional (1D) LIDAR component observations, simulated two-dimensional (2D) LIDAR component observations, and / or simulated proximity sensor observations. Such observations can be simulated with high fidelity in a simulated environment and can be better converted to real observations compared to, for example, RGB images. In addition, such observations can be simulated with high fidelity even if the simulated environment is simulated with relatively low fidelity. In addition, the physical properties of the robot can be simulated in a robot simulator, so that precise robot motion can be simulated through simple depth perception, which can achieve training of the low-level policy model to generate low-level actions that avoid obstacles and are efficient.

[0012] Therefore, various embodiments enable the use of simple depth observations (e.g., from 1D LIDAR, 2D LIDAR, and / or (one or more) proximity sensors) to train low-level policy models in a simulated environment. Using depth observations from real robots (and optionally without any RGB image observations), such low-level policy models can be effectively used on real robots to achieve safe and efficient low-level control of these robots. In addition, the low-level control generated using the low-level policy model is also based on the high-level actions determined using the high-level policy model. As described above, such high-level policy models can be trained using real-world observations (which may include RGB image observations and / or other higher fidelity observations) and supervised training. Through the collaborative use and training of both high-level policy models and low-level policy models, high-level actions can be determined using higher-fidelity real-world observations (as well as target labels and optional lower-fidelity observations), while low-level actions are determined using lower-fidelity real-world observations (as well as determined high-level actions). This can be achieved by separating the two models while training them collaboratively (e.g., by using high-level actions when training the low-level policy model, but not necessarily using high-level actions generated using the high-level policy model) and utilizing the two models collaboratively.

[0013] The above description is provided only as an overview of some embodiments disclosed herein. These and other embodiments will be described in more detail herein.

[0014] Other embodiments may include at least one transitory or non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., central processing unit(s) (CPUs), graphics processing units(s) (GPUs), and / or tensor processing units(s) (TPUs)) to perform methods such as one or more methods described above and / or elsewhere herein. Still other embodiments may include a system of one or more computers and / or one or more robots including one or more processors operable to execute stored instructions to perform methods such as one or more methods described above and / or elsewhere herein.

[0015] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are contemplated as part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 An example environment is shown in which implementations disclosed herein may be implemented.

[0017] Figure 2 is a flow chart illustrating an example method of training a high-level policy model according to various embodiments disclosed herein.

[0018] Figure 3 is a flow chart illustrating an example method for training a low-level policy model according to various embodiments disclosed herein.

[0019] Figure 4 is a flow chart illustrating an example method of utilizing a high-level policy model and a low-level policy model in navigating a mobile robot to a navigation goal.

[0020] Figure 5 An example architecture of a robot is schematically depicted.

[0021] Figure 6 An example architecture of a computer system is schematically depicted. DETAILED DESCRIPTION

[0022] Embodiments disclosed herein include a high-level policy model for long-range planning of mobile robot navigation. It is trained to generate a high-level action output based on the current robot observation and target label, and the high-level action output indicates the best high-level action of the mobile robot that enables the mobile robot to get closer to the navigation target. The best high-level action can be one of a plurality of discrete high-level actions of a defined high-level action space. As a non-limiting example, the discrete high-level actions can include or be limited to general navigation instructions such as "forward", "turn left" and "turn right". The embodiment also includes a low-level policy model, which is used to generate a low-level action output based on the best high-level action and the current robot observation (which can be the same or different from the observation used for the high-level policy model), and the low-level action output defines a corresponding low-level action that can be performed on the robot in a safe, robust and efficient manner. The low-level action can be a low-level action in a defined low-level action space (such as a continuous robot motion space). At the same time, the low-level action avoids obstacles in its vicinity, so it does not execute the high-level command verbatim.

[0023] The two policy models have complementary properties. The high-level policy model is trained given the deployment environment, allowing it to be used for planning in the deployment environment. However, the high-level policy model cannot be used to accurately guide the robot. The low-level policy model is not trained given the environment, generating low-level actions without understanding the environment. However, the low-level policy model can be used to accurately and safely move the robot.

[0024] Before referring to the drawings, an overview of a specific embodiment of the technology disclosed herein is provided. It should be understood that the disclosure herein is not limited to such embodiments, and additional embodiments are disclosed herein (e.g., in the Summary of the Invention, the remainder of the Detailed Description, and the Claims).

[0025] In some embodiments, the high-level policy model takes as input: observation data, such as an RGB image x (or its embedding) and a binary proximity indicator p∈{0,1}; and a target label, such as a one-hot vector g∈{0,1} k, which represents one of k possible target locations in the environment. The proximity indicator can be, for example, the output of a radar reading, and can indicate whether a collision is imminent. For example, it can be defined as 1 if there is an object within 0.3m or other threshold, and 0 otherwise. As described above, when the target label is a one-hot vector, the one-hot value in the vector has semantic meaning (e.g., different one-hot values ​​are used for different navigation targets), but does not necessarily have any relevance to the deployment environment for which the high-level policy model is trained. Additional or alternative target labels can be utilized in various embodiments, such as word-embedded target labels as semantic descriptors of the navigation target, image-embedded target labels as images of the navigation target, and / or other target labels that provide semantic meaning of the navigation target.

[0026] In some embodiments, the output action space of the high-level policy model is defined as and includes multiple discrete actions, such as forward, left turn, and / or right turn. Forward motion may be intended to be, for example, 1 meter; and turn may be intended to be, for example, fifteen degrees. However, note that one or more of these values ​​are approximate (e.g., at least forward motion) because their semantics are established, for example, during training of a low-level policy model.

[0027] Using the above notation, the high-level policy model is trained to output a value v(a, x; g) that estimates progress toward goal g, defined as the negative change in distance to g if action a is taken when observing x. This value function can be used to estimate which action moves the robot closest to the goal:

[0028]

[0029] The above value function can be implemented using a recurrent neural network (RNN) that takes as input the concatenation and transformation of the observation x, the target label g, and the proximity p: υ(a,x;g)=LSTM(MLP2(ResNet50(x),p,MLP1(g)). The RNN can be, for example, a single-layer long short-term memory (LSTM) network model or other memory network model (e.g., a gated recurrent unit (GRU)). The image embedder can be a neural network model for processing an image and generating a dense (relative to pixel size) embedding of the image, such as a ResNet50 network. The target label g can be, for example, a one-hot vector at k possible positions, and / or other target labels, such as the target labels described above. The MLP in the preceding notation l represents an l-layer perceptron with ReLU. The size of the above perceptron and LSTM network models can be set to, for example, 2048 or other values.

[0030] Some (certain) actions in may potentially be performed literally without any risk of collision. For example, "rotate left" and "rotate right" may be performed without any risk of collision. Therefore, in some embodiments, for such actions, they may optionally be implemented using the corresponding default low-level actions defined for the corresponding specific high-level actions. However, Other actions (one or more) in the process may potentially cause a collision, such as a "forward" action. A separate low-level policy model may be optionally trained and used to perform such actions (one or more) (e.g., a "forward" action).

[0031] The inputs to the low-level policy model can be, for example, 1D LIDAR readings, 2D LIDAR readings, and / or proximity sensor readings. Such (one or more) readings, while low fidelity (e.g., compared to 3D LIDAR and / or RGB images), are able to capture obstacles, which is sufficient for short-term safe control. Low-level action space can be continuous and can optionally be defined by the kinematics of the robot. As a non-limiting example, for a differential drive mobile base, the action space It can be a 4-dimensional real-valued vector of the torsion values ​​of the two wheels (the linear and angular velocity of each wheel).

[0032] In some embodiments, the low-level policy model can be a convolutional neural network (CNN) that can process the last n LIDAR readings (and / or other readings) as input, where n can be greater than 1 in various embodiments. For example, the last three readings x, x can be processed. -1 ,x -2 , and since they are optionally one-dimensional (e.g., 1D LIDAR or proximity sensors), they can be concatenated into images, where the second dimension is time. The output generated using the low-level policy model can be a value in the above low-level action space. More formally, the generated low-level action a low It can be expressed as:

[0033] a low =ConvNet(concat(x -2 ,x -1 ,x))

[0034] in, And where ConvNet is a CNN model, for example a CNN model with the following 4 layers: conv([7,3,16],5)→conv([5,1,20],3)→fc(20)→, where conv(k,s) represents a convolution with kernel k and stride s, and fc(d) is a fully connected layer with output dimension d.

[0035] In some embodiments, the training of the high-level policy model can utilize real images X from the deployed world obtained by traversal. Images can be captured, for example, by (one or more) monocular cameras (e.g., RGB images), (one or more) stereo cameras (e.g., RGBD images), and / or other higher fidelity visual components. These images represent the state of the robot in the world and can be organized in the form of a graph, with edges representing actions that move the robot from one state to another. In some of these embodiments, the images are based on images captured via an rig of six cameras (or other visual components) organized in a hexagonal shape. The rig moves along an environment (e.g., corridors and spaces) and captures a set of images every 1m (or other distance). The rig can be mounted to, for example, a mobile robot base that can be optionally manually guided along the environment, and / or can be mounted to a human and / or non-robot base that guides along the environment.

[0036] After the images are captured, they can optionally be stitched into 360 degree panoramas, which can be cropped in any direction to obtain images of the desired field of view (FOV). This can allow the creation of observations with the same properties (e.g., FOV) as a robot camera. For example, a FOV of 108 degrees and 90 degrees along the width and height, respectively, can be used to mimic a robot camera with the same FOV. Each panorama can be cropped every X degrees to obtain Y separate images. For example, each panorama can be cropped every 15 degrees to obtain 24 separate images. In addition, edges can be defined between images, where the edges represent actions. For example, two rotation actions "turn left" and "turn right" can be represented, which move the robot to the next left or right image, respectively, at the same location.

[0037] The pose of the image can also be estimated and assigned to the image. For example, the Cartographer Positioning API and / or other (one or more) techniques can be used to estimate the pose of the image. The estimation of the pose can be based only on locally correct SLAM and loop closure. Therefore, the high accuracy required for a global geometric map is not required, and further, no mapping of the surrounding environment is required.

[0038] Action(s) may also be defined between images from different panoramas. For example, a "forward" action may be defined as an edge between two images, where the "forward" action is from the current image to a nearby image by moving in the direction of the current view. A nearby image may be an image that is ideally a fixed distance (e.g. 1.0 m) from the current view. However, there is no guarantee that the image has been captured at that new position. Therefore, if there are images captured within a certain range of fixed distances (e.g. from 0.7 m to 1.0 m), an action may still be considered possible (and the corresponding images utilized).

[0039] Images organized as graphs, whose edges represent actions that move a robot from one state to another, can have relatively high visual fidelity. Furthermore, traversals defined by the graph can cover most of the specified space of a deployment environment. However, high-level actions capture coarse motion. Therefore, they can be used to express navigation paths but cannot be robustly executed on a robot. Therefore, in various embodiments, images and high-level actions are only used to train a high-level policy model.

[0040] Unlike recent reinforcement learning (RL) based methods, the training for training high-level policy models can be formulated as a supervised learning problem. For goal-driven navigation, optimal paths can be generated (and used as supervisory signals) by employing a shortest path algorithm (and / or other path optimization algorithms) or by artificial demonstrations of feasible navigation paths. These paths can be used as supervision at each step of policy execution (when present). Since supervised learning has lower sample complexity, it has an advantage over RL in terms of efficiency.

[0041] To define the training loss, consider a set of navigation paths P = {p1, ..., p N}. These paths can be defined on a graph that organizes the image. It can be the set of all shortest paths to the target generated by the shortest path planner.

[0042] For a target g, a starting state x (e.g., a starting image), and a path d(x,g;p) represents the distance from x to g along p if both the source and target are on the path in that order. If one or both are not on the path, the distance is infinite. From x to g The shortest path in can be thought of as:

[0043]

[0044] Using d, if we apply a high-level action a in state x, we can define progress toward the goal as:

[0045]

[0046] Among them, x ′ is the image reached after taking action a.

[0047] The loss trains the high-level policy model to generate an output value as close to y as possible. In many embodiments, the RNN model is used as the high-level policy model, and the loss is defined on the entire navigation path. If the navigation path is represented as x = (x1, ..., x T ), the loss can be expressed as:

[0048]

[0049] Wherein, the model υ can be, for example, v(a,x;g)=LSTM(MLP2(ResNet50(x),p,MLP1(g)) defined above. Stochastic gradient descent can optionally be used to update the RNN model based on the loss, where, at each step of training, a navigation path can be generated and the above loss is formulated to perform gradient updates. These paths are generated using the current high-level policy model and a random starting point. At the beginning of training, using the high-level policy model results in random actions being performed, so the navigation path is random. As training progresses, the navigation path becomes more meaningful, and the above loss focuses on the situations that will be encountered during inference.

[0050] In various embodiments, the low-level policy model is trained in one or more synthetic environments, such as a synthetic environment including several corridors and rooms. (One or more) synthetic environments can be generated using a 2D layout that can be elevated in 3D by extending the walls upward. (One or more) synthetic environments may be different from the deployment environment optionally used in the training of the high-level policy model. The observations used in training the low-level policy model may be relatively low-fidelity observations, such as 1D depth images, 2D depth images, and / or other low-fidelity observations. Due to their simplicity, these observations, although low in fidelity relative to the observations used in training the high-level policy model, can be simulated with high fidelity and the trained model transferred to a real robot. In addition, the physics of the robot can be simulated in a simulated environment using a physics engine (such as the PyBullet physics engine). Therefore, precise robot motion can be simulated by simple depth perception, which is sufficient to train low-level obstacle avoidance control that can be transferred to the real world.

[0051] In various embodiments, continuous deep Q-learning (DDPG) is used to train the low-level policy model. For example, the policy may be to perform a "go forward" action without colliding with an object. With such a policy and for a robot with differential drive, the reward R(x, a) required by DDPG for a given action a in state x may be highest if the robot moves straight as fast as possible without colliding:

[0052]

[0053] Among them, v lin (a) and v ang (a) shows the linear and angular velocities of the differential actuator after applying the current action (in the current state omitted for brevity). If the action does not result in a collision, the reward depends on how fast the robot moves (R lin =1.0) and straightness (R ang =-0.8). If there is a collision, the robot suffers a large negative reward R collision = -1.0. Whether a collision exists can be easily determined in a simulation environment.

[0054] The employed DDPG algorithm may utilize a critic network that approximates the Q-value for a given state x and action a.

[0055] Now turning to the attached figure, Figure 1 An example environment is shown in which embodiments disclosed herein may be implemented. Figure 1 Included are low-level policy models 156 and high-level policy models 154. High-level policy models 154 may be trained by high-level policy trainer 124. As described herein, high-level policy trainer 124 may utilize supervised training data 152 and supervised learning when training high-level policy models.

[0056] The low-level policy model 156 can be trained by the low-level policy trainer 126 (which can use the DDPG algorithm). When training the low-level policy model using reinforcement learning, the low-level policy trainer 126 can interact with the simulator 180, which simulates the simulated environment and the simulated robot interacting in the simulated environment.

[0057] Robot 110 is also Figure 1 , and is an example of a physical (i.e., real-world) mobile robot that can utilize high-level policy models and low-level policy models trained according to embodiments disclosed herein when performing a robotic navigation task. Additional and / or alternative robots may be provided, such as Figure 1The robot 110 shown is an additional robot that differs in one or more aspects. For example, a mobile forklift robot, an unmanned aerial vehicle ("UAV"), and / or a humanoid robot may be used instead of or in addition to the robot 110.

[0058] The robot 110 includes a base 113 having wheels 117A, 117B disposed on opposite sides thereof for moving the robot 110. For example, the base 113 may include one or more motors for driving the wheels 117A, 117B of the robot 110 to achieve a desired direction, speed, and / or acceleration of movement of the robot 110.

[0059] The robot 110 also includes a vision component 111 that can generate observation data related to the shape, color, depth, and / or other features of (one or more) objects within the line of sight of the vision component 111. The vision component 111 can be, for example, a monocular camera, a stereo camera, and / or a 3D LIDAR component. The robot 110 also includes an additional vision component 112, which can generate observation data related to the shape, color, depth, and / or other features of (one or more) objects within the line of sight of the vision component 112. The vision component 112 can be, for example, a proximity sensor, a one-dimensional (1D) LIDAR component, or a two-dimensional (2D) LIDAR component. In various embodiments, the vision component 111 generates higher fidelity observations (relative to the vision component 112).

[0060] The robot 110 also includes one or more processors that, for example, implement the high-level engine 134 and the low-level engine 136 (described below) and provide control commands to actuators and / or other operating components of the actuators based on low-level actions generated using the low-level policy model (and based on outputs generated using the high-level policy model 154). The robot 110 also includes robot arms 114A and 114B with corresponding end effectors 115A and 115B, each of which is in the form of a gripper with two opposing "fingers" or "digits". Although specific gripping end-effectors 115A, 115B are shown, additional and / or alternative end-effectors may be utilized, such as alternative impact-type gripping end-effectors (e.g., impact-type gripping end-effectors with gripping “plates,” impact-type gripping end-effectors with more or fewer “digital fingers” / “claws”); “intrusive” gripping end-effectors; “retractive” gripping end-effectors; or “contact” gripping end-effectors or non-grasping end-effectors. Additionally, although in Figure 1 A particular placement of vision components 111 and 112 is shown in FIG. 1 , but additional and / or alternative placements may be utilized.

[0061] As described above, the processor(s) of the robot 110 may implement the high-level engine 134 and the low-level engine 136, which once trained, operate using the corresponding high-level policy model 154 and the low-level policy model 156. The high-level engine 134 may process the observation data 101 and the target label 102 using the high-level policy model 154 to generate the high-level action 103. The observation data 101 may include, for example, the current observation from the vision component 111 (and optionally the current observation from the vision component 112 and / or other (one or more) sensors). The target label 102 may be, for example, a one-hot vector, a word embedding of a semantic descriptor of a navigation target, an image embedding of an image of the navigation target, and / or other target labels that provide semantic meaning of the navigation target.

[0062] The low-level engine 136 processes the high-level action 103 and the additional observation data 104 using the low-level policy model 156 to generate a low-level action 105. The additional observation data 104 can be, for example, the current observation from the vision component 112. The low-level action 105 is provided to the control engine 142, which can also be implemented by the (one or more) processors of the robot 110, which generates corresponding control commands 106 that are provided to the (one or more) actuators 144 to cause the robot 110 to implement the low-level action 105. This process can continue, each time relying on new current observation data 101 and new current additional observation data 104, until the navigation target is reached. By continuous execution, navigation of the robot 110 to the target corresponding to the target tag 102 can be achieved.

[0063] Now go to Figure 2 , a flowchart is provided that illustrates an example method 200 for training a high-level policy model according to various embodiments disclosed herein. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. The system may include one or more components of one or more computer systems, such as one or more processors. In addition, although the operations of method 200 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0064] At block 202 , the system begins high-level policy model training.

[0065] At block 204, the system generates a target label for the navigation target. For example, the target label can be a semantically meaningful one-hot vector or other target labels described herein.

[0066] At block 206, the system selects real observation data for a starting pose. For example, the real observation data may include a real RGB image from a deployment environment in which the advanced policy model will be deployed.

[0067] At block 208, the system generates a corresponding value for each of the N high-level actions based on processing the real observation data and the target label using the high-level policy model. For example, the system can generate a first metric for the forward action, a second metric for the right turn action, and a third metric for the left turn action.

[0068] At block 210 , the system selects the action with the best corresponding value from the corresponding values ​​generated at block 208 .

[0069] At block 212, the system selects new real observation data for the new pose after the selected action is performed. For example, the system can select a new real image based on the edge of the graph of the new real image being organized by the selected action as being related to the observation data of block 206. For example, for a right turn action, an image from the same location but to the right by X degrees can be selected at block 412. In addition, for example, for a forward action, an image 1 meter away from the image of the observation data of block 206 and along the same direction as the image of the observation data of block 206 can be selected.

[0070] At block 214, the system generates and stores a baseline truth value for the selected action. The system can generate the baseline truth value based on a comparison of: (A) the distance from the previous pose to the navigation target along a baseline truth path (e.g., the shortest path from an optimizer or a human demonstration path); and (B) the distance from the new pose to the navigation target along the baseline truth path. In an initial iteration of block 214, the previous pose will be the starting pose. In future iterations, the previous pose will be the new pose determined in the iteration of block 212 immediately preceding the most recent iteration of block 212.

[0071] At block 216 , the system generates corresponding values ​​for each of the N actions based on processing the new real observation data and the target labels using the high-level policy model.

[0072] At block 218, the system selects the action with the best corresponding value. The system then proceeds to block 220 and determines whether to continue the current supervision episode. If so, the system returns to block 212 and performs another iteration of blocks 212, 214, 216, and 218. The system may determine to continue the current supervision episode if the goal specified by the navigation goal has not been reached, if a threshold amount of iterations of blocks 212, 214, 216, and 218 have not been performed, and / or if other criteria have not been met.

[0073] If at block 220 the system determines not to continue the current supervision episode (e.g., the navigation goal has been reached), the system proceeds to block 222. At block 222, the system generates a loss based on a comparison of: (A) the generated value of the selected action (generated at block 208 and at iteration(s) of block 216); and (B) the generated ground truth value (generated at iteration(s) of block 214).

[0074] At block 224 , the system then updates the high-level policy model based on the losses.

[0075] At block 226, the system determines whether training of the high-level policy model is complete. If not, the system returns to block 204 and performs another iteration of blocks 204-224. If so, the system proceeds to block 228, and training of the high-level policy model ends. The determination of block 226 may be based, for example, on whether a threshold amount of episodes have been executed and / or other factor(s).

[0076] Figure 3 300 is a flowchart illustrating an example method 300 for training a low-level policy model according to various embodiments disclosed herein. For convenience, the operations of the flowchart are described with reference to a system performing the operations. The system may include one or more components of one or more computer systems, such as one or more processors. In addition, although the operations of method 300 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0077] At block 302 , the system begins low-level policy model training.

[0078] At box 304, the system obtains the current high-level action and the current simulated observation. The current high-level action can be, for example, a forward action and / or other (one or more) high-level actions for which the low-level policy model is being trained. The current simulated observation can be a simulated observation from a simulated 1D LIDAR component, a simulated 2D LIDAR component, and / or other simulated components. Optionally, at box 304, the system also obtains N previous simulated observations, such as the two last simulated observations (in addition to the current simulated observation).

[0079] At block 306, the system processes the current high-level action and the current simulated observation using the low-level policy model to generate a low-level action output defining a low-level robot action. In some embodiments, the system also processes N previous observations (if any), such as the two last observations (in addition to the current observation), when generating the low-level action output.

[0080] At block 308, the system controls the simulated robot based on the low-level robot actions. The simulated robot may be controlled in a simulator that simulates the robot using a physics engine and also simulates the environment.

[0081] At block 310, the system determines a reward based on simulation data obtained after the robot is simulated based on the low-level robot motion control. The reward may be determined based on a reward function, such as a reward function that penalizes robot collisions while rewarding faster speeds and / or shorter distances when reaching a navigation goal. For example, a reward function may heavily penalize motion that results in a collision while rewarding collision-free motion based on the speed and / or straightness of the motion.

[0082] At block 312, the system updates the low-level policy model based on the reward. In some implementations, block 312 is performed after each iteration of block 310. Although for simplicity, Figure 3 304, 306, 308, and 310. In these other embodiments, the low-level policy model is updated based on the rewards from the multiple iterations. For example, in these other embodiments, multiple iterations of blocks 304, 306, 308, and 310 may be performed during a simulation episode (or during multiple simulation episodes). For example, when multiple iterations of blocks 304, 306, 308, and 310 are performed during a simulation episode, the current simulated observation at a non-initial iteration of block 304 may be the simulated observation resulting from the most recent iteration of performing block 308—and the last (one or more) observations optionally processed at block 306 may be the current observation of the most recent (one or more) previous iteration of block 304. Multiple iterations of blocks 304, 306, 308, and 310 may be performed iteratively during a simulation episode until one or more conditions occur, such as a threshold amount of iterations, a collision of the simulated robot with an environmental object (determined in block 310), and / or other (one or more) conditions. Thus, in various implementations, block 312 may be performed in a batch manner, and the model may be updated based on multiple rewards determined during consecutive simulation episodes.

[0083] At block 314, the system determines whether training of the low-level policy model is complete. If not, the system returns to block 304 and performs another iteration of blocks 304-312. If so, the system proceeds to block 316, and training of the low-level policy model ends. The determination of block 314 may be based, for example, on whether a threshold amount of episodes have been executed and / or other factor(s).

[0084] Figure 44 is a flow chart illustrating an example method 400 for using a high-level policy model and a low-level policy model when navigating a mobile robot to a navigation target. For convenience, the operations of the flow chart are described with reference to a system that performs the operations. The system may include one or more components of one or more computer systems, such as one or more processors of a robot. In addition, although the operations of method 400 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0085] At block 402, the system begins robotic navigation.

[0086] At box 404, the system identifies a target label for a navigation target in the environment. The target label can be a semantically meaningful one-hot vector, a word embedding of a semantic descriptor of the navigation target, a target label that is an image embedding of an image of the navigation target, and / or other target labels that provide semantic meaning for the navigation target. The target label can be generated based on user interface input and / or based on output from a higher-level task planner that identifies the navigation target. For example, a target label of "trash can" can be generated based on a spoken user interface input of "navigate to trash can." For example, the target label can be based on an image of "trash can" that is identified based on a spoken user interface input and / or based on a word embedding of "trash can."

[0087] At block 406, the system obtains current observation data based on outputs from the robot component(s). For example, the current observation data may include a current image captured by a camera of the robot, and optionally a current proximity sensor reading of a proximity sensor of the robot.

[0088] At block 408, the system processes the current observation data and the target label using the trained high-level policy model to generate a high-level action output. For example, the high-level action output may include a corresponding metric for each of the N individual high-level actions.

[0089] At block 410, the system selects a high-level action based on the high-level action outputs. For example, the system can select the high-level action with the "best" metric (e.g., the highest metric when a higher metric indicates the best high-level action).

[0090] At block 412, the system determines whether the high-level action can be implemented without utilizing the low-level policy model. For example, an action (or actions) such as "turn left" or "turn right" can optionally be implemented without utilizing the low-level policy model, while other actions (or actions) such as "go forward" require utilizing the low-level policy model.

[0091] If at box 412, the system determines that the high-level action can be achieved without using the low-level policy model, the system proceeds to box 412 and selects a low-level action for the high-level action. For example, if the high-level action is "turn right", the default low-level action of "turn right" can be selected.

[0092] If at box 412, the system determines that the high-level action cannot be performed without utilizing the low-level policy model, the system proceeds to box 414 and uses the trained low-level policy model to process the current additional observations to generate a low-level action output that defines the low-level action. For example, if the high-level action is "forward", the current additional observation data (optionally, as well as the previous N additional observation data instances) can be processed to generate a low-level action output that defines the low-level action. When generating the low-level action output, the high-level action of "forward" can also be optionally processed together with the current additional observation data. For example, the additional observation data can include depth readings from the robot's 1D LIDAR component. Although referred to as "additional" observation data in this article, in various embodiments, when generating the high-level action output, the current additional observation data of box 412 can also be processed at box 408 together with other current observation data.

[0093] At block 416 , the system controls the actuator(s) of the mobile robot to cause the mobile robot to implement the low-level action of block 412 or block 414 .

[0094] At block 418, the system determines whether the navigation target indicated by the target tag has been reached. If not, the system returns to block 406 and performs another iteration of blocks 406-416 using the new current observation data. If so, the system proceeds to block 420, and navigation to the navigation target ends. Another iteration of method 400 may be performed in response to identifying a new navigation target in the environment.

[0095] Figure 5 An example architecture of a robot 525 is schematically depicted. The robot 525 includes a robot control system 560, one or more operating components 540a-540n, and one or more sensors 542a-542m. The sensors 542a-542m may include, for example, vision sensors, light sensors, pressure sensors, pressure wave sensors (e.g., microphones), proximity sensors, accelerometers, gyroscopes, thermometers, barometers, etc. Although the sensors 542a-542m are depicted as being integrated with the robot 525, this is not meant to be limiting. In some embodiments, the sensors 542a-542m may be located external to the robot 525, such as as a stand-alone unit.

[0096] The operating components 540a-540n may include, for example, one or more end effectors and / or one or more servomotors or other actuators to achieve movement of one or more components of the robot. For example, the robot 525 may have multiple degrees of freedom, and each actuator may control actuation of the robot 525 within one or more degrees of freedom in response to a control command. As used herein, the term actuator includes a mechanical or electrical device (e.g., a motor) that produces motion, in addition to any (one or more) drivers that may be associated with the actuator and convert a received control command into one or more signals for driving the actuator. Thus, providing a control command to an actuator may include providing a control command to a driver that converts the control command into an appropriate signal for driving an electrical or mechanical device to produce the desired motion.

[0097] The robot control system 560 can be implemented in one or more processors (such as a CPU, GPU, and / or other controller(s)) of the robot 525. In some embodiments, the robot 525 can include a "brainbox" that can include all or various aspects of the control system 560. For example, the brainbox can provide real-time data bursts to the operating components 540a-540n, where each real-time burst includes a set of one or more control commands that indicate, among other things, motion parameters (if any) for each of one or more of the operating components 540a-540n. In some embodiments, the robot control system 560 can perform one or more aspects of the method 400 described herein.

[0098] As described herein, in some embodiments, all or various aspects of the control commands generated by the control system 560 when performing a robotic task can be based on the utilization of the trained low-level and high-level policy models described herein. Figure 5 525 as an integral part of the robot 525, but in some embodiments, all or various aspects of the control system 560 may be implemented in a component that is separate from but in communication with the robot 525. For example, all or various aspects of the control system 560 may be implemented on one or more computing devices (such as the computing device 610) that are in wired and / or wireless communication with the robot 525.

[0099] Figure 66 is a block diagram of an example computing device 610 that can be optionally used to perform one or more aspects of the techniques described herein. The computing device 610 typically includes at least one processor 614 that communicates with a plurality of peripheral devices via a bus subsystem 612. These peripheral devices may include a storage subsystem 624 (e.g., including a memory subsystem 625 and a file storage subsystem 626), a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices allow a user to interact with the computing device 610. The network interface subsystem 616 provides an interface to an external network and is connected to corresponding interface devices in other computing devices.

[0100] The user interface input device 622 may include a keyboard, a pointing device (such as a mouse, trackball, touch pad, or graphic tablet), a scanner, a touch screen incorporated into a display, an audio input device (such as a voice recognition system, a microphone), and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways to input information into the computing device 610 or over a communication network.

[0101] User interface output device 620 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, for example, via an audio output device. Typically, the use of the term "output device" is intended to include all possible types of devices and approaches to output information from computing device 610 to a user or another machine or computing device.

[0102] The storage subsystem 624 stores programs and data structures that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 624 may include logic to perform selected aspects of one or more methods described herein.

[0103] These software modules are typically executed by the processor 614 alone or in combination with other processors. The memory 625 used in the storage subsystem 624 may include multiple memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution and a read-only memory (ROM) 632 in which fixed instructions are stored. The file storage subsystem 626 can provide persistent storage for program and data files, and can include a hard drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of certain embodiments may be stored by the file storage subsystem 626 in the storage subsystem 624, or in other machines accessible to the processor(s) 614.

[0104] The bus subsystem 612 provides a mechanism for the various components and subsystems of the computing device 610 to communicate with each other as intended. Although the bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple busses.

[0105] The computing device 610 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, for the purpose of illustrating some embodiments, the computing device 610 is described in detail below. Figure 6 The description of computing device 610 depicted in FIG. 6 is intended only as a specific example. Figure 6 Many other configurations of computing device 610 may have more or fewer components than the computing device depicted in FIG.

[0106] In some embodiments, a method for navigating a mobile robot in an environment is provided, the method comprising identifying a target label of a navigation target in the environment, and navigating the mobile robot to the navigation target. Navigating the mobile robot to the navigation target comprises: in each of multiple iterations during navigation to the navigation target: obtaining corresponding current observation data, the corresponding current observation data is based on the corresponding current output of a sensor component of the mobile robot; processing the corresponding current observation data and the target label using a trained high-level policy model to generate a corresponding high-level action output; using the corresponding high-level action output to select a corresponding specific high-level action from a plurality of discrete high-level actions in a defined high-level action space; obtaining corresponding current additional observation data, the corresponding current additional observation data is based on the corresponding current additional output of an additional sensor component of the mobile robot; processing the corresponding current additional observation data and the corresponding specific high-level action using a trained low-level policy model to generate a corresponding low-level action output; and controlling one or more actuators of the mobile robot based on the corresponding low-level action output so that the mobile robot implements the corresponding low-level action. The corresponding low-level action output defines a corresponding low-level action of a defined low-level action space, and the defined low-level action space defines the robot motion more finely than the high-level action space.

[0107] These and other embodiments may include one or more of the following features. The discrete high-level actions of the defined high-level action space may not have any definition of one or more parameters of the robot motion defined in the low-level action space. The discrete high-level actions of the defined high-level action space may not have any definition of any speed of the robot motion, and the low-level action space may define one or more speeds of the robot motion. The low-level action space may be a continuous action space. Each corresponding low-level action may define one or more corresponding linear speeds and / or one or more corresponding angular speeds. For example, the mobile robot may include a first wheel, and each of the corresponding low-level actions may define at least one corresponding linear speed of one or more corresponding linear speeds of the first wheel. The sensor component may be a camera and / or the additional sensor component may be a proximity sensor, a one-dimensional (1D) LIDAR component, or a two-dimensional (2D) LIDAR component. The sensor component may be a camera, each corresponding current output may be a corresponding current image, and each corresponding current observation data may be a corresponding embedding of the corresponding current image, which is generated by processing the current image using an image embedding model. When generating each corresponding high-level action output, the trained high-level policy model may also be used to process the corresponding additional observation data and the corresponding current observation data and the target label. When generating each corresponding low-level action output, the corresponding current observation data may not be processed using a trained low-level policy model. The trained high-level policy model may be a recurrent neural network (RNN) model and / or may be trained using supervised learning. The trained low-level policy model may be trained using reinforcement learning. For example, the trained low-level policy model may be trained using a reward signal, which is generated based on an output from a robot simulator that simulates the navigation of a simulated robot in a simulated environment. When generating each corresponding low-level action output, the corresponding current additional observation data from one or more immediately preceding iterations in the iteration may also be processed together with the corresponding current additional observation data. The target label may include a unique hot vector having a unique hot value, which is assigned based on the location of a navigation target in the environment, the classification of an object, or the embedding of an object image. In each of the multiple iterations, the method may further include: determining that the corresponding specific high-level action is a specific high-level action that can cause a collision; and in response to determining that the specific high-level action is a specific high-level action that can cause a collision, the corresponding current additional observation data and the corresponding specific high-level action may be processed using a trained low-level policy model to generate a corresponding low-level action output.

[0108] In some embodiments, a method for navigating a mobile robot in an environment is provided, the method comprising identifying a target label of a navigation target in the environment, and navigating the mobile robot to the navigation target. Navigating the mobile robot to the navigation target comprises: in each of each iteration during navigation to the navigation target: obtaining corresponding current observation data, the corresponding current observation data is based on the corresponding current output from the sensor component of the mobile robot; processing the corresponding current observation data and the target label using a trained high-level policy model to generate a corresponding high-level action output; using the corresponding high-level action output to select a corresponding specific high-level action from a plurality of discrete high-level actions in a defined high-level action space; determining whether the corresponding specific high-level action is a specific high-level action that can cause a collision; when determining that the corresponding specific high-level action is not a specific high-level action that can cause a collision: controlling one or more actuators of the mobile robot based on a corresponding default low-level action defined for the corresponding specific high-level action; and when determining that the corresponding specific high-level action is a specific high-level action that can cause a collision: generating a corresponding low-level action output using a trained low-level policy model, the corresponding low-level action output being based on the high-level action and optimized according to the low-level policy model to reach the navigation target as quickly as possible without collision.

[0109] In some embodiments, a method for training a high-level policy model and a low-level policy model for collaborative use in the automatic navigation of a mobile robot is provided. The method includes performing supervised training of the high-level policy model to train the high-level policy model to generate a corresponding high-level action output based on processing corresponding observation data and corresponding target labels of corresponding navigation targets in the environment, and the corresponding high-level action output indicates which high-level action of multiple discrete high-level actions will result in the closest movement to the corresponding navigation target. Performing supervised training includes: using real images captured around the real environment as part of the input to be processed by the high-level policy model during supervised training; and during supervised training, using the benchmark truth navigation path in the real environment to generate a loss as a supervisory signal. The method further includes performing reinforcement training of the low-level policy model to train the low-level policy model to generate a corresponding low-level action output based on processing corresponding additional observation data and corresponding high-level actions, and the corresponding low-level action output indicates a specific implementation of a high-level action that is more finely defined than the high-level action. Performing reinforcement training includes: using simulated data generated by a robot simulator, generating rewards based on a reward function; and using rewards to update the low-level policy model. The reward function penalizes the robot for collisions, while optionally rewarding faster speed and / or shorter distance to the navigation goal.

Claims

1. A method implemented by one or more processors, the method comprising: Identify the target labels of the robot's robotic task; obtaining current observation data based on current outputs from a sensor assembly of the robot; Processing the current observation data and the target label using a trained high-level policy model to generate a high-level action output; using the high-level action output to select a particular high-level action from a plurality of discrete high-level actions in a defined high-level action space; processing the specific high-level action using the trained low-level policy model to generate a low-level action output, Wherein, the low-level action output defines a low-level action of a defined low-level action space, wherein, when generating the low-level action output, the trained low-level policy model is not used to process the current observation data, and wherein the defined low-level action space defines the robot motion more finely than the high-level action space; and One or more actuators of the robot are controlled based on the low-level motion output so that the robot implements the low-level motion.

2. The method according to claim 1, wherein: The discrete high-level actions of the defined high-level action space are free from any definition of one or more parameters of the robot motion defined in the low-level action space.

3. The method according to claim 1, wherein: The discrete high-level actions of the defined high-level action space do not define any speed of robot motion, and the low-level action space defines one or more speeds of robot motion.

4. The method according to claim 1, wherein: The low-level action space is a continuous action space.

5. The method according to claim 1, wherein: The specific high-level action is one or more words that describe the specific high-level action.

6. The method according to claim 1, wherein: Each corresponding low-level motion defines one or both of the following: one or more corresponding linear velocities and one or more corresponding angular velocities.

7. The method according to claim 1, wherein: The target label is based on a semantic descriptor of the robotic task.

8. The method according to claim 7, wherein: The target label is the word embedding of the semantic descriptor.

9. The method according to claim 7, wherein: Identifying the target tag includes identifying the semantic descriptor based on the semantic descriptor being included in a user interface input.

10. The method according to claim 9, wherein: The user interface input is a spoken user interface input.

11. The method according to claim 1, wherein: The target labels are based on images of the robotic task.

12. The method according to claim 1, wherein: The sensor component is a camera.

13. The method according to claim 1, further comprising: obtaining current additional observation data, the current additional observation data being based on current additional outputs from additional sensor components of the robot; Wherein, when generating the low-level action output, the method includes processing the current additional observation data and the specific high-level action, and using a trained low-level policy model to generate the low-level action output.

14. The method according to claim 13, wherein: The additional sensor component is a proximity sensor, a one-dimensional (1D) LIDAR component or a two-dimensional (2D) LIDAR component.

15. The method according to claim 1, wherein: The trained low-level policy model is trained using reinforcement learning.

16. The method according to claim 1, wherein: The target label is a one-hot vector.

17. The method according to claim 1, further comprising: Determining that the specific advanced action is a specific advanced action that can cause a collision; In response to determining that the specific high-level action is a specific high-level action that can cause a collision, processing the specific high-level action using the trained low-level policy model is performed to generate the low-level action output.

18. A system comprising: Memory, which stores instructions; One or more processors operable to execute the instructions to: Identify the target labels of the robot's robotic task; obtaining current observation data based on current outputs from a sensor assembly of the robot; Processing the current observation data and the target label using a trained high-level policy model to generate a high-level action output; using the high-level action output to select a particular high-level action from a plurality of discrete high-level actions in a defined high-level action space; processing the specific high-level action using the trained low-level policy model to generate a low-level action output, Wherein, the low-level action output defines a low-level action of a defined low-level action space, wherein, when generating the low-level action output, the trained low-level policy model is not used to process the current observation data, and wherein the defined low-level action space defines the robot motion more finely than the high-level action space; and One or more actuators of the robot are controlled based on the low-level motion output so that the robot implements the low-level motion.

19. The system of claim 18, wherein: The one or more processors belong to the robot.

20. A method for training a high-level policy model and a low-level policy model for collaborative use in autonomous navigation of a mobile robot, the method comprising: Performing supervised training of the high-level policy model to train the high-level policy model to generate a corresponding high-level action output based on processing corresponding observation data and a corresponding target label of a corresponding navigation target in an environment, wherein the corresponding high-level action output indicates which of a plurality of discrete high-level actions will result in a motion closest to the corresponding navigation target, wherein performing the supervised training comprises: using real images captured around the real environment as part of the input to be processed by the high-level policy model during the supervised training, and During the supervised training, a loss is generated using a ground truth navigation path in a real environment as a supervisory signal; performing enhanced training of the low-level policy model to train the low-level policy model to generate corresponding low-level action outputs based on processing corresponding additional observation data and corresponding high-level actions, wherein the corresponding low-level action outputs indicate a specific implementation of a high-level action that is more finely defined than the high-level action, wherein performing the enhanced training comprises: generating a reward based on a reward function using simulation data generated by a robot simulator; and The low-level policy model is updated using the reward; wherein the reward function penalizes robot collisions while rewarding faster speed and / or shorter distance in reaching a navigation goal.

21. A method implemented by one or more processors, the method comprising: Identify the target labels of the robot's robotic task; obtaining current observation data based on current outputs from a sensor assembly of the robot; Processing the current observation data and the target label using a trained high-level policy model to determine a specific high-level action in a defined high-level action space; Processing the specific high-level action using the trained low-level policy model to generate a low-level action output, wherein the low-level action outputs define a low-level action of a defined low-level action space, and wherein, when generating the low-level action output, the trained low-level policy model is not used to process the current observation data; and One or more actuators of the robot are controlled based on the low-level motion output so that the robot implements the low-level motion.

22. The method according to claim 21, wherein: The defined high-level action space is free of any definition of one or more parameters of the robot motion defined in the low-level action space.

23. The method according to claim 21, wherein: The defined high-level action space does not define any speeds of robot motion, and the low-level action space defines one or more speeds of robot motion.

24. The method according to claim 21, wherein: The low-level action space is a continuous action space.

25. The method according to claim 21, wherein: The specific high-level action is one or more words that describe the specific high-level action.

26. The method according to claim 21, wherein: Each corresponding low-level motion defines one or both of the following: one or more corresponding linear velocities and one or more corresponding angular velocities.

27. The method of claim 21, wherein: The target label is based on a semantic descriptor of the robotic task.

28. The method according to claim 27, wherein: The target label is the word embedding of the semantic descriptor.

29. The method according to claim 27, wherein: Identifying the target tag includes identifying the semantic descriptor based on the semantic descriptor being included in a user interface input.

30. The method of claim 29, wherein: The user interface input is a spoken user interface input.

31. The method of claim 21, wherein: The target labels are based on images of the robotic task.

32. The method of claim 21, wherein: The sensor component is a camera.

33. The method of claim 31 , further comprising: obtaining current additional observation data, the current additional observation data being based on current additional outputs from additional sensor components of the robot; Wherein, when generating the low-level action output, the method includes processing the current additional observation data and the specific high-level action, and using a trained low-level policy model to generate the low-level action output.

34. The method of claim 33, wherein: The additional sensor component is a proximity sensor, a one-dimensional (1D) LIDAR component or a two-dimensional (2D) LIDAR component.

35. The method of claim 21, wherein: The trained low-level policy model is trained using reinforcement learning.

36. The method of claim 21, wherein: The target label is a one-hot vector.

37. The method of claim 21, further comprising: Determining that the specific advanced action is a specific advanced action that can cause a collision; In response to determining that the specific high-level action is a specific high-level action that can cause a collision, processing the specific high-level action using the trained low-level policy model is performed to generate the low-level action output.