Autonomous task performance based on visual embeddings

By employing pixel matching and pose increment calculation methods, the problem of recognition and task execution accuracy under changing conditions in traditional vision systems is solved, achieving higher robustness and autonomy.

CN114097004BActive Publication Date: 2026-05-01TOYOTA RESEARCH INSTITUTE INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TOYOTA RESEARCH INSTITUTE INC
Filing Date
2020-06-10
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional vision systems are limited in their accuracy when recognizing and performing tasks due to insufficient robustness of unique features in images, especially when they are struggling to effectively recognize and respond to changing conditions.

Method used

A pixel-based matching method is adopted, which identifies the pixel descriptors of the image and matches them with the keyframe image. Combined with pose increment calculation, the accuracy and robustness of the vision system are improved.

Benefits of technology

It improves the vision system's ability to recognize and perform tasks under changing conditions, enhances the robot's autonomous navigation and manipulation capabilities in the environment, and reduces dependence on initial posture and position.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114097004B_ABST
    Figure CN114097004B_ABST
Patent Text Reader

Abstract

A method for controlling a robotic device is presented. The method includes capturing an image corresponding to a current view of the robotic device. The method also includes identifying a keyframe image that contains a first set of pixels that match a second set of pixels of the image. The method also includes performing, by the robotic device, a task corresponding to the keyframe image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit of U.S. Patent Application No. 16 / 570,618, filed September 13, 2019, entitled “AUTONOMOUS TASK PERFORMANCE BASEDON VISUAL EMBEDDINGS”, which claims the benefit of U.S. Provisional Patent Application No. 62 / 877,792, filed July 23, 2019, entitled “KEYFRAME MATCHER”, U.S. Provisional Patent Application No. 62 / 877,791, filed July 23, 2019, entitled “VISUAL TEACH AND REPEAT FOR MANIPULATION-TEACHING VR”, and U.S. Provisional Patent Application No. 62 / 877,793, filed July 23, 2019, entitled “VISUALIZATION”, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] Certain aspects of this disclosure generally relate to robotic devices, and more specifically, to systems and methods for interacting with an environment by using stored memories of that environment. Background Technology

[0004] Autonomous agents (e.g., vehicles, robots, drones, etc.) and semi-autonomous agents use machine vision to analyze regions of interest in their surrounding environment. In operation, an autonomous agent can rely on a trained neural network to identify objects within regions of interest in images of its surrounding environment. For example, a neural network can be trained to recognize and track objects captured by one or more sensors, such as light detection and ranging (LIDAR) sensors, sonar sensors, red-green-blue (RGB) cameras, RGB-depth (RGB)-D cameras, etc. Sensors can be coupled to or communicate with devices (such as autonomous agents). Object detection applications for autonomous agents can analyze sensor image data to detect objects (e.g., pedestrians, cyclists, other vehicles, etc.) in the surrounding scene from the autonomous agent.

[0005] In traditional systems, autonomous agents are trained to perform tasks using a feature-based approach. That is, they perform tasks based on features extracted from an image of the robot's current view. The features of the image are limited to unique elements of objects within the image. If the image lacks many features, the accuracy of traditional object detection systems can be limited. Improving the accuracy of vision systems is desired. Summary of the Invention

[0006] In one aspect of this disclosure, a method for controlling a robotic device is disclosed. The method includes capturing an image corresponding to a current view of the robotic device. The method also includes identifying a keyframe image containing a first set of pixels that matches a second set of pixels in the image. The method further includes having the robotic device perform a task corresponding to the keyframe image.

[0007] In another aspect of this disclosure, a non-transitory computer-readable medium having non-transitory program code recorded thereon is disclosed. The program code is used to control a robotic device. The program code is executed by a processor and includes program code for capturing an image corresponding to the current view of the robotic device. The program code also includes program code for identifying a keyframe image containing a first set of pixels that matches a second set of pixels in the image. The program code further includes program code for the robotic device to perform a task corresponding to the keyframe image.

[0008] Another aspect of this disclosure relates to a robotic device. The device has a memory and one or more processors coupled to the memory. The one or more processors are configured to capture an image corresponding to a current view of the robotic device. The one or more processors are also configured to identify a keyframe image containing a first set of pixels that matches a second set of pixels in the image. The one or more processors are further configured to perform a task corresponding to the keyframe image.

[0009] This has provided a fairly broad overview of the features and technical advantages of this disclosure in order to better understand the detailed embodiments described below. Additional features and advantages of this disclosure will now be described. Those skilled in the art will understand that this disclosure can be readily used as the basis for modifying or designing other structures for achieving the same purposes as this disclosure. Those skilled in the art will also recognize that such equivalent constructions do not depart from the teachings of this disclosure as set forth in the appended claims. The novel features, as well as further objects and advantages considered characteristic of this disclosure in both its organization and manner of operation, will be better understood from the following description when considered in conjunction with the accompanying drawings. However, it should be clearly understood that each drawing is provided for illustrative and descriptive purposes only and is not intended to be a limitation of this disclosure. Attached Figure Description

[0010] The features, nature, and advantages of this disclosure will become more apparent from the detailed description of the embodiments set forth below when used in conjunction with the accompanying drawings, in which the same reference numerals are indicated accordingly throughout the text.

[0011] Figure 1 Examples of operator control of robotic equipment according to various aspects of this disclosure are illustrated.

[0012] Figure 2A Examples of robotic devices operating in an environment according to various aspects of this disclosure are illustrated.

[0013] Figure 2B Examples of images captured by robotic devices according to various aspects of this disclosure are illustrated.

[0014] Figure 2C Examples of keyframes according to various aspects of this disclosure are illustrated.

[0015] Figure 2D Examples of keyframe matching according to various aspects of this disclosure are illustrated.

[0016] Figure 3 Examples of visual outputs generated based on information provided from the operation of a robotic device in an environment, according to various aspects of this disclosure, are illustrated.

[0017] Figure 4 The diagram illustrates an example of the hardware implementation of a robot system according to various aspects of this disclosure.

[0018] Figure 5 A graphical sequence illustrating the behaviors taught according to various aspects of this disclosure is provided.

[0019] Figure 6 Software modules for a robot system according to various aspects of this disclosure are described.

[0020] Figure 7 Methods for controlling robotic devices according to various aspects of this disclosure are described. Detailed Implementation

[0021] The specific embodiments described below with reference to the accompanying drawings are intended as descriptions of various configurations and are not intended to represent the only configuration in which the concepts described herein can be practiced. The specific embodiments include particular details to provide a thorough understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts can be practiced without these particular details. In some cases, well-known structures and components are shown in block diagram form to avoid obscuring these concepts.

[0022] Based on the teachings described herein, those skilled in the art will understand that the scope of this disclosure is intended to cover any aspect of this disclosure, whether implemented independently of or in combination with any other aspect of this disclosure. For example, an apparatus may be implemented using any number of the described aspects, or a method may be practiced. Furthermore, the scope of this disclosure is intended to cover such apparatuses or methods practiced using structures, functions, or structures and functions other than those described in this disclosure. It should be understood that any aspect of this disclosure may be embodied by one or more elements of the claims.

[0023] The word “exemplary” is used in this document to mean “serving as an example, instance, or illustration.” Any aspect described as “exemplary” in this document is not necessarily to be construed as superior to or better than the others.

[0024] While specific aspects are described herein, numerous variations and arrangements of these aspects fall within the scope of this disclosure. Although some benefits and advantages of preferred aspects have been mentioned, the scope of this disclosure is not intended to be limited to specific benefits, uses, or objectives. Rather, various aspects of this disclosure are intended to be broadly applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated by way of example in the accompanying drawings and the following description of preferred aspects. The detailed description and accompanying drawings are merely illustrative and not limiting of this disclosure, the scope of which is defined by the appended claims and their equivalents.

[0025] Autonomous and semi-autonomous agents can perform tasks in response to detected objects. For example, an agent can navigate an environment, identify objects in the environment, and / or interact with objects in the environment. In this application, a robot or robotic device refers to an autonomous or semi-autonomous agent. For simplicity, a robotic device is referred to as a robot.

[0026] The various aspects of this disclosure are not limited to a particular type of robot. Various types of robotic devices are envisioned, having one or more vision systems that provide visual / visual output. The visual output can be provided to the operator via video feed.

[0027] According to various aspects of this disclosure, the robot is capable of manipulating movement. That is, the robot is capable of changing the position of the end effector based on joint configuration. In one configuration, the robot is equipped with automated whole-body control and planning. This allows a human operator to seamlessly demonstrate end effector movement in the task space within virtual reality (VR), with little or no concern for kinematic constraints or the robot's posture.

[0028] The robot includes one or more field-of-view red-green-blue and depth (RGB-D) sensors on a translation / tilt gimbal, providing crucial context for human operators performing tasks in virtual reality. An RGB-D image is a combination of an RGB image and its corresponding depth image. The depth image is an image channel where each pixel relates to the distance between the corresponding object in the image plane and the RGB image.

[0029] Robots can navigate an environment based on images obtained from video feeds (e.g., image feeds). Robots can also perform tasks based on images from video feeds. In traditional vision systems, robots use feature-based methods to recognize images (e.g., images of objects).

[0030] In traditional systems, robots are trained to perform tasks using feature-based methods. That is, they perform tasks based on features extracted from an image in the robot's current view. Feature-based methods identify a test image (e.g., the current image) by finding similarities between the test image and a reference image.

[0031] Specifically, traditional vision systems compare the features of a test image with the features of a reference image. The test image is recognized when the features match those of the reference image. However, in some cases, traditional robots cannot perform the same task when the object's origin and / or orientation / position is inconsistent with the programmed or taught task.

[0032] Features refer to distinctive (e.g., unique) characteristics of an image (e.g., corners, edges, high-contrast areas, low-contrast areas, etc.). An image can include multiple features. Descriptors encode the features of an image. Feature vectors can be examples of descriptors. In traditional vision systems, descriptors are limited to encoding unique characteristics of an image (e.g., objects in an image).

[0033] In most cases, descriptors are robust to image transformations (e.g., localization, scaling, brightness, etc.). That is, traditional vision systems can recognize objects under changing conditions (such as changing lighting conditions). Nevertheless, robustness is limited because descriptors are limited to encoding only unique characteristics of the image. To improve a robot's ability to respond to image recognition tasks, it is desirable to improve the robustness of the target vision system.

[0034] Various aspects of this disclosure relate to assigning descriptors to pixels of an image. That is, unlike conventional vision systems, various aspects of this disclosure do not limit descriptors to unique features. Therefore, the accuracy of the vision system can be improved. Specifically, conventional vision systems are limited to comparing different features (e.g., characteristics) of a test image with different features of a reference image. Images include more pixels than features. Therefore, by comparing pixels rather than features, the accuracy of the comparison is improved.

[0035] In one configuration, the vision system compares the pixels of a test image with the pixels of a reference image. The reference image can be referred to as a keyframe. Keyframes can be obtained during training and are invariant to pose and image transformations. When the number of matching pixels exceeds a threshold, the vision system determines that the test image matches a keyframe.

[0036] In one configuration, the vision system compares pixel descriptors to identify matching pixels in a test image and keyframes. Pixel descriptors include pixel-level information and depth information. Pixel-level information includes things like the pixel's RGB value and the context of the pixel / surrounding pixels within the image.

[0037] The context of a pixel includes information used to disambiguate the pixel. The size of the context depends on the pixel and is learned during training. For example, a pixel in the middle of a plain white wall will require more context than a pixel on a highly textured logo on a bottle. The robot learns how much context is needed without having to manually create this functionality. Depth information indicates the distance from the surface corresponding to the pixel to the sensor used to obtain an image of the surface.

[0038] During training, one or more sensors acquire image data of the scene as the robot performs a task. Keyframes are generated from the acquired image data. When the robot is performing a specific task, a keyframe can be conceptualized as a memory of the environment. That is, a keyframe can be set as an anchor point for a position within the task or environment. After testing, when the robot matches the current image with a keyframe, a training function corresponding to that keyframe can be executed.

[0039] Descriptors can be assigned to pixels in keyframes. A descriptor can be a vector or array of values. For example, a descriptor can be a hexadecimal vector. Descriptors (e.g., pixel descriptors) can be used to identify matching images.

[0040] After training, the robot can be initialized to perform one or more learning tasks. While performing commanded tasks, the robot can collect real-time image data. As the robot traverses and / or manipulates objects within its environment, the real-time image data is mapped pixel-by-pixel to pixels within keyframes. The correspondence between the current image and the keyframe is determined based on the similarity of pixel descriptors.

[0041] That is, the robot compares the values ​​of pixel descriptors in the current image with those in the keyframe. The confidence level in the comparison indicates the probability that a pixel in the current image matches a pixel in the keyframe. The confidence level in a match can be determined using the rigid body assumption. That is, the confidence level can be determined based on the pairwise distances between points on the object in the current image and the keyframe. The confidence level increases when the distance between points on the object remains constant, regardless of the poses of the two images. The pose of the image refers to the robot's pose when the image was captured. The confidence level decreases when the distance between points on the object varies between images with different poses.

[0042] The robot can also visualize the confidence level of a match using diagonal / parallel correspondences (e.g., non-cross-correspondences) between multiple pixels in the current image data and keyframes. This visualization can be transmitted to a remote operations center. Cross-correspondences indicate pixel mismatches between the current image data and keyframes.

[0043] In one configuration, the robot is placed in an environment and image data of the scene is collected. The robot is then controlled (e.g., via a virtual reality interface) to perform tasks. In virtual reality, the human operator's field of vision is limited to the robot's field of vision during training, improving the robot's ability to perform tasks autonomously.

[0044] Robots can be taught to perform actions / tasks by parameterizing operator-performed behaviors through a virtual reality interface. For example, a virtual reality interface could include the use of paddles, handheld controllers, paintbrush tools, wiping tools, and / or placement tools manipulated by an operator wearing a headset depicting a VR environment. Thus, instead of teaching direct task-space motion, the human operator teaches a set of parameterized primitives (or behaviors). These parameterized primitives combine collision-free motion planning with hybrid (position and force) Cartesian control to reduce the number of parameters taught during execution and provide robustness.

[0045] Parametric behavior refers to learning a task by breaking it down into a small number of discrete behaviors. Each behavior is defined by a set of parameters, such as joint angle changes and rotation angles. The values ​​of these parameters can be configured and updated based on the robot's behavior while performing the task.

[0046] Parameterized behaviors can be learned and extracted from a learning task and combined with other tasks to form larger tasks. Parameterized behaviors (such as opening a door with a rotary handle) can be implemented to perform the opening of any door handle (e.g., a door handle requiring thirty (30) degrees of rotation or a door handle requiring sixty (60) degrees or more of rotation). For example, the number of degrees of rotation can be a parameter defining the parameterized behavior of opening a door with a rotary handle.

[0047] To perform a task, the robot can be placed (relative to its initial position during training) in the same or similar environment. The robot can be placed in different initial positions and optionally in different initial poses (e.g., adjusting joint angles to different initial positions). The robot can be assigned to perform the same task (e.g., a set of parameterized behaviors) (such as picking up a bottle, opening a cabinet, and placing the bottle inside the cabinet) (without human operator control). For example, the robot can perform the same task by updating the behavioral parameters taught during virtual reality control sequences. Parameters can be updated based on the robot's current pose and / or localization, relative to the pose and / or localization used during training.

[0048] As discussed, to update parameters, the robot captures an initial image of the scene and maps pixels and / or descriptors from the new image to keyframes. This mapping defines the relative transformation between the new image and the keyframes. The relative transformation provides the mapping from keyframes to the new image. The relative transformation can be defined by changes in the robot's x-axis position, y-axis position, z-axis position, roll, pitch, and yaw. The relative transformation can be used to update the parameters of parameterized behavior from the learned parameters to the observed conditions.

[0049] Relative transformations can be applied to parameterized behaviors. By applying relative transformations to parameterized behaviors, a robot can perform the same task as previously taught, even if the initial position and / or pose has changed. The robotic system can continuously map pixels and / or dense neural network descriptors from the current scene to those from keyframes, allowing parameterized behaviors to be adjusted at intervals (e.g., continuously). For example, relative transformations can be applied to taught actions (such as pulling open a drawer, opening a door, picking up a cup or bottle, etc.) defined by a set of parameterized behaviors.

[0050] In some respects, actions can be scene-wide and / or object-specific. For example, the action of picking up a bottle could include navigating to the bottle using keyframes that are scene-wide. Once near the bottle, bottle-specific keyframes could be analyzed independently of the environment.

[0051] Navigation behaviors can be used to move a robot from one point to another. This allows the robot to locate objects (such as bottles) that could be anywhere in the environment, and then, during training for the "pick up" action, perform tasks such as "pick up" the bottle regardless of its location. Manipulation behaviors are used to move individual parts of the robot (e.g., the torso and / or arms) to make contact with a desired object.

[0052] Figure 1 An example illustrating how an operator 100 controls a robot 106 according to various aspects of this disclosure is provided. Figure 1 As shown, operator 100 is equipped with a vision system 102 and a motion controller 104 (e.g., a gesture tracking system) for controlling robot 106. In this example, operator 100 can control robot 106 to perform one or more tasks. Figure 1 In the example, robot 106 is trained in kitchen 108. Various aspects of this disclosure are not limited to training robot 106 in kitchen 108; other environments are also envisioned.

[0053] The vision system 102 can provide video feeds and capture the gaze of the operator 100. The operator 100 can be located at a location remote from the location of the robot 106. In this example, the robot 106 is located in the kitchen 108 and the operator 100 is located at a different location than the kitchen 108 (such as the robot control center 114).

[0054] Vision system 102 can provide video feeds for the positioning of robot 106. For example, vision system 102 can provide a view of kitchen 108 based on the forward-looking perspective of robot 106. Other perspectives (such as a 360-degree view) can be provided. Perspectives are provided via one or more vision sensors (such as cameras) of robot 106. Vision system 102 is not limited to... Figure 1 The head-mounted device shown. The vision system 102 can also be a monitor 110, an image projector, or another system capable of displaying video feeds from the robot 106.

[0055] One or more actions of robot 106 can be controlled via motion controller 104. For example, motion controller 104 captures gestures from operator 100 and robot 106 mimics the captured gestures. Operator 100 can control the movement, limb movements, and other actions of robot 106 via motion controller 104. Various aspects of this disclosure are not limited to capturing operator 100's gestures via motion controller 104. Other types of gesture capture systems are conceivable. Operator 100 can control robot 106 via wireless connection 112. Additionally, robot 106 can provide feedback (such as video feeds) to operator 100 via wireless connection 112.

[0056] Robot 106 can be trained to navigate in a specific environment (e.g., kitchen 108) and / or similar environments. Furthermore, robot 106 can be trained to perform tasks on objects in the specific environment and / or similar objects in any environment. For example, robot 106 can be trained to open / close drawers in kitchen 108. Training can be performed only on drawers in kitchen 108 and / or similar drawers in any environment (such as drawers in another kitchen).

[0057] Figure 2A Examples of robot 200 in environment 202 according to various aspects of this disclosure are illustrated. For clarity, Figure 2A A top-down approach is provided for environment 202. For example... Figure 2A As shown, environment 202 includes a dining table 204, a sink 206, a drawer 208 with a spoon 218, and a counter 210.

[0058] exist Figure 2AIn the example, the open drawer 208 is within the robot 200's field of view 224 (e.g., gaze). The open drawer 208 will be displayed in the current image data captured by the robot 200's sensors. In one configuration, the robot 200 compares the current image data with keyframe data to determine whether an action should be performed.

[0059] Figure 2B This illustrates an example of the current image 252 captured by robot 200 during the test. (See example image 252.) Figure 2B As shown, based on the field of view 224 of robot 200, the current image 252 includes a sink 206 and a counter 210. In one configuration, robot 200 compares the pixel descriptors of the current image 252 with the pixel descriptors of keyframes to determine subsequent actions and / or to determine the robot's localization in the environment.

[0060] Figure 2C This illustrates an example of keyframe 260 captured by robot 200 during training. (See example...) Figure 2C As shown, keyframe 260 includes a sink 206, a counter 210, a drawer 208, and a spoon 218. Keyframe 260 can be captured in an environment (e.g., positioning) that is different from or the same as the environment in which the robot 200 captures the current image 252. The angle and / or depth of keyframe 260 can be different from or the same as the angle and / or depth of the current image 252. For simplicity, in Figure 2B , Figure 2C and Figure 2D In the example, keyframe 260 and the current image 252 have the same angle and depth.

[0061] In one configuration, the user annotates the area of ​​keyframe 260 as the anchor region. For example, in Figure 2C In the image, spoon 218 and drawer 208 are movable, dynamic objects. In contrast, the top of the faucet or the edge of the sink are static objects. Therefore, users can annotate static objects, allowing the pixels of these static objects to be used to determine if the current image matches a keyframe. Annotating the pixels of static objects (e.g., anchor points) can improve the accuracy of pose increments.

[0062] Figure 2D Examples of comparing the current image 252 with keyframe 260 according to various aspects of this disclosure are illustrated. In one configuration, robot 200 compares the pixel descriptor (e.g., the value of the pixel descriptor) of the current image 252 with the pixel descriptor of keyframe 260. Line 270 between the current image 252 and keyframe 260 illustrates the matching pixels between the two images.

[0063] When two or more lines 270 do not intersect, the image can be matched. When two or more lines 270 intersect, a mismatch can be identified. The current image 252, keyframe 260, and lines 270 can be output to the monitor in the operation center. Figure 2D The number of pixels compared is for illustrative purposes. Various aspects of this disclosure are not limited to simply comparing... Figure 2D The number of pixels being compared.

[0064] As discussed, after training, the robot can be initialized to perform one or more learning tasks. For example, the robot's task could be to navigate and / or manipulate objects in an environment. The robot collects real-time image data while performing commanded tasks. This real-time image data is then compared to keyframes.

[0065] In one configuration, the confidence level of matching pixels between the current image data and keyframes is visualized. The display can also visualize comparisons of pixel descriptors with various thresholds. For example, if the image depth is greater than a threshold, the image may be rejected.

[0066] When one or more keyframes match the current image data, actions associated with the matched keyframes can be executed. For example, in Figure 2D In this context, upon determining that the current image 252 matches keyframe 260, the robot can perform a task (such as opening or closing a drawer). As another example, the robot can turn a faucet on / off. Alternatively, actions can be queued for execution once a previous action is completed. As yet another example, the robot can determine its current location based on a keyframe and use that current location to navigate along a path to another location. New keyframes can be referenced when the robot moves within the environment or changes the position of objects in the environment.

[0067] In one configuration, one or more pixels are filtered before comparing pixel descriptors. Functions can be executed to filter outliers—for example, the Random Sample Consensus (RANSAC) function can be used to filter out pixel outliers by randomly sampling observation data. Filtering pixels reduces noise and improves accuracy.

[0068] As discussed, the depth and / or angle of the current image may differ from the depth and / or angle of the matching keyframe. Therefore, in one configuration, the relative transformation between the matching pixels of the current image data and the pixels of the keyframe is determined. That is, the pose increment can be determined based on the comparison between the current image and the keyframe. The increment refers to the change in the robot's pose from the keyframe to the current image.

[0069] Pose increments can be determined based on depth information in pixel descriptors. Depth information refers to the distance from the sensor to the object corresponding to the pixel. Pose increments refer to the (x,y,z) coordinate, roll, pitch, and yaw transformations between the current image and the keyframe. In other words, pose increments explain how the image is transformed based on the sensor movement from the keyframe to the current image.

[0070] The accuracy of the pose increment increases with the number of matching pixels. In this disclosure, the number of matching pixels can be greater than the number of matching features. Therefore, various aspects of this disclosure improve the accuracy of determining the pose increment compared to conventional systems that rely solely on feature comparison of images. The pose increment can be determined using a least-squares function.

[0071] In other words, the attitude can be determined by minimizing the squared reprojection error. Specifically, the pixels in the keyframe should be the closely matching pixels found in the current image. Minimizing the squared reprojection error can be expressed as a nonlinear least squares problem, which can be solved for six unknown parameters (x, y, z, roll, pitch, yaw). The RANSAC loop can be used to reduce sensitivity to outliers.

[0072] Pose increments (e.g., relative transformations) between one or more values ​​of a pixel's pixel descriptor can be used to update parameters of the behavior to be performed. For example, if the pose increment indicates that the robot is positioned one foot away from an object compared to the position defined in a keyframe, the behavior can be updated to describe the robot's current position (e.g., pose). The position defined in the keyframe refers to the location used to train the robot to perform the task.

[0073] Figure 3 Examples of output 300 from a vision system according to various aspects of this disclosure are illustrated. For example... Figure 3 As shown, output 300 includes the current image 302, keyframe 308, histogram 306, and threshold window 304. Output 300 can be displayed at a location remote from the robot's current position. For example, output 300 can be displayed at a control center that controls the robot's operation.

[0074] Current image 302 is an image from the robot's current field of view. Keyframe 308 is a set of keyframes compared to current image 302. Histogram 306 illustrates the confidence score used to match current image 302 with each keyframe in this set of keyframes.

[0075] During training, for navigation tasks, the robot captures keyframes at certain points along the route (from start to finish). These keyframes captured along the route serve as a set of keyframes for the navigation task. During testing, the robot compares the current image with this set of keyframes to identify its current location.

[0076] Based on the number of matching pixels (e.g., pixel descriptors), the robot determines the match confidence of the current image with a specific keyframe. When the match confidence is greater than a threshold, the robot performs a task associated with the matching keyframe. The task may include localization (e.g., determining the robot's current position). Figure 3 In the example, the robot determines its location on the route corresponding to the location of the keyframe with the maximum match confidence of 310.

[0077] Histogram 306 can vary depending on the task. Figure 3 Histogram 306 in the table is used for navigation tasks. If the robot is performing a manipulation task (e.g., object manipulation), histogram 306 will be different. A set of keyframes for a manipulation task can have fewer keyframes compared to a set used for navigation tasks.

[0078] For example, to open a cabinet, a robot can be trained to open it if the cabinet is closed or partially open. In this example, the histogram would include two keyframe matching bars in the bar chart. One bar would indicate the confidence level of the current image matching a keyframe of a closed cabinet, while the other would indicate the confidence level of the current image matching a keyframe of a partially open cabinet.

[0079] The pose confidence window 304 illustrates the confidence level in the pose matching from the pose matcher. In one configuration, before determining the pose increment, the robot determines whether multiple criteria meet one or more thresholds. These criteria can be based on depth, pose, clique, error, and other elements.

[0080] In one configuration, the robot determines the number of pixels with depth values. For some surfaces (such as glass or glossy surfaces), the robot may be unable to determine depth. Therefore, pixels corresponding to these surfaces may not have depth values. The number of pixels with depth values ​​can be determined before filtering pixels. If the number of pixels with depth values ​​is greater than a threshold, a depth criterion is met. A pose confidence window 304 may include a bar graph that compares bars 320 representing the number of pixels with depth values ​​with bars 322 representing the threshold. The bars may be color-coded.

[0081] Alternatively, the bar chart of the pose confidence window 304 may include a comparison of bars representing the number of features aligned between the current image and keyframes with bars representing an alignment threshold. To determine the accuracy of the match between the keyframe and the current image, the robot can apply a pose transformation and determine how many features of an object in the current image are aligned with features of an object in the keyframe, and vice versa. If a feature is associated with a non-static object, the feature may not be aligned. For example, a keyframe may include a cup that is no longer present in the current image. Therefore, the features of the cup will not be aligned between images. The number of aligned features can be called the clique inlier. When the clique inlier is greater than a threshold, pose increments can be generated.

[0082] Alternatively, the bar chart of the attitude confidence window 304 may include a comparison of bars representing the number of features used to determine the attitude with bars representing the attitude threshold. That is, a certain number of features should be used to determine the attitude. If the number of features used to determine the attitude is less than the attitude threshold, the attitude increment may not be calculated. The number of features used to determine the attitude can be referred to as the attitude normality.

[0083] The bar chart of the pose confidence window 304 may also include bars comparing the root mean square (RMS) error of the pose increment to a threshold. If the RMS error is less than the RMS error threshold, the pose increment is likely satisfactory. If the RMS error is greater than the RMS error threshold, the user or robot may not use the pose increment to perform the task.

[0084] The attitude confidence window 304 is not limited to a bar chart; other graphs or images can be used. The criteria for the attitude confidence window 304 are not limited to those discussed; other criteria can be used.

[0085] As discussed, aspects of this disclosure relate to mobile manipulation hardware and software systems capable of autonomously performing human-level tasks in real-world environments after being taught tasks using demonstrations from humans in virtual reality. In one configuration, a mobile manipulation robot is used. The robot may include full-body task-space hybrid position / force control. Furthermore, as discussed, parameterized primitives linked to robustly learned dense visual embedding representations of the scene are taught to the robot. Finally, a task graph of the taught behavior can be generated.

[0086] The aspects of this disclosure are not about programming or training a robot to recognize a fixed set of objects or perform predefined tasks, but rather about enabling a robot to learn new objects and tasks from human demonstrations. The robot can autonomously perform the learned tasks under naturally varying conditions. The robot does not use prior object models or maps and can be taught, based on an example, to associate a given set of behaviors with arbitrary scenes and objects. The vision system is trained offline on existing supervised and unsupervised datasets, and the rest of the system can function normally without additional training data.

[0087] Compared to conventional systems that teach motion in direct task space, various aspects of this disclosure teach a set of parameterized behaviors. These behaviors combine collision-free motion planning and hybrid (position and force) Cartesian end-effector control, minimizing the taught parameters and providing robustness during execution.

[0088] In one configuration, task-specific, learning-intensive visual pixel-by-pixel embeddings are computed. These pixel-by-pixel embeddings link parameterized behavior to the scene. Due to this linking, the system can handle diverse environments with high robustness by sacrificing generalization to new situations.

[0089] Task behaviors can be taught independently using visual input conditions and based on successful exit criteria. These behaviors can be linked together in a dynamic task graph. Because the behaviors are linked, the robot can reuse behaviors to perform task sequences.

[0090] The robot may include multiple degrees of freedom (DOF). For example, the robot may include 31 DOFs, divided into five subsystems: chassis, torso, left arm, right arm, and head. In one configuration, the chassis includes four drive and steerable wheels (e.g., eight DOFs in total) to enable “pseudo-integrity” mobility. The drive / steering actuator assembly may include various motors and gearboxes. The torso may include five DOFs (yaw-pitch-pitch-pitch-yaw). Each arm may include seven DOFs. The head may be a translation / tilt head with two DOFs. Each arm may also include a single DOF gripper with underactuated fingers. The aspects of this disclosure are not limited to the robot discussed above. Other configurations are contemplated. In one example, the robot includes custom tools (such as sponges or swiffer tools).

[0091] In one configuration, force / torque sensors are integrated with the robot to measure the forces acting on its interaction with the environment. For example, force / torque sensors can be placed at the wrist of each arm. Sensing sensors can be integrated into the head to provide a wide field of view, while also enabling the robot and human to perform tasks in VR scenarios.

[0092] Various aspects of this disclosure provide several levels of abstraction for controlling robots. In one configuration, the lowest control level provides real-time coordinated control of the DOF (Domain of Flight) of all robots. Real-time control may include joint control and component control. Joint control enables low-level device communication and exposes device commands and states in a generic manner. Furthermore, joint control supports actuators, force sensors, and inertial measurement units. Joint control can be configured at runtime to support different robot variations.

[0093] Component control can handle higher-level coordination of a robot by dividing it into multiple components (e.g., right arm, head, etc.) and providing a set of parameterized controllers for each component. Component control can provide controllers for: joint position and velocity; joint admittance; camera appearance; chassis position and velocity; and hybrid task space pose, velocity, and admittance control.

[0094] End-effector taskspace control provides another level of abstraction for controlling the robot. This level of abstraction addresses the robot's posture to achieve these desired movements. The whole-body inverse kinematics (IK) for hybrid Cartesian control is formulated as a quadratic program and solved. Components may be subject to linear constraints of joint position, velocity, acceleration, and gravitational torque.

[0095] Full-body IK can be used for motion planning of Cartesian pose targets. In one configuration, the occupied environment voxels are equipped with spheres and bladders. Voxel collision constraints are added to the quadratic planning IK to prevent collisions between the robot and the world. Motion planning can be performed using a rapidly-exploring random tree (RRT), sampling in Cartesian space with quadratic planning IK as the turning function between nodes.

[0096] Planning in Cartesian space leads to natural and direct motion. Using quadratic programming (IK) as the steering function improves the reliability of the planning because the same controller can be used for both planning and execution, reducing potential discrepancies between the two. Similarly, motion planning for joint position targets combines RRT with a component-controlled joint position controller that acts as the steering function.

[0097] The next level of abstraction defines parameterized behaviors. In a configuration, parameterized behaviors are primitive actions that can be parameterized and ordered together to accomplish a task. Behaviors can include, but are not limited to: manipulation actions (such as grabbing, lifting, releasing, pulling, retracting, rubbing, joint movement, and direct control); navigation actions (such as driving using speed commands, entering using position commands, and following a path using active obstacle avoidance); and other auxiliary actions (such as looking and stopping).

[0098] Each action can have one or more actions of different types (such as joint or Cartesian movements of one or more robot parts). Each action can use different control strategies (such as position, velocity, or admittance control), and motion planning can also be used to avoid external obstacles. The robot's motion, whether or not it uses motion planning, avoids self-collisions and satisfies motion control constraints.

[0099] Each behavior can be parameterized by different actions, which in turn will have their own parameters. For example, the grasping behavior can include four parameters: gripper angle, 6D approach, grasp, and (optionally) gripper lifting posture. In this example, these parameters define the following predefined sequence of actions: (1) opening the gripper to the desired gripper angle; (2) planning and executing a collision-free path for the gripper to the 6D approach posture; (3) moving the gripper to the 6D grasping posture and stopping upon contact; (4) closing the gripper; and (5) moving the gripper to the 6D lifting posture.

[0100] The ultimate level of control abstraction is the task. In one configuration, a task is defined as a sequence of actions that enables the robot to manipulate and navigate in a human environment. (See Task Diagram) Figure 5 It is a directed (cyclic or acyclic) graph with different tasks as nodes and different transition conditions as edges, including fault detection and fault recovery. Edge conditions include the state of each action execution, checking the object in hand using force / torque sensors, voice commands, and keyframe matching for handling different objects and environments.

[0101] According to various aspects of this disclosure, a perception pipeline is designed to provide a robot with an understanding of its surrounding environment. The perception pipeline also provides the robot with the ability to identify actions to be taken given a taught task. In one configuration, a fused RGB-D image is created by projecting multiple depth images onto a wide-field-of-view image (e.g., a wide-field-of-view left image) of a high-resolution color stereo pair. The system runs a set of deep neural networks to provide various pixel-level classifications and feature vectors (e.g., embeddings). Based on visual features recalled from the taught sequence, the pixel-level classifications and feature vectors are accumulated into a temporal 3D voxel representation. The pixel-level classifications and feature vectors can be used to recall actions to be performed.

[0102] In one configuration, object categories are not defined. Furthermore, no model of the object or environment is assumed. Instead of explicitly detecting and segmenting objects and explicitly estimating 6-DOF object poses, it can generate dense pixel-level embeddings for a variety of tasks. Reference embeddings from the teaching sequence can be used to perform each behavior classification or pose estimation.

[0103] The trained model can be fully convolutional. In one configuration, pixels in the input image are mapped to points in the embedding space. The embedding space is assigned a metric, which is implicitly defined by the loss function, and the training process is defined by the model output. The trained model can be used for a variety of tasks.

[0104] In one configuration, given a single annotated example, the trained model detects all objects of a semantic class. Objects of a semantic class can be detected by comparing embeddings on the annotation with embeddings seen in other regions. The model can be trained using a discriminative loss function.

[0105] A model can be trained to determine object instances. This model identifies and / or counts individual objects. The model can be trained to predict a vector (2D embedding) at each pixel. This vector can point to the centroid of the object containing that pixel. At runtime, pixels pointing to the same centroid can be grouped to segment the scene. Runtime execution can be performed in 3D.

[0106] The model can also be trained on 3D correspondences. This model generates view- and lighting-invariant per-pixel embeddings so that any view of a given 3D point in the scene maps to the same embedding. A loss function can be used to train this model.

[0107] The pixel-wise embeddings (and depth data) of each RGB-D frame are fused into a dynamic 3D voxel map. First- and second-order position, color, and embedding statistics are accumulated for each voxel. The expiration of dynamic objects is based on the back projection of the voxel onto the depth image. The voxel map is segmented using standard graph segmentation based on semantic and instance labels and geometric proximity. The voxel map is also collapsed into a 2.5D map with stereoscopic and traversability classification statistics.

[0108] Voxel maps are used for collision-free full-body motion planning, while 2.5D maps are used for collision-free chassis motion. For 3D collision checking, a greedy approach can be used to group voxels in the map into vesicles. The segmented objects can be used by behaviors to attach objects to the hand when grasping.

[0109] A one-shot teaching approach can be used to teach a robot, enabling it to recognize features in a scene (or a specific manipulated object) that are highly correlated with features from previously taught task records. These features are saved as keyframes throughout the task as the user demonstrates it. A keyframe can be an RGB image containing a multidimensional embedding with per-pixel depth (if valid).

[0110] Assuming the current image is sufficiently similar to the reference image present during teaching, the embedding acts as a feature descriptor that allows for the establishment of per-pixel correspondences at runtime. Since depth exists at (most) every pixel, the correspondences can be used to solve for the incremental pose between the current and reference images. Inliers can be detected using Euclidean constraints, and the 6-DOF pose can be solved by applying a Levenberg-Marquardt least squares function with RANSAC.

[0111] Incremental pose is used as a correction, which can be applied to adapt the learned sequence of behaviors to the current scene. Because embeddings can be defined at each pixel, keyframes can be as wide as including every pixel in the image, or as narrow as using only pixels from a user-defined mask. As discussed, users can define masks by selectively annotating regions of the image as task-related or on objects.

[0112] In addition to visual sensing, in one configuration, the robot collects and processes audio input. Audio provides another set of embeddings as input for teaching the robot. For example, the robot acquires audio input by asking questions and understanding spoken language responses from humans. Custom keyword detection modules can be used to understand the speech responses.

[0113] The robot can use a fully-convolutional keyword-spotting model to understand a custom wakeword, a set of objects (e.g., "cup" or "bottle"), and a set of locations (e.g., "cabinet" or "refrigerator"). In one configuration, the model listens for wakewords at regular intervals (e.g., every 32 milliseconds). When a wakeword is detected, the robot searches for either an object to detect or a keyword to locate. During training, noise is artificially added to make the recognition more robust.

[0114] As discussed, to teach the robot tasks, the operator uses a set of VR modes. Each behavior can have a corresponding VR mode for setting and commanding specific parameters of that behavior. Depending on the type of parameter, each behavior mode can include customized visualizations to assist in setting each parameter. For example, when setting parameters for a door-pulling motion, the hinge axis is marked and visualized as a line, and the candidate pulling postures of the gripper are restricted to fall on an arc around the hinge. Several practical VR modes are used to assist the teaching process (such as resuming behavior, annotating the environment with relevant objects, and repositioning the virtual robot, camera images, and menus in the VR world).

[0115] During execution, the robot's pose and environmental components may differ from those used during training. Feature matching can be used to find features in the environment similar to the taught features. Pose increments can be built from the matched feature correspondences. The user-taught behavior is transformed by the computed pose increments. In one configuration, multiple keyframes are passed to the matching problem. The best matching keyframe is selected based on the number of correspondences.

[0116] Figure 4 This is a diagram illustrating examples of the hardware implementation of a robot system 400 according to various aspects of this disclosure. The robot system 400 may be a component of an autonomous or semi-autonomous system (such as a vehicle, robot device 428, or other equipment). Figure 4 In the example, robot system 400 is a component of robot device 428. Robot system 400 can be used to control the actions of robot device 428 by inferring the intentions of the operator of robot device 428.

[0117] The robot system 400 can be implemented using a bus architecture, typically represented by bus 430. Depending on the specific application and overall design constraints of the robot system 400, bus 430 may include any number of interconnect buses and bridges. Bus 430 links together various circuits including one or more processors and / or hardware modules represented by processor 420, communication module 422, positioning module 418, sensor module 402, motion module 426, memory 424, keyframe module 408, and computer-readable medium 414. Bus 430 may also link various other circuits (such as timing sources, peripherals, voltage regulators, and power management circuits well known in the art), which will not be described further.

[0118] Robot system 400 includes a transceiver 416 coupled to a processor 420, a sensor module 402, a keyframe module 408, a communication module 422, a positioning module 418, a movement module 426, a memory 424, and a computer-readable medium 414. The transceiver 416 is coupled to an antenna 434. The transceiver 416 communicates with various other devices via a transmission medium. For example, the transceiver 416 can receive commands via transmission from an operator of robot device 428. As discussed herein, the operator may be located at a position remote from the positioning of robot device 428. As another example, the transceiver 416 can transmit information about keyframe matching from keyframe module 408 to the operator.

[0119] Robot system 400 includes a processor 420 coupled to a computer-readable medium 414. The processor 420 performs processing, including the execution of software stored on the computer-readable medium 414 that provides the functionality according to this disclosure. When executed by the processor 420, the software causes the robot system 400 to perform various functions described for a particular device, such as robot device 428 or any of modules 402, 408, 414, 416, 418, 420, 422, 424, 426. The computer-readable medium 414 can also be used to store data manipulated by the processor 420 during software execution.

[0120] Sensor module 402 can be used to acquire measurements via various sensors, such as first sensor 406 and second sensor 404. First sensor 406 may be a vision sensor (such as a stereo camera or RGB camera) for capturing 2D images. Second sensor 404 may be a ranging sensor (such as a LiDAR sensor, RADAR sensor, or RGB-D sensor). Of course, the aspects of this disclosure are not limited to the sensors described above, as other types of sensors (such as, for example, thermal sensors, sonar sensors, and / or laser sensors) may also be considered for use with any of sensors 404 and 406. Measurements from first sensor 406 and second sensor 404 can be processed by one or more of processor 420, sensor module 402, communication module 422, positioning module 418, movement module 426, and memory 424 in conjunction with computer-readable medium 414 to implement the functions described herein. In one configuration, data captured by first sensor 406 and second sensor 404 can be transmitted to an operator as a video feed via transceiver 416. The first sensor 406 and the second sensor 404 can be coupled to or communicate with the robot device 428.

[0121] The positioning module 418 can be used to determine the location of the robot device 428. For example, the positioning module 418 can use a Global Positioning System (GPS) to determine the location of the robot device 428. The communication module 422 can be used to facilitate communication via transceiver 416. The communication module 422 can be configured to provide communication capabilities via various wireless protocols, such as WiFi, LTE, 3G, etc. The communication module 422 can also be used to communicate with other components of the robot device 428 that are not modules of the robot system 400.

[0122] The mobility module 426 can be used to facilitate the movement of the robot device 428 and / or its components (e.g., limbs, hands, etc.). As an example, the mobility module 426 can control the movement of limbs 438 and / or wheels 432. As another example, the mobility module 426 can communicate with the power source of the robot device 428 (such as a motor or battery).

[0123] The robot system 400 also includes a memory 424 for storing data related to the operation of the robot device 428 and the keyframe module 408. These modules may be software modules running in the processor 420, residing in / stored in the computer-readable medium 414 and / or the memory 424, one or more hardware modules coupled to the processor 420, or some combination thereof.

[0124] The keyframe module 408 can communicate with the sensor module 402, transceiver 416, processor 420, communication module 422, positioning module 418, motion module 426, memory 424, and computer-readable medium 414. In one configuration, the keyframe module 408 receives input (e.g., image data) corresponding to the current view of the robotic device 428 from one or more sensors 404, 406. The keyframe module 408 can assign pixel descriptors to each pixel or a group of pixels in the input.

[0125] Keyframe module 408 identifies a keyframe image having a first set of pixels that matches the second set of pixels in the input. Keyframe module 408 can match the input pixel descriptors with the pixel descriptors of the keyframe image. The keyframe image can be stored in memory 424, in a network storage device (not shown), or in another location. Keyframe images can be captured and stored while training the robot device 428 to perform actions. Pixel descriptors of the keyframe image can be assigned during training.

[0126] The keyframe module 408 also controls the robot device 428 to perform actions corresponding to the keyframe images. The actions can be adjusted based on the relative transformation between the input and keyframe images (e.g., pose increments). The keyframe module can determine the relative transformation based on changes in distance and / or pose in the input and keyframe images. The actions of the robot device 428 can be controlled via communication with the motion module 426.

[0127] Figure 5 Examples of graph sequences 500 for the taught behavior according to various aspects of this disclosure are illustrated. For example... Figure 5 As shown, the graph sequence 500 includes a start node 502 and an end node 504. The graph sequence 500 can branch or loop based on sensed visual input, audio input, or other conditional input.

[0128] For example, such as Figure 5As shown, after starting node 502, the robot can perform the "listen_for_object" behavior. In this example, the robot determines whether it has sensed visual or audio input corresponding to a cup or a bottle. In this example, different sequences of actions are performed based on whether the sensed input corresponds to a cup or a bottle. Various aspects of this disclosure are not limited to... Figure 5 The behavior shown is as described.

[0129] Figure 6 Examples of software modules for a robotic system according to various aspects of this disclosure are illustrated. Figure 6 The software modules can be used Figure 4 One or more components of the hardware system (such as processor 420, communication module 422, positioning module 418, sensor module 402, motion module 426, memory 424, keyframe module 408, and computer-readable medium 414). Various aspects of this disclosure are not limited to... Figure 6 The module.

[0130] like Figure 6 As shown, the robot can receive audio data 604 and / or image data 602. Image data 602 can be an RGB-D image. Audio network 606 can listen for wake words at certain intervals. Audio network 606 receives raw audio data 604 to detect wake words and extract keywords from raw audio data 604.

[0131] A neural network (such as a dense embedding network 608) receives image data 602. Image data 602 can be received at intervals. The dense embedding network 608 processes the image data 602 and outputs an embedding 610 for the image data 602. The embedding 610 and the image data 602 can be combined to generate a voxel map 626. The embedding 610 can also be input to a keyframe matcher 612.

[0132] The keyframe matcher 612 compares the embedding 610 with multiple keyframes. When the embedding 610 corresponds to the embedding of a keyframe, a matching keyframe is identified. The embedding 610 may include pixel descriptors, depth information, and other information.

[0133] Task module 614 can receive one or more task graphs 616. Task module 614 provides a response to a request from keyframe matcher 612. Keyframe matcher 612 matches tasks with matching keyframes. Tasks can be determined from task graphs 616.

[0134] The task module 614 can also transmit behavior requests to the behavior module 618. The behavior module 618 provides behavior status to the task module 614. Additionally, the behavior module 618 can request information about matching keyframes and corresponding tasks from the keyframe matcher 612. The keyframe matcher 612 provides this information to the behavior module 618. The behavior module 618 can also receive voxels from the voxel graph 626.

[0135] In one configuration, the behavior module 618 receives motion planning from the motion planner 620 in response to a motion planning request. The behavior module 618 also receives component status from the component control module 622. In response to receiving the component status, the behavior module 618 transmits component commands to the component control module 622. Finally, the component control module 622 receives joint status from the joint control module 624. In response to receiving the joint status, the component control module 622 transmits joint commands to the joint control module 624.

[0136] Figure 7 A method 700 for controlling a robotic device according to one aspect of this disclosure is described. For example... Figure 7 As shown, in an optional configuration, at block 702, the robotic device captures one or more keyframe images while being trained to perform a task. The task may include interaction (such as opening a door) or navigation (such as navigating in an environment).

[0137] After training, at block 704, the robotic device captures an image corresponding to the current view. Images can be captured from one or more sensors, such as RGB-D sensors. Pixel descriptors can be assigned to each pixel in the first group of pixels in the keyframe image and the second group of pixels in the image. The pixel descriptors provide pixel-level information as well as depth information. The user can select the region corresponding to the first group of pixels.

[0138] At block 706, the robotic device identifies a keyframe image with a first set of pixels that matches a second set of pixels in the image. The robotic device can compare the image with multiple keyframes. When the corresponding pixel descriptors match, the robotic device can determine that the first set of pixels matches the second set of pixels.

[0139] In an optional configuration, at block 708, the robotic device determines a difference in at least one of the following: distance, pose, or a combination thereof between the keyframe images and the images. This difference may be referred to as the pose increment. At block 710, the robotic device performs a task related to the keyframe images. In an optional configuration, at block 712, the task is adjusted based on the determined differences between the keyframe images and the images.

[0140] The various operations described above can be performed by any suitable device capable of performing the corresponding function. This device may include various hardware and / or software components and / or modules, including but not limited to circuits, application-specific integrated circuits (ASICs), or processors. Typically, in the operations illustrated in the figures, these operations may have corresponding devices with similar numbering, plus functional components.

[0141] As used herein, the term "determine" encompasses a variety of actions. For example, "determine" can include operations, calculations, processing, derivation, investigation, searching (e.g., searching in a table, database, or other data structure), ascertaining, etc. Additionally, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), etc. Furthermore, "determine" can include parsing, selecting, picking, building, etc.

[0142] As used herein, the phrase “at least one” in the list of items refers to any combination of those items, including a single member. For example, “at least one of a, b, or c” is intended to cover: a, b, c, ab, ac, bc, and abc.

[0143] The various illustrative logic blocks, modules, and circuits described herein can be implemented or performed using a processor configured according to this disclosure, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The processor can be a microprocessor, a controller, a microcontroller, or a state machine specifically configured as described herein. As described herein, the processor can also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor), multiple microprocessors, one or more microprocessors combined with a DSP core, or other such special configurations.

[0144] The steps of the methods or algorithms described in this disclosure can be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module can reside in a storage device or machine-readable medium, including random access memory (RAM), read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store required program code in the form of instructions or data structures and is accessible by a computer. The software module can include a single instruction or multiple instructions and can be distributed across several different code segments, different programs, and across multiple storage media. The storage medium can be coupled to the processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be integrated into the processor.

[0145] The methods disclosed herein include one or more steps or actions for implementing the methods. The method steps and / or actions may be interchanged with each other without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of a particular step and / or action may be modified without departing from the scope of the claims.

[0146] The described functionality can be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system within the device. The processing system can be implemented using a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus may include any number of interconnect buses and bridges. The bus can link together various circuits, including processors, machine-readable media, and bus interfaces. The bus interface can be used to connect network adapters, etc., to the processing system via the bus. The network adapter can be used to implement signal processing functions. For certain aspects, a user interface (e.g., keyboard, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also link various other circuits well known in the art (such as timing sources, peripherals, voltage regulators, power management circuits, etc.), and therefore will not be described further.

[0147] The processor can be responsible for managing the bus and processing, including the execution of software stored on machine-readable media. Whether referring to software, firmware, middleware, microcode, hardware description languages, or others, software should be interpreted as device instructions, data, or any combination thereof.

[0148] In hardware implementations, machine-readable media can be part of a processing system separate from the processor. However, as those skilled in the art will readily understand, machine-readable media or any portion thereof can be external to the processing system. By way of example, machine-readable media may include transmission lines, carrier waves modulated by data, and / or device-separate computer products, all of which can be accessed by the processor via a bus interface. Alternatively or additionally, machine-readable media or any portion thereof may be integrated into the processor (e.g., in cases where it may have caches and / or dedicated register files). Although the various components discussed may be described as having a specific location (e.g., local components), they can also be configured in various ways (e.g., specific components are configured as part of a distributed computing system).

[0149] The processing system may be configured with one or more microprocessors providing processor functionality and external memory providing at least a portion of machine-readable medium, all linked together with other supporting circuitry via an external bus architecture. Alternatively, the processing system may include one or more neuromorphic processors for implementing the neuron and nervous system models described herein. As another option, the processing system may be implemented using an application-specific integrated circuit (ASIC), in which at least a portion of the processor, bus interface, user interface, supporting circuitry, and machine-readable medium is integrated into a single chip, or has one or more field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuitry, or any combination of circuitry capable of performing the various functions described throughout this disclosure. Those skilled in the art will recognize how best to implement the described functions of the processing system, given the specific application and the overall design constraints imposed on the system as a whole.

[0150] Machine-readable media may include multiple software modules. Software modules may include transmission modules and receiving modules. Each software module may reside in a single storage device or be distributed across multiple storage devices. For example, a software module may be loaded from a hard disk drive into RAM when a triggering event occurs. During the execution of a software module, the processor may load certain instructions into a cache to improve access speed. One or more cache lines may then be loaded into a special-purpose register file for processor execution. When the functionality of a software module is referred to below, it will be understood that this functionality is implemented by the processor when executing instructions from that software module. Furthermore, it should be understood that various aspects of this disclosure lead to improvements in the functionality of processors, computers, machines, or other systems implementing these aspects.

[0151] If implemented in software, these functions can be stored or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media includes both computer storage media and communication media, including any storage medium that facilitates the transfer of a computer program from one place to another. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared (IR), radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) are all included in the definition of medium. As used herein, disks include compact discs (CDs), laser discs, optical discs, digital multifunction discs (DVDs), floppy disks, and Blu-ray discs. A disk, which typically reproduces data magnetically, and another disk, which uses lasers, reproduces data optically. Therefore, in some aspects, a computer-readable medium may include a non-transitory computer-readable medium (e.g., a tangible medium). Furthermore, in other aspects, a computer-readable medium may include a transient computer-readable medium (e.g., a signal). Combinations of the above should also be included within the scope of computer-readable media.

[0152] Therefore, certain aspects may include computer program products for performing the operations presented herein. For example, such computer program products may include computer-readable media on which instructions are stored (and / or encoded) that can be executed by one or more processors to perform the operations described herein. For certain aspects, computer program products may include packaging materials.

[0153] Furthermore, it should be understood that modules and / or other suitable means for performing the methods and techniques described herein may be downloaded and / or otherwise obtained by the user terminal and / or base station at the time of application. For example, such a device may be coupled to a server to facilitate the transmission of means for performing the methods described herein. Alternatively, the various methods described herein may be provided via a storage device, such that the user terminal and / or base station can obtain the various methods when the storage device is coupled to or provided to the device. Furthermore, any other suitable techniques for providing the methods and techniques described herein to the device may be utilized.

[0154] It should be understood that the claims are not limited to the precise configuration and components described above. Various modifications, alterations, and variations may be made to the arrangement, operation, and details of the methods and apparatus described above without departing from the scope of the claims.

Claims

1. A method for performing a task using a robotic device, comprising: Capture a test image corresponding to the current view of the robot device in the test environment; The keyframe image is identified based on a comparison between one or more pixel descriptors associated with the keyframe image and one or more pixel descriptors associated with the test image, the keyframe image being a reference image captured by the robotic device in a training environment; Based on the identification of the keyframe, the difference between a first distance, a first pose, or a first combination thereof between the robot device and a first object in the keyframe image and a second distance, a second pose, or a second combination thereof between the robot device and a second object in the image is determined; Parameters associated with the task, which is associated with the keyframe image, are adjusted based on the difference between determining one of the first distance, the first pose, or a first combination thereof, and the second distance, the second pose, or a second combination thereof; and The task is performed by the robotic device based on adjustments to the parameters.

2. The method according to claim 1, wherein: The first pose corresponds to the pose of the robot device relative to a first object in the keyframe image; and The second pose corresponds to the pose of the robotic device relative to the second object in the image.

3. The method according to claim 1, wherein: The first distance corresponds to the distance between the robot device and the first object in the keyframe image; and The second distance corresponds to the distance between the robotic device and the second object in the image.

4. The method according to claim 1, wherein: Each pixel descriptor in the one or more pixel descriptors associated with the keyframe image and each pixel descriptor in the one or more pixel descriptors associated with the test image are associated with a set of values ​​corresponding to pixel-level information and depth information.

5. The method according to claim 1, wherein, The task includes one or both of interacting with objects or navigating in an environment.

6. The method according to claim 1, wherein, The parameter is associated with one or more of the speed or location associated with the task.

7. The method according to claim 1, wherein, The robotic device is trained to perform the task based on human demonstrations.

8. An apparatus for performing a task via a robotic device, comprising: processor; as well as A memory coupled to the processor and storing instructions that, when executed by the processor, are operable to cause the device to: Capture a test image corresponding to the current view of the robot device in the test environment; The keyframe image is identified based on a comparison between one or more pixel descriptors associated with the keyframe image and one or more pixel descriptors associated with the test image, the keyframe image being a reference image captured by the robotic device in a training environment; Based on the identification of the keyframe, the difference between a first distance, a first pose, or a first combination thereof between the robot device and a first object in the keyframe image and a second distance, a second pose, or a second combination thereof between the robot device and a second object in the image is determined; Parameters associated with the task, which is associated with the keyframe image, are adjusted based on the difference between determining one of the first distance, the first pose, or a first combination thereof, and the second distance, the second pose, or a second combination thereof; and The task is performed by the robotic device based on adjustments to the parameters.

9. The apparatus according to claim 8, wherein: The first pose corresponds to the pose of the robot device relative to a first object in the keyframe image; and The second pose corresponds to the pose of the robotic device relative to the second object in the image.

10. The apparatus according to claim 8, wherein: The first distance corresponds to the distance between the robot device and the first object in the keyframe image; and The second distance corresponds to the distance between the robotic device and the second object in the image.

11. The apparatus according to claim 8, wherein: Each pixel descriptor in the one or more pixel descriptors associated with the keyframe image and each pixel descriptor in the one or more pixel descriptors associated with the test image are associated with a set of values ​​corresponding to pixel-level information and depth information.

12. The apparatus according to claim 8, wherein, The task includes one or both of interacting with objects or navigating in an environment.

13. The apparatus according to claim 8, wherein, The parameter is associated with one or more of the speed or location associated with the task.

14. The apparatus according to claim 8, wherein, The robotic device is trained to perform the task based on human demonstrations.

15. A non-transitory computer-readable medium having program code recorded thereon for performing tasks via a robotic device, the program code being executed by a processor and including means for: Program code for capturing test images corresponding to the current view of the robot device in the test environment; Program code for identifying the keyframe image based on a comparison between one or more pixel descriptors associated with the keyframe image and one or more pixel descriptors associated with the test image, the keyframe image being a reference image captured by the robotic device in a training environment; Program code for determining, based on the identification of the keyframe, the difference between a first distance, a first pose, or a first combination thereof between a robot device and a first object in the keyframe image, and a second distance, a second pose, or a second combination thereof between a robot device and a second object in the image; Program code for adjusting parameters associated with a task, which is associated with the keyframe image, based on determining the difference between one of the first distance, the first pose, or a first combination thereof and one of the second distance, the second pose, or a second combination thereof; as well as Program code used to perform the task via the robotic device by adjusting the parameters.

16. The non-transitory computer-readable medium according to claim 15, wherein: The first pose corresponds to the pose of the robot device relative to a first object in the keyframe image; and The second pose corresponds to the pose of the robotic device relative to the second object in the image.

17. The non-transitory computer-readable medium according to claim 15, wherein: The first distance corresponds to the distance between the robot device and the first object in the keyframe image; and The second distance corresponds to the distance between the robotic device and the second object in the image.

18. The non-transitory computer-readable medium according to claim 15, wherein: Each pixel descriptor in the one or more pixel descriptors associated with the keyframe image and each pixel descriptor in the one or more pixel descriptors associated with the test image are associated with a set of values ​​corresponding to pixel-level information and depth information.

19. The non-transitory computer-readable medium according to claim 15, wherein, The task includes one or both of interacting with objects or navigating in an environment.

20. The non-transitory computer-readable medium according to claim 15, wherein, The parameter is associated with one or more of the speed or location associated with the task.

Citation Information

Patent Citations

  • Robot navigation method based on global map and robot using navigation method to navigate

    CN108072370A

  • Robot device, method of controlling the same, computer program, and robot system

    US20160375585A1