Autonomous Task Execution Based on Perspective Embedding

By using pixel-level information and depth data to identify matching keyframes, the method improves the accuracy of task execution for autonomous agents, overcoming limitations of feature-based approaches.

JP7693648B2Active Publication Date: 2025-06-17TOYOTA JIDOSHA KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022503936
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-13
Filing Date
2020-06-10
Publication Date
2025-06-17
Estimated Expiration
2040-06-10

AI Technical Summary

Technical Problem

Conventional autonomous agents rely on feature-based approaches for task execution, which can be limited by the uniqueness and quantity of features in an image, leading to reduced accuracy in object detection and task execution.

Method used

The method involves capturing an image corresponding to the current vision of a robotic device, identifying a keyframe image with matching pixels, and executing a task associated with the keyframe image, using pixel-level information and depth data for improved matching accuracy.

Benefits of technology

This approach enhances the accuracy of the vision system by comparing pixel descriptors rather than just unique features, allowing the robotic device to execute tasks more reliably even with varying environmental conditions and object positions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693648000001
    Figure 0007693648000001
  • Figure 0007693648000002
    Figure 0007693648000002
  • Figure 0007693648000003
    Figure 0007693648000003
Patent Text Reader

Abstract

A method for controlling a robotic device is presented. The method includes capturing an image corresponding to a current vision of the robotic device. The method also includes identifying a keyframe image having a first set of pixels that matches a second set of pixels of the image. The method further includes performing, by the robotic device, a task corresponding to the keyframe image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 877,792, filed on June 23, 2019, entitled "Key Frame Matcher", U.S. Provisional Patent Application No. 62 / 877,791, filed on June 23, 2019, entitled "Visual Instruction and Repetition for Operations - Instruction VR", and U.S. Provisional Patent Application No. 62 / 877,793, filed on June 23, 2019, entitled "Visualization", and claims the benefit of U.S. Patent Application No. 16 / 570,618, filed on September 13, 2019, entitled "Autonomous Task Execution Based on Viewpoint Embedding", the content of which is incorporated herein by reference.

[0002] Certain aspects of the present disclosure generally relate to robotic devices, and more particularly to systems and methods that use environmental memory to interact with the environment.

Background Art

[0003] Autonomous agents (e.g., vehicles, robots, drones, etc.) and semi - autonomous agents use machine vision to analyze regions of interest in the surrounding environment. During operation, an autonomous agent may rely on a neural network trained to identify objects present in the regions of interest in an image of the surrounding environment. For example, the neural network may be trained to identify and track objects captured by one or more sensors such as a light detection and ranging (LIDAR) sensor, a sonar sensor, an RGB camera, an RGB - D camera, etc. The sensors may be connected to or in communication with a device such as an autonomous agent. An object detection application for an autonomous agent may analyze sensor image data to detect objects (e.g., pedestrians, people on bicycles, other vehicles, etc.) from the scenery surrounding the autonomous agent.

[0004] In a conventional system, autonomous agents are trained to execute tasks using a feature-based approach. That is, tasks are executed based on features extracted from the current visual image of the robotic device. The features of the image are limited to the unique features of the objects in the image. If the image does not have many features, the accuracy of the conventional object detection system can be limited. It is desirable to improve the accuracy of the vision system.

Summary of the Invention

[0005] In one aspect of the present disclosure, a method of controlling a robotic device is disclosed. The method includes capturing an image corresponding to the current vision of the robotic device. The method also includes identifying a keyframe image comprising a first set of pixels that matches a second set of pixels of the image. The method further includes executing a task for which the robotic device image corresponds to the keyframe image.

[0006] In another aspect of the present disclosure, a non-transitory computer-readable medium storing non-transitory program code is disclosed. The program code is for controlling a robotic device. The program code includes program code for capturing an image corresponding to the current vision of the robotic device when executed by a processor. The program code also includes program code for identifying a keyframe image comprising a first set of pixels that matches a second set of pixels of the image. The program code further includes program code for executing a task for which the robotic device corresponds to the keyframe image.

[0007] Another aspect of the present disclosure relates to a robotic device. The device has a memory and one or more processors connected to the memory. The processor captures an image corresponding to the current vision of the robotic device. The processor also identifies a keyframe image comprising a first set of pixels that matches a second set of pixels of the image. The processor further executes a task corresponding to the keyframe image.

[0008] The features and technical advantages of the present disclosure have been broadly outlined above in order for the following detailed description to be better understood. Additional features and advantages of the present disclosure will be described below. It should be understood by those skilled in the art that the present disclosure can be readily used as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. It should also be recognized by those skilled in the art that such equivalent configurations do not depart from the teachings of the present disclosure as defined by the appended claims. The new features that are considered to be characteristics of the present disclosure will be better understood from the following description when considered in conjunction with the accompanying drawings, along with further objectives and advantages. However, it should be clearly understood that each drawing is provided for purposes of illustration and description only and is not intended to define the limits of the present disclosure.

Brief Description of the Drawings

[0009] The functions, nature, and advantages of the present disclosure will become more apparent from the detailed description provided below when considered in combination with the corresponding drawings in which like reference characters refer to the whole.

[0010]

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Mode for Carrying Out the Invention

[0011] The following detailed description related to the accompanying drawings is intended to explain various configurations and is not intended to present a single configuration for implementing the concepts described in this specification. The detailed description includes specific details for the purpose of providing a complete understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts can be implemented without these specific details. In some cases, well-known structures and components are shown in block diagrams to avoid obscuring such concepts.

[0012] Based on the teachings, it should be understood by those skilled in the art that the scope of the present disclosure is intended to include any aspect of the present disclosure, whether implemented independently or in combination with other aspects of the present disclosure. For example, an apparatus may be implemented using any number of the aspects disclosed, or a method may be implemented. In addition, the scope of the present disclosure is intended to include, in addition to the various aspects disclosed in the present disclosure, such apparatus or methods implemented using other structures and functions, or structures and functions. It should be understood that any aspect of the present disclosure may be embodied by one or more elements of the claims.

[0013] As used herein, the term "exemplary" is used in the sense of "serving as an example, instance, or illustration". Any aspect of the specification described as "exemplary" should not necessarily be understood as being preferred or advantageous compared to other aspects.

[0014] While specific embodiments are described in this specification, numerous variations and substitutions to these embodiments are included within the scope of the present disclosure. Although some benefits and advantages of the preferred embodiments are described, the scope of the present disclosure is not intended to be limited to any specific benefit, use, or purpose. Rather, the embodiments of the present disclosure are intended to be widely applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated by way of example in the figures and description of the preferred embodiments. The detailed description and drawings are for the purpose of describing the present disclosure only and not for the purpose of limitation, and the scope of the present disclosure is defined by the appended claims and their equivalents.

[0015] Autonomous agents and semi-autonomous agents may perform tasks on detected objects. For example, an agent may navigate through an environment, identify objects in the environment, and / or interact with objects in the environment. In the present application, a robot or robotic device is an autonomous agent or a semi-autonomous agent. For simplicity, a robotic device is referred to as a robot.

[0016] Aspects of the present disclosure are not limited to a particular type of robot. Various types of robotic devices having one or more vision systems that provide visual / visualization output are contemplated. The visual output may be provided to an operator as a supplied video.

[0017] According to aspects of the present disclosure, a robot is capable of performing a movement operation. That is, a robot has the ability to change the position of an end effector as a function of a joint configuration. In one configuration, a robot is equipped with a function of automatically controlling and planning its entire body. As a result, a human operator can perform seamless end effector operations in the task space in virtual reality (VR) with little or no concern for kinematic constraints or the posture of the robot.

[0018] The robot is equipped with one or more vision RGB-D (red-green-blue and depth) sensors on a pan / tilt head and can provide important context to a human operator in virtual reality. An RGB-D image is a combination of an RGB image and a corresponding depth image. A depth image is an image channel in which each pixel is associated with the distance between the image plane and the corresponding object on the RGB image.

[0019] The robot may travel in the environment based on an image obtained from a provided video of the environment (e.g., a provided image). The robot may also perform a task based on an image from the provided video. In a conventional vision system, the robot identifies an image (e.g., an image of an object) by a feature point-based approach.

[0020] In a conventional system, the robot is trained to perform tasks by a feature-based approach. That is, the task is performed based on features extracted from the current visual image of the robot. In a feature point-based approach, a test image (e.g., the current image) is identified based on the similarity between the test image and a reference image.

[0021] In particular, in a conventional vision system, the features of a test image are compared with the features of a reference image. The test image is identified when the features match the features of the reference image. In some cases, a conventional robot cannot perform the same task if the starting point and / or the direction / position of an object do not match a programmed or taught task.

[0022] A feature is an obvious (e.g., unique) characteristic of an image (e.g., a corner, an edge, a high-contrast region, a low-contrast region, etc.). An image may include multiple features. A descriptor encodes the features in an image. A feature vector may be an example of a descriptor. In a conventional vision system, the descriptor was limited to encoding unique features of an image (e.g., an object in the image).

[0023] In most cases, the descriptor is robust to image transformations (e.g., localization, scale, brightness, etc.). That is, conventional vision systems can identify objects under various conditions such as various lighting conditions. Nevertheless, since the descriptor is limited to encoding unique features of the image, the robustness is limited. To improve the ability of a robot to identify an image and execute a task, it is desirable to improve the robustness of the vision system of the image.

[0024] Aspects of the present disclosure relate to assigning descriptors to pixels of an image. That is, in contrast to conventional vision systems, in aspects of the present disclosure, the descriptor is not limited to unique features. Thus, the accuracy of the vision system can be improved. In particular, conventional vision systems were limited to comparing obvious features (e.g., characteristics) of a test image with obvious features of a reference image. An image contains more pixels than features. Therefore, by comparing pixels instead of features, the accuracy of the comparison is increased.

[0025] That is, in one configuration, the vision system compares the pixels of a test image with the pixels of a reference image. The reference image may be referred to as a keyframe. The keyframe may be acquired during training and is invariant to pose and image transformations. The vision system determines that the test image matches the keyframe when the number of matching pixels is greater than a threshold.

[0026] In one configuration, the vision system compares the pixel descriptors of the pixels to identify the matching pixels between the test image and the keyframe. The pixel descriptor includes pixel-level information and depth information. The pixel-level information includes information such as the RGB values of the pixel and the context of the pixel among the image / surrounding pixels.

[0027] The context of a pixel contains information to eliminate the ambiguity of the pixel. The size of the context depends on the pixel and is learned during training. For example, a pixel in the center of a plain white wall may require more context compared to a highly textured logo on a bottle. The robot learns how much context is needed without having to create this functionality manually. Depth information corresponds to the pixel and indicates the distance to the sensor used to acquire the image of that surface from the surface.

[0028] During training, when the robot executes a task, one or more sensors acquire image data of the scene. Keyframes are generated from the acquired image data. The keyframes may be conceptualized as a memory of the environment when the robot was performing a specific task. That is, the keyframes may be set as an anchor to a task or location in the environment. After testing, when the current image matches the keyframe by the robot, the trained function corresponding to the keyframe may be executed.

[0029] Descriptors may be assigned to the pixels of the keyframe. The descriptor may be a vector or array of values. For example, the descriptor may be a vector with 16 elements. The descriptor (e.g., pixel descriptor) may be used to identify matching images.

[0030] After training, the robot may be initialized to execute one or more learned tasks. The robot may collect raw image data when executing the commanded task. The raw image data is mapped at the pixel level to the pixels in the keyframe when the robot passes through and / or manipulates objects in the environment. The correspondence between the current image and the keyframe is determined based on the similarity of the pixel descriptors.

[0031] That is, the robot compares the values of the pixel descriptors of the current image with those of the keyframe. The confidence level of the comparison indicates the likelihood that the pixels of the current image match the pixels of the keyframe. A rigid body assumption may be used to determine the confidence level of the match. That is, the confidence level may be determined based on the distances between pairs of points on the object in the current image and the keyframe. The confidence level increases when the distances between points on the object remain constant regardless of the poses of both images. The pose in the image is the pose of the robot when the image was taken. If the distances between points on the object are different between images with different poses, the confidence level decreases.

[0032] The robot may visualize the confidence level of the match using the diagonal / parallel correspondence (e.g., non-crossing relative relationship) between the current image data and multiple pixels of the keyframe. The visualized result may be transmitted to a remote operation center. The crossing relationship indicates a non-match between the current image data and the keyframe.

[0033] In one configuration, the robot is located in an environment and image data of the scenery is collected. Then the robot is controlled (e.g., through a virtual reality interface) to perform a task. By restricting the view of a human operator in virtual reality to the view of the robot during training, the ability of the robot to perform a task alone is improved.

[0034] The robot may be instructed to execute actions / tasks by parameterizing the actions performed by the operator through a virtual reality interface. For example, the virtual reality interface may include the use of paddles, hand-held controllers, paintbrush tools, wiping tools, and / or placement tools that the operator operates while wearing a headset that renders the VR environment. Thus, a human operator teaches a set of parameterized primitives (or actions) rather than directly teaching actions in the task space. The parameterized primitives combine a collision-free motion plan and hybrid (position and force) Cartesian control to reduce the parameters to be taught and provide robustness during execution.

[0035] Parameterized motion is to learn by dividing the task into a small number of separate chunks of motion. Each motion is defined by a set of parameters such as joint angle changes, rotation angles, etc. The values of these parameters may be configured and updated based on the situation of the robot when executing the task.

[0036] Parameterized motions can be learned and extracted from one learned task and combined with other tasks to form larger tasks. A parameterized motion such as opening a door with a rotating handle may be implemented as performing the action of opening any door handle (e.g., a door that requires a 30-degree rotation, or a door that requires a 60-degree rotation, or more rotations). For example, the rotation angle can be one of the parameters that define the parameterized motion for opening a door with a rotating door handle.

[0037] To execute a task, the robot may be located in the same or a similar environment (relative to the initial position during training). The robot may be located at different starting points and may assume relatively different initial poses (e.g., adjusted to starting positions with different joint angles). The robot may be tasked with performing the same task (e.g., a set of parameterized actions), such as picking up a bottle, opening a cabinet, and placing the bottle in the cabinet, without control by a human operator. For example, the robot may perform the same task by updating the parameters of the actions taught in a virtual reality-controlled sequence. The parameters may be updated based on the current pose and / or position of the robot as compared to the pose and / or position used during training.

[0038] As discussed, to update the parameters, the robot takes an initial image of the scene and maps pixels and / or descriptors from the new image to the keyframe. The mapping defines the relative transformation between the initial image and the keyframe. The keyframe can be mapped to the new image using the relative transformation. The relative transformation may be defined by changes in the position of the robot on the x-axis, y-axis, and z-axis, roll, pitch, and yaw. The relative transformation may be used to update the parameters of the parameterized actions from the taught parameters to the observed situation.

[0039] Relative transformations may be applied to parameterized operations. By applying relative transformations to parameterized operations, a robot can perform the same task as previously taught even if its initial position and / or orientation changes. The robot system may continuously map pixel and / or dense neural network descriptors from the current scene to those from keyframes, spacing them out (e.g., continuously) from the parameterized operations. For example, relative transformations may be applied to taught actions defined by a set of parameterized operations such as pulling out and opening a drawer, opening a door, picking up a cup or a bottle, etc.

[0040] In some aspects, an action may be related to the overall scene and / or object-specific. For example, to perform the action of picking up a bottle, it may be necessary to use keyframes related to the overall scene to drive to the location of the bottle. Once approaching the bottle, keyframes specific to the bottle may be analyzed independently of the environment.

[0041] The driving operation may be used to move the robot from one location to another. This may enable the robot to identify the position of an object that can be located anywhere in the environment during the training of the "pick up" action and perform tasks such as "pick up" regardless of the position of the bottle. The manipulating operation may be used to move parts of the robot (e.g., the torso and / or arms) to contact the desired object.

[0042] FIG. 1 shows an example in which an operator 100 controls a robot 106 according to an aspect of the present disclosure. As shown in FIG. 1, the operator 100 includes a vision system 102 and an operation controller 104 (e.g., a gesture following system) for controlling the robot 106. In this example, the operator 100 may control the robot 106 to perform one or more tasks. In the example of FIG. 1, the robot 106 is trained in the kitchen 108. Aspects of the present disclosure are not limited to training the robot 106 in the kitchen 108, and other environments are also considered.

[0043] The vision system 102 may not only capture the vision of the operator 100 but also provide a supply video. The operator 100 may be at a position away from the robot 106. In this example, the robot 106 is located in the kitchen 108, and the operator 100 is located at a location different from the kitchen 108, such as the robot control center 114.

[0044] The vision system 102 may provide a supply video of the robot 106. For example, the vision system 102 may provide a view of the kitchen 108 based on the front view of the robot 106. Other viewpoints, such as a 360° view, may be provided. The viewpoint is provided using one or more vision sensors, such as a video camera of the robot 106. The vision system 102 is not limited to the headset as shown in FIG. 1. The vision system 102 may be a monitor 110, an image projector, or any other device capable of displaying the supply video from the robot 106.

[0045] One or more actions of the robot 106 may be controlled via the motion controller 104. For example, the motion controller 104 may capture the gestures of the operator 100, and the robot 106 may imitate the captured gestures. The operator 100 may control the movement, limb movements, and other actions of the robot 106 via the motion controller 104. Aspects of the present disclosure are not limited to capturing the gestures of the operator 100 via the motion controller 104. Other types of gesture capture systems are also conceivable. The operator 100 may control the robot 106 via the wireless connection 112. Additionally, the robot 106 may provide feedback such as a feed video to the operator 100 via the wireless connection 112.

[0046] The robot 106 may be trained to travel in a specific environment (e.g., the kitchen 108) and / or similar environments. Additionally, the robot 106 may be trained to perform tasks on objects in a specific environment and / or similar objects in any environment. For example, the robot 106 may be trained to open and close the drawers in the kitchen 108. The training may be implemented only for the drawers in the kitchen 108 and / or for similar drawers in any environment, such as the drawers in another kitchen.

[0047] FIG. 2A shows an example of a robot 200 in an environment 202 according to an aspect of the present disclosure. For clarity, FIG. 2A is a top view of the environment 202. As shown in FIG. 2A, the environment 202 includes a dining table 204, a sink 206, a drawer 208 containing spoons 218, and a counter 210.

[0048] In the example of FIG. 2A, the open drawer 208 is within the field of view 224 (e.g., in the line of sight) of the robot 200. The open drawer 208 may be displayed in the current image data captured by the sensors of the robot 200. In one configuration, the robot 200 compares the current image data with keyframe data to determine whether an action should be executed.

[0049] Figure 2B shows an example of a current image 252 taken during training by robot 200. As shown in Figure 2B, the current image 252 of the field of view 224 of robot 200 includes sink 206 and counter 210. In one configuration, to determine subsequent actions and / or to determine the position of the robot in the environment, robot 200 compares the pixel descriptors of the current image 252 with the pixel descriptors of the keyframe.

[0050] Figure 2C shows an example of a keyframe 260 taken during training by robot 200. As shown in Figure 2C, keyframe 260 includes sink 206, counter 210, drawer 208, and spoon 218. Keyframe 260 may be taken in an environment different from, or the same environment (e.g., position) as, the environment in which robot 200 captured the current image 252. The angle and / or depth of keyframe 260 may be different from, or the same as, the angle and / or depth of the current image 252. For simplicity, in Figures 2B, 2C, and 2D, the angle and depth of keyframe 260 and the current image 252 are the same.

[0051] In one configuration, the user annotates areas considered to be anchor regions in keyframe 260. For example, in Figure 2C, spoon 218 and drawer 208 are movable, dynamic objects. In contrast, the upper part of the faucet or the edge of the sink is a static object. Thus, the user may annotate the static objects so that the pixels of the static objects are used to determine whether the current image matches the keyframe. Annotating the pixels of static objects (e.g., anchors) can improve the accuracy of the pose delta.

[0052] Figure 2D shows an example of comparing the current image 252 with the keyframe 260 according to an aspect of the present disclosure. In one configuration, robot 200 compares the pixel descriptors (e.g., the values of the pixel descriptors) of the current image 252 with the pixel descriptors of the keyframe 260. Line 270 between the current image 252 and the keyframe 260 indicates the pixels that match between the two images.

[0053] If two or more lines 270 do not intersect, the images may match. On the other hand, if two or more lines 270 intersect, a mismatch may be identified. The current image 252, key frame 260, and lines 270 may be output to the display of the control center. The number of pixels compared in FIG. 2D is used for illustrative purposes. Aspects of the present disclosure are not limited to only comparing the number of pixels compared in FIG. 2D.

[0054] As discussed, after training, the robot may be initialized to perform one or more learned tasks. For example, the robot may be tasked with traveling in the environment and / or manipulating objects in the environment. The robot collects raw image data when performing the commanded task. The raw image data is compared to the key frame.

[0055] In one configuration, the confidence of the pixels that match between the current image data and the key frame is visualized. The display may also visualize the comparison of the pixel descriptor with various thresholds. For example, if the depth in the image is greater than the threshold, the image may be rejected.

[0056] If one or more key frames match the current image data, an action associated with the matching key frame may be executed. For example, in FIG. 2D, in determining whether the current image 252 matches the key frame 260, the robot may perform a task such as opening or closing a drawer. As another example, the robot may open or close a faucet. Additionally, or alternatively, when the previous action is completed, the execution of the action may be added to the queue. As another example, the robot may determine its current position based on the key frame and use it to travel a path from the current position to another position. As the robot moves in the environment or the position of an object in the environment changes, a new key frame may be referenced.

[0057] In one configuration, one or more pixels are filtered before comparing the pixel descriptors. A function may be executed to filter outliers. For example, a random sample consensus (RANSAC) function may be used to filter pixel outliers by randomly sampling the observed data. Filtering the pixels reduces noise and improves accuracy.

[0058] As discussed, the depth and / or angle of the current image may differ from the depth and / or angle of the matching keyframe. Thus, in one configuration, a relative transformation between the matching pixels of the current image data and the pixels of the keyframe is determined. That is, a pose delta may be determined based on a comparison of the current image and the keyframe. The delta is the change in the robot's pose from the keyframe to the current image.

[0059] The pose delta may be determined based on the depth information of the pixel descriptor. The depth information is the distance from the sensor to the object corresponding to the pixel. The pose delta is the change in (x, y, z) coordinates, roll, pitch, and yaw between the current image and the keyframe. That is, the pose delta explains how the image has changed from the keyframe to the current image based on the movement of the sensor.

[0060] The accuracy of the pose delta increases in relation to an increase in the number of matching pixels. In the present disclosure, the number of matching pixels may be greater than the number of matching features. Thus, the aspects of the present disclosure improve the accuracy of determining the pose delta compared to conventional systems that compare images based only on features. A least squares function may be used to determine the pose delta.

[0061] That is, the posture may be determined by minimizing the reprojection error squared. In particular, the pixels in the keyframe should be the pixels that exactly match those seen in the current image. The minimization of the reprojection error squared can be represented as a non-linear least squares problem solved for six unknown parameters (x, y, z, roll, pitch, yaw). A RANSAC loop may be used to reduce the sensitivity to outliers.

[0062] For updating the parameters of the operation to be performed, a pose delta (e.g., relative transformation) between one or more numerical values of the pixel descriptors of the pixels may be used. For example, if the pose delta indicates that the robot is 1 foot (about 30.5 centimeters) away from the object compared to the position defined at the keyframe, the operation may be updated considering the current position (e.g., pose) of the robot. The position defined at the keyframe is the position used when training the robot to perform the task.

[0063] Figure 3 shows an example of an output 300 from a vision system according to an aspect of the present disclosure. As shown in Figure 3, the output 300 includes a current image 302, a keyframe 308, a histogram 306, and a threshold window 304. The output 300 may be displayed at a position remote from the current position of the robot. For example, the output 300 may be displayed at a control center that controls the operation of the robot.

[0064] The current image 302 is an image from the current vision of the robot. The keyframe 308 is one keyframe from a set of keyframes that are compared with the current image 302. The histogram 306 shows the confidence scores regarding the match from the current image 302 to each keyframe in the set of keyframes.

[0065] During training, in the driving task, the robot captures key frames at specific points from start to goal along the path. The key frames captured along the path are used as a set of key frames in the driving task. During testing, the robot compares the current image with the set of key frames to identify its current position.

[0066] Based on the number of matching pixels (e.g., pixel descriptors), the robot determines the confidence level of the match from the current image to a specific key frame. If the confidence level of the match is higher than the threshold, the robot executes the task associated with the matching key frame. The task may include position determination (e.g., determining the current position of the robot). In the example of Figure 3, the robot determines that it is at the position corresponding to the position of the key frame with the highest confidence level of 310 along the path.

[0067] The histogram 306 may vary according to the task. The histogram 306 in Figure 3 is for the driving task. The histogram 306 may be different when the robot is performing an operation (e.g., object manipulation) task. The set of key frames for the operation task may have fewer key frames compared to the set of key frames for the driving task.

[0068] For example, when opening a cabinet, the robot may be trained to open it when the cabinet is closed or slightly open. In this example, the histogram may have two key frame match bar graphs in the bar graph. One bar graph indicates the confidence level that the current image matches the key frame of the closed cabinet, and the other bar graph may indicate the confidence level that the current image matches the key frame of the partially open cabinet.

[0069] The pose confidence window 304 indicates the confidence in pose consistency obtained from the pose matcher. In one configuration, before determining the pose delta, the robot determines whether a number of criteria satisfy one or more thresholds. The criteria may be based on depth, pose, creak, error, and other factors.

[0070] In one configuration, the robot determines the number of pixels having a depth value. For surfaces such as glass or shiny surfaces, the robot may not be able to determine the depth. Thus, the pixels corresponding to these surfaces may not have a depth value. The number of pixels having a depth value may be determined before filtering the pixels. If the number of pixels having a depth value is greater than a threshold, the depth criterion is satisfied. The pose confidence window 304 may include a bar graph comparing a bar graph 320 of the number of pixels having a depth value and a bar graph 322 indicating the threshold. The bar graph may be color-coded.

[0071] In addition, or alternatively, the bar graph of the pose confidence window 304 may include a comparison of a bar graph of the number of features that align between the current image and the keyframe and a bar graph of an alignment threshold. To determine the accuracy of the match between the keyframe and the current image, the robot may apply a pose transformation and determine how many features of the objects in the current image align with the features of the objects in the keyframe, and vice versa. If the features are associated with non-static objects, the features may not align. For example, the keyframe may include a cup that no longer exists in the current image. Thus, the features of the cup may not align between the images. The number of aligned features may be called a creak inlier. A pose delta may be generated if the creak inlier is greater than a threshold.

[0072] Additionally, or alternatively, the bar graph of the pose confidence window 304 may include a comparison of the bar graph of the number of features used to determine the pose and the bar graph of the pose threshold. That is, a certain number of features should be used to determine the pose. If the number of features used to determine the pose is less than the pose threshold, the pose delta may not be calculated. The number of features used to determine the pose may be referred to as pose inliers.

[0073] The bar graph of the pose confidence window 304 may also include a comparison of the bar graph of the root mean square (RMS) error of the pose delta and the threshold. If the RMS error is less than the RMS error threshold, the pose delta may be acceptable. If the RMS error is greater than the RMS error threshold, the user or robot may not use the pose delta to execute the task.

[0074] The pose confidence window 304 is not limited to a bar graph, and other graphs or images may be used. The criteria for the pose confidence window 304 are not limited to the discussed criteria, and other criteria may be used.

[0075] As discussed, aspects of the present disclosure relate to a mobile manipulation hardware and software system capable of autonomously performing human-level tasks in a real-world environment after being demonstrated a task by a human within virtual reality. In one configuration, a mobile manipulation robot is used. The robot may include whole-body task space hybrid position / force control. Additionally, as discussed, parameterized primitives linked to a dense visual embedding representation of a robustly learned landscape are taught to the robot. Finally, a task graph of the taught actions may be generated.

[0076] Rather than programming or training a robot to recognize a set of fixed objects or perform a predefined task, aspects of the present disclosure allow a robot to learn new objects and tasks from human demonstrations. The learned tasks may be autonomously executed by the robot under naturally varying conditions. The robot can be taught, from a single example, to associate a given set of actions with any landscape and objects without using previous object models or maps. The vision system may be trained offline using existing labeled and unlabeled datasets, and the rest of the system may function without additional training data.

[0077] In contrast to conventional systems that directly teach actions in task space, aspects of the present disclosure teach a parameterized set of actions. These actions combine a collision-free motion plan with a hybrid (position and force) Cartesian control of the end effector to minimize the taught parameters and provide robustness during execution.

[0078] In one configuration, an embedding for trained dense vision specialized for a task is computed. This pixel-wise embedding links the parameterized actions to the landscape. This link allows the system to handle a wide variety of environments with high robustness in exchange for generalization to new situations.

[0079] The actions of the task may be taught independently using visual input conditions and success-based end conditions. The actions may be linked together within a dynamic task graph. Since the actions are linked, the robot may reuse actions to execute a task sequence.

[0080] The robot may have multiple degrees of freedom (DOF). For example, the robot may have 31 degrees of freedom (DOF) divided into five subsystems: a chassis, a torso, a left arm, a right arm, and a head. In one configuration, the chassis includes four drive-steerable wheels (e.g., a total of 8 degrees of freedom) that achieve "quasi-holonomic" mobility. The drive / steer actuator package may include various motors and gearheads. The torso may have 5 degrees of freedom (yaw, pitch, pitch, pitch, yaw). Each arm may have 7 degrees of freedom. The head may have 2 degrees of freedom for pan / tilt. Each arm may also include a 1-degree-of-freedom gripper with under-driven fingers. Aspects of the present disclosure are not limited to the robots discussed above. Other configurations are possible. In one example, the robot may include a custom tool such as a sponge or a mop.

[0081] In one configuration, the robot is integrated with a force / torque sensor for measuring the interaction force with the environment. For example, the force / torque sensor may be arranged at the wrist of each arm. The head may be integrated with a perception sensor that provides a wide field of view and also provides a VR context for a human or the robot to perform tasks.

[0082] Aspects of the present disclosure provide several levels of abstraction for robot control. In one configuration, at the lowest control level, real-time coordinated control of all degrees of freedom of the robot is provided. Real-time control may include joint control and component control. Joint control implements low-level device communication and exposes device commands and states in a general form. In addition, joint control supports actuators, force sensors, and inertial measurement devices. Joint control may be configured at runtime to support different robots.

[0083] By component control, the robot is divided into components (such as the right arm, head, etc.), and by providing a set of parameterized controllers for each component, higher-level collaborative actions of the robot can be handled. Component control may provide controllers for joint position and velocity, joint admittance, camera vision, vehicle position and velocity, and posture, velocity, and admittance control in the hybrid task space.

[0084] Task space control of the end effector enables the abstraction of robot control in another dimension. This level of abstraction solves the problem of the robot posture for achieving the desired motion. The inverse kinematics (IK) of the whole body for hybrid Cartesian control is formed and solved as a quadratic program. There may be linear constraints on the components regarding joint position, velocity, acceleration, and gravitational torque.

[0085] The IK of the whole body may be used for the motion plan to reach the posture goal in Cartesian coordinates. In one configuration, occupied voxels of the environment are fitted with spheres or capsules. To avoid collisions between the robot and the world, collision constraints of the voxels are added to the quadratic program of the IK. In the quadratic program of the IK, sampling in Cartesian space is performed as a steering function between nodes, and the motion plan may be performed using a rapidly-exploring random tree (RRT).

[0086] The plan in Cartesian space results in natural and direct motions. By using the quadratic program of the IK as the steering function, the reliability of the plan can be improved, and the same controller may be used for both the plan and execution to reduce the discrepancy between them. Similarly, the motion plan towards the joint position goal and the joint position controller by component control acting as the steering function are combined to use the RRT.

[0087] Parameterized operations are defined at the next level of abstraction. In one configuration, the parameterized operations are primitive actions that can be parameterized and combined to accomplish tasks. Operations can include manipulation actions such as grasping, lifting, placing, pulling, retracting, wiping, direct control, driving actions such as moving joints, driving by speed command, driving by position command, path following while performing active obstacle avoidance, and preparatory actions such as visually stopping, without limitation.

[0088] Each operation can have one or more different types of actions, such as the movement of one or more joints of a robot part or movement in Cartesian coordinates. Each action can use different control methods such as position, speed, or admittance control, and can choose to use an operation plan to avoid external obstacles. The operations of the robot, whether or not using an operation plan, avoid self-collisions and satisfy the operation control constraints.

[0089] Each operation is parameterized by different actions, or alternatively the actions may have their own parameters. For example, the grasping operation may consist of four parameters: the gripper angle, the 6D approach, the grasp, and the (optional) posture of the gripper when lifted. In this example, these parameters define the following predefined sequence of actions: (1) open the gripper to the desired gripper angle; (2) plan and execute a collision-free path to the 6D approach posture; (3) move the gripper to the 6D grasping posture and stop when contact is made; (4) close the gripper; and (5) move the gripper to the 6D lift pose.

[0090] The abstraction of the final level of control is the task. In one configuration, the task is defined as a sequence of actions that enables the robot to perform operations and navigate in a human environment. The task graph (see Figure 5) is a valid, periodic or aperiodic graph with different tasks as nodes, different movement situations as edges, and including anomaly detection and recovery from anomalies. The edge situations include the execution status of each action for handling different objects and environments, the inspection of the object in hand using a force / torque sensor, voice commands, and the matching with keyframes.

[0091] According to an aspect of the present disclosure, a perception pipeline for the robot to understand the surrounding environment is designed. The perception pipeline also enables the robot to recognize which action to take based on the taught task. In one configuration, a fused RGB-D image is created by projecting a plurality of depth images of a high-resolution color stereo pair onto one field-of-view image (e.g., the left image of a wide field of view). The system runs a set of deep neural networks to provide various pixel-level classifications and feature vectors (e.g., embeddings). Based on the visual features called from the taught sequence, the pixel-level classifications and feature vectors are accumulated into a temporary 3D voxel representation. The pixel-level classifications and feature vectors may be used to call the actions to be performed.

[0092] In one configuration, the category of the object is not defined. Additionally, or the model or environment of the object is not assumed. Instead of explicitly detecting and segmenting the object and explicitly estimating the six-degree-of-freedom pose, a dense pixel-level embedding may be generated for more general tasks. The reference embedding from the taught sequence may be used to perform action classification or pose estimation.

[0093] The trained model may be a complete convolutional type. In one configuration, the pixels of the input image are each mapped to a point in the embedding space. The embedding space is given a metric that is implicitly defined by the loss function defined by the output of the model and the training procedure. The trained model may be used for various tasks.

[0094] In one configuration, if an example with one annotation is given, the trained model detects all objects in the semantic class. The objects in the semantic class may be detected by comparing the embeddings in the annotation with the embeddings in other regions. The model may be trained by a discriminative loss function.

[0095] The model may be trained to determine object instances. This model identifies and / or counts independent objects. The model may be trained to predict the vector (2D embedding) of each pixel. The vector may indicate the centroid of the object containing that pixel. At runtime, pixels that point to the same centroid may be grouped as segments of the scene. The execution at runtime may be performed in 3D.

[0096] The model may be trained for 3D correspondences. This model provides, for each pixel, an embedding that is invariant to views and lighting such that the view of any 3D point in the scene is mapped to the same embedding. This model may be trained using a loss function.

[0097] The embedding (and depth data) for each pixel of each RGB-D frame is fused into a dynamic 3D voxel map. Each voxel accumulates the statistics of the position, color, and embedding in the first and second orders. The expiration time of the dynamic object is based on the back-projection of the voxel into the depth image. The voxel map is segmented using standard graph segmentation based on semantic and instance labels, as well as geometric approximation. The voxel map is dimensionally reduced to a 2.5D map with elevation and traversability classification statistics.

[0098] The 2.5D map is used for the motion of the vehicle body without collision, while the voxel map is used for the whole-body motion planning without collision. For the inspection of collisions in 3D, the voxels in the map may be grouped into capsules using a greedy method. The segmented objects may be used for the motion to attach to the hand when the object is grasped.

[0099] The robot may be trained by a one-shot learning approach to recognize features in the scenery (or of a specific manipulation object) that are highly related to the features recorded in the tasks taught in the past. When the task is demonstrated by the user, the features are saved in the form of keyframes throughout the task. The keyframe may be an RGB image containing a multi-dimensional embedding with depth (if available) for each pixel.

[0100] The embedding functions as a feature descriptor that can establish a pixel-by-pixel correspondence relationship at runtime under the assumption that the current image is sufficiently similar to the reference image that existed during teaching. Since depth exists in (almost) all pixels, the correspondence relationship can be used to solve the pose delta between the current image and the reference image. Euclidean constraints may be used to detect inliers, and the Levenberg-Marquardt least squares function is applied together with RANSAC to solve for the 6-degree-of-freedom pose.

[0101] The pose deltas serve as the role of corrections applicable to adapt the sequence of the taught actions to the current landscape. Since the embeddings may be defined for each pixel, the keyframe may be as broad as to include all pixels in the image, or as narrow as to use only the pixels within a mask defined by the user. As discussed, the user may define the mask by selectively annotating regions in the image as being related to a task or being on an object.

[0102] In addition to visual sensing, in one configuration, the robot collects and processes voice input. The voice provides a set of another embedding as an input for teaching the robot. As an example, the robot asks questions and obtains voice input by understanding the spoken language of the responses from humans. The voice responses may be understood using a custom keyword detection module.

[0103] The robot may utilize a fully convolutional keyword spotting model to understand custom wake words, a set of objects (e.g., "mug" or "bottle"), and a set of locations (e.g., "cabinet" or "refrigerator"). In one configuration, the model listens for wake words at a certain interval, e.g., 32ms. Once a wake word is detected, the robot pays attention to whether object or location keywords are detected. During training, artificial noise is added to make the recognition more robust.

[0104] As discussed, to teach a task to a robot, an operator uses a set of VR modes. Each action may have a corresponding VR mode to set and command the parameters specific to that action. Each action mode may include a customized visualization according to the type of parameter to assist in setting each parameter. For example, when setting the parameters of the motion of pulling a door, the hinge axis is labeled and visualized as a line, and the posture candidates for pulling the gripper are restricted to an arc centered on the hinge. To assist the teaching process, several utility VR modes such as motion restoration, annotation of the environment with related objects, repositioning of the virtual robot, camera images, and menus of the VR world are used.

[0105] During execution, the posture of the robot and the parts in the environment may be different from those used during training. Feature matching may be used to discover features in an environment similar to that taught. The posture delta may be established from the correspondence of the matched features. The actions taught by the user may be changed by the calculated posture delta. In one configuration, multiple keyframes are passed to the matching problem. Based on the number of correspondences, the best-matching keyframe is selected.

[0106] FIG. 4 is a diagram showing an example of a hardware implementation of a robot system 400 according to an aspect of the present disclosure. The robot system 400 may be a component of an autonomous or semi-autonomous system such as a vehicle, a robot device 428, or other devices. In the example of FIG. 4, the robot system 400 is a component of the robot device 428. The robot system 400 may be used to control the actions of the robot device 428 by inferring the intentions of the operator of the robot device 428.

[0107] The robot system 400 may be implemented by a bus architecture generally represented as bus 430. The bus 430 may include any number of interconnecting buses and bridges depending on the particular application of the robot system 400 and overall design constraints. The bus 430 connects various circuits such as one or more processors and / or hardware modules represented as processor 420, communication module 422, position module 418, sensor module 402, movement module 426, memory 424, keyframe module 408, and computer-readable medium 414. The bus 430 may also connect various other circuits known to those skilled in the art, such as a timing source, peripherals, a voltage regulator, and a power management circuit, and thus will not be described further herein.

[0108] The robot system 400 includes a transceiver 416 connected to the processor 420, a sensor module 402, a keyframe module 408, a communication module 422, a position module 418, a movement module 426, a memory 424, and a computer-readable medium 414. The transceiver 416 is connected to an antenna 434. The transceiver 416 communicates with various devices via a transmission medium. For example, the transceiver 416 may receive instructions from an operator of the robot device 428 via communication. As discussed herein, the operator may be located at a position remote from the robot device 428. As another example, the transceiver 416 may transmit information regarding keyframe matching from the keyframe module 408 to the operator.

[0109] The robot system 400 includes a processor 420 connected to a computer-readable medium 414. The processor 420 performs processing including the execution of software stored in the computer-readable medium 414 and providing functions according to the present disclosure. When the software is executed by the processor 420, the robot system 400 causes various functions described for specific devices such as the robot device 428 or modules 402, 408, 414, 416, 418, 420, 422, 424, 426 to be executed. The computer-readable medium 414 may also be used to store data that is operated on by the processor 420 when the software is executed.

[0110] The sensor module 402 may be used to obtain measurement values via different sensors such as a first sensor 406 and a second sensor 404. The first sensor 406 may be a visual sensor such as a stereo camera or an RGB camera for taking 2D images. The second sensor 404 may be a ranging sensor such as a LiDAR sensor, a RADAR sensor, or an RGB-D sensor. Of course, aspects of the present disclosure are not limited to the above sensors, and for example, other types of sensors such as temperature, sound waves, and / or lasers may also be considered as either of the sensors 404, 406. The measurement values from the first sensor 406 and the second sensor 404 may be processed by one or more of the processor 420, the sensor module 402, the communication module 422, the position module 418, the movement module 426, and the memory 424 in conjunction with the computer-readable medium 414 to implement the functions described herein. In one configuration, the data captured by the first sensor 406 and the second sensor 404 may be transmitted to the operator as a supplied video via the transceiver 416. The first sensor 406 and the second sensor 404 may be connected to the robot device 428 or may be in communication with the robot device 428.

[0111] The position module 418 may be used to determine the position of the robot device 428. For example, the position module 418 may use the Global Positioning System (GPS) to determine the position of the robot device 428. The communication module 422 may be used to facilitate communication via the transceiver 416. The communication module 422 may provide communication capabilities via different wireless protocols such as WiFi, Long Term Evolution (LTE), 3G, etc. The communication module 422 may also be used to communicate with other components of the robot device 428 that are not modules of the robot system 400.

[0112] The movement module 426 may be used to facilitate the movement of the robot device 428 and / or components of the robot device 428 (such as limbs, hands, etc.). For example, the movement module 426 may control the movement of the limbs 438 and / or the wheels 432. As another example, the movement module 426 may be in communication with a power source of the robot device 428 such as an engine or a battery.

[0113] The robot system 400 also includes a memory 424 for storing data related to the operation of the robot device 428 and the keyframe module 408. The modules may be software modules executed within the processor 420, those resident / stored in the computer-readable medium 414 and / or the memory 424, one or more hardware modules connected to the processor 420, or a combination thereof.

[0114] The keyframe module 408 may communicate with the sensor module 402, transceiver 416, processor 420, communication module 422, position module 418, movement module 426, memory 424, and computer-readable medium 414. In one configuration, the keyframe module 408 receives input (e.g., image data) corresponding to the current vision of the robotic device 428 from one or more sensors 404, 406. The keyframe module 408 may assign pixel descriptors to each pixel, or set of pixels, in the input.

[0115] The keyframe module 408 identifies a keyframe image having a set of first pixels that match a set of second pixels in the input. The keyframe module 408 may match the pixel descriptors of the input to the pixel descriptors of the keyframe image. The keyframe image may be stored in the memory 424, in network storage (not shown), or in another location. The keyframe image may be captured and stored while the robotic device 428 is being trained to perform an action. The pixel descriptors of the keyframe image may be assigned during training.

[0116] The keyframe module 408 also controls the robotic device 428 to perform an action corresponding to the keyframe image. The action may be adapted based on a relative transformation (e.g., pose delta) between the input and the keyframe image. The keyframe module may determine the relative transformation based on changes in distance and / or pose of the input and the keyframe image. The action of the robotic device 428 may be controlled via communication with the movement module 426.

[0117] FIG. 5 shows an example of a graphical sequence 500 of an instructed operation according to an aspect of the present disclosure. As shown in FIG. 5, the graphical sequence 500 includes a start node 502 and an end node 504. The graphical sequence 500 may branch or loop based on sensed visual input, audio input, or other conditions.

[0118] For example, as shown in FIG. 5, after the start node 502, the robot may perform the operation of "listen_for_object". In this example, the robot determines whether it has sensed visual or audio input corresponding to a cup or a bottle. In this example, different operation sequences are executed based on whether the sensed input corresponds to a cup or a bottle. The aspects of the present disclosure are not limited to the operations shown in FIG. 5.

[0119] FIG. 6 shows an example of software modules for a robot system according to an aspect of the present disclosure. The software modules of FIG. 6 may use one or more components of the hardware system of FIG. 4 such as a processor 420, a communication module 422, a position module 418, a sensor module 402, a movement module 426, a memory 424, a keyframe module 408, and a computer-readable medium 414. The aspects of the present disclosure are not limited to the modules shown in FIG. 6.

[0120] As shown in FIG. 6, the robot may receive audio data 604 and / or image data 602. The image data 602 may be an RGB-D image. The audio network 606 may listen for wake words at certain intervals. The audio network 606 receives the raw audio data 604 to detect wake words and extract keywords from the raw audio data 604.

[0121] A neural network such as a dense embedding network 608 receives the image data 602. The image data 602 may be received at certain intervals. The dense embedding network 608 processes the image data 602 and outputs an embedding 610 of the image data 602. The embedding 610 and the image data 602 may be combined to generate a voxel map 626. The embedding 610 may also be input to a keyframe matcher 612.

[0122] The keyframe matcher 612 compares the embedding 610 with a plurality of keyframes. When the embedding 610 corresponds to the embedding of a keyframe, the matching keyframe is identified. The embedding 610 may include a pixel descriptor, depth information, and other information.

[0123] The task module 614 may receive one or more task graphs 616. The task module 614 provides a response to a request from the keyframe matcher 612. The keyframe matcher 612 matches the task to the matching keyframe. The task may be determined from the task graph 616.

[0124] The task module 614 may also send an operation request to the operation module 618. The operation module 618 provides an operation status to the task module 614. In addition, the operation module 618 may request information about the matching keyframe and the corresponding task from the keyframe matcher 612. The keyframe matcher 612 provides information about the matching keyframe and the corresponding task to the operation module 618. The operation module 618 may receive voxels from the voxel map 626.

[0125] In one configuration, the operation module 618 receives an operation plan from the operation planner 620 in response to an operation plan request. The operation module 618 also receives the part status from the part control module 622. The operation module 618 sends a part command to the part control module 622 in response to receiving the part status. Finally, the part control module 622 receives the joint status from the joint control module 624. The part control module 622 sends a joint command to the joint control module 624 in response to receiving the joint status.

[0126] FIG. 7 shows a method 700 for controlling a robot device according to an aspect of the present disclosure. As shown in FIG. 7, optionally, at block 702, when the robot device is trained to perform an action, it captures one or more keyframe images. The task may include interactions such as opening a door or traveling like traveling in the environment.

[0127] After training, at block 704, the robot device captures an image corresponding to the current vision. The image may be captured by one or more sensors such as an RGB-D sensor. A pixel descriptor may be assigned to each pixel of the first set of pixels of the keyframe image and the second set of pixels of the image. The pixel descriptor provides pixel-level information in addition to depth information. The area corresponding to the first set of pixels may be selected by the user.

[0128] At block 706, the robot device also identifies a keyframe image having a first set of pixels that matches the second set of pixels of the image. The robot device may compare the image with a plurality of keyframes. The robot device may determine whether the first set of pixels matches the second set of pixels when the respective pixel descriptors match.

[0129] Optionally, at block 708, the robot device determines at least one difference in distance, pose, or a combination thereof between the keyframe image and the image. The difference may be referred to as a pose delta. At block 710, the robot device performs a task associated with the keyframe image. Optionally, at block 712, the task is adjusted based on the determined differences between the keyframe image and the image.

[0130] The various operations of the method described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. When there are operations shown in the figures, these operations may have corresponding functional components assigned generally similar numbers.

[0131] As used herein, "determining" includes a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, searching (e.g., searching within a table, database, or other structure), probing, etc. Additionally, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Further, "determining" may include resolving, selecting, choosing, establishing, etc.

[0132] As used herein, the phrase "at least one of" refers to any combination of items from a list of items, including a single item. For example, "at least one of a, b, or c" is intended to include a, b, c, a-b, a-c, b-c, a-b-c.

[0133] The various illustrative logical blocks, modules, and circuits described in connection with the present disclosure may be implemented or executed by a processor configured according to the present disclosure, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a Field Programmable Gate Array signal (FPGA) or other programmable logic device (PLD), discrete gates or transistor logic, discrete hardware components, or any combination of the foregoing designed to perform the functions described herein. The processor may be a microprocessor, a controller, a microcontroller, or a state machine configured as described in the present specification. The processor may also be implemented as a combination of computing devices, such as, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with a DSP core, or other special configurations described herein.

[0134] The steps or algorithms of the methods described in connection with the present disclosure may be embodied directly in hardware, in software modules executed by a processor, or in a combination of the two. The software modules may be present in a storage device, or a machine-readable medium, including any medium that can be used to carry or store the desired program code in the form of instructions or data structures and that is accessible by a computer, such as a random access memory (RAM), a read-only memory (ROM), a flash memory, an erasable programmable read only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a removable disk, a CD-ROM, or other optical disk storage device, a magnetic disk storage device, or other magnetic storage device, or any other medium. The software modules may comprise a single instruction, or may comprise a number of instructions, and may be distributed over several different code segments, among different programs, and across several storage media. The storage medium may be connected to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor.

[0135] The methods disclosed herein include one or more steps or actions for realizing the disclosed methods. The steps and / or actions of the methods may be interchanged with one another without departing from the scope of the claims. In other words, the order and / or use of specific steps and / or actions may be changed without departing from the scope of the claims, unless a specific order of the steps or actions is specified.

[0136] The described functionality may be implemented by hardware, software, firmware, or any combination thereof. When implemented in hardware, an example of a hardware configuration may include a processing system in the apparatus. The processing system may be implemented using a bus architecture. The bus may include any number of interconnecting buses and bridges depending on the specific application of the processing system and overall design constraints. The bus may connect various circuits including a processor, a machine-readable medium, and a bus interface. The bus interface may be used to connect, among other things, a network adapter to the processing system via the bus. The network adapter may be used to implement signal processing functions. In certain aspects, a user interface (e.g., keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also connect various other circuits known to those of ordinary skill in the art, such as a timing source, peripherals, voltage control, power management circuits, etc., and thus will not be described further.

[0137] The processor may be responsible for processing including management of the bus and execution of software stored in the machine-readable medium. Software shall be construed to mean instructions, data, or any combination thereof, regardless of the terminology used such as software, firmware, middleware, microcode, hardware description language, or other.

[0138] In a hardware implementation, the machine-readable medium may be part of a processing system separate from the processor. However, as will be readily understood by those skilled in the art, the machine-readable medium, or any portion thereof, may be external to the processing system. For example, the machine-readable medium may include a communication line, a carrier wave modulated by data, and / or a computer product detached from the device, all of which may be accessed by the processor via a bus interface. Alternatively, or in addition, the machine-readable medium, or a portion thereof, may be integrated into the processor, such as when a cache and / or a special register file may exist. Although the various components discussed have been described as having a particular location, such as local components, they may be configured in various ways, such as specific components configured as part of a distributed computing system.

[0139] The processing system may be composed of one or more microprocessors providing processor functionality, and at least a portion of the machine-readable medium and an external memory, all of which may be connected through a support circuit by an external bus architecture. Alternatively, the processing system may include one or more neuromorphic processors for implementing the neuron models and neural system models described herein. As another alternative, the processing system may be implemented by an application-specific integrated circuit (ASIC) having a processor, a bus interface, a user interface, a support circuit, and at least a portion of the machine-readable medium integrated into a single chip, or one or more Field Programmable Gate Arrays (FPGAs), Programmable Logic Devices (PLDs), controllers, state machines, gate logic, individual hardware components, or other suitable circuits, or any combination of circuits capable of performing the various functions described within this disclosure. Those skilled in the art will recognize how best to implement the functions of the described processing system based on the specific application and the overall design constraints imposed on the system as a whole.

[0140] The machine-readable medium may comprise a number of software modules. The software modules may include a transmission module and a reception module. Each software module may reside within a single storage device or may be distributed across a plurality of storage devices. For example, a software module may be loaded from a hard drive to RAM when a triggering event occurs. During execution of the software module, the processor may load some instructions into the cache to increase the access speed. One or more cache lines may then be loaded into a special-purpose register file for execution by the processor. By referring to the following functions of the software module, it will be understood that the functions are performed by the processor when the software module executes instructions. Further, it should be understood that aspects of the present disclosure improve the functions of a processor, computer, machine, or other system implementing such aspects.

[0141] When implemented in software, the functions may be stored or transmitted on a computer-readable medium as one or more instructions or code. The computer-readable medium includes both a computer storage device and a communication medium including any storage device that facilitates transfer of a computer program from one place to another. Additionally, any connection is appropriately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared (IR), radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as IR, radio, and microwave are included in the definition of the medium. As used herein, disk and disc include compact disc (CD), laser disc (registered trademark), optical disc, digital versatile disc (DVD), floppy (registered trademark) disc, Blu-ray (registered trademark) disc, where disk typically magnetically reproduces data and disc optically reproduces data using a laser. Thus, in some aspects, the computer-readable medium may comprise a non-transitory computer-readable medium (e.g., a tangible medium). Additionally, in other aspects, the computer-readable medium may comprise a transitory computer-readable medium (e.g., a signal). The above combinations should be included within the scope of the computer-readable medium.

[0142] Accordingly, certain aspects may comprise a computer program product for performing the operations presented herein. For example, such a computer program product may comprise a computer-readable medium storing (and / or encrypting) instructions executable by one or more processors for performing the operations described herein. In certain aspects, the computer program product may include packaging materials.

[0143] Furthermore, it should be understood that the modules and / or other suitable means for implementing the methods and techniques described herein may be downloadable and / or retrievable by the user terminal and / or the base station as necessary. For example, such devices can be connected to a server to facilitate the transfer of means for implementing the methods described herein. Alternatively, the various methods described herein can be provided via a storage means in such a manner that the user terminal and / or the base station can obtain the various methods by connecting a storage means to the device or providing a storage means to the device. Furthermore, any other technique for providing the methods and techniques described herein to the device can be used.

[0144] It should be understood that the claims are not limited to the exact configurations and components shown above. Various modifications, changes, and variations can be made to the arrangements, operations, and details of the methods and devices described above without departing from the scope of the claims. The invention disclosed in this specification includes the following aspects. 〔Aspect 1〕 Taking an image corresponding to the current vision of a robot device, Identifying a keyframe image having a set of first pixels that matches a set of second pixels of the image, Executing a task corresponding to the keyframe image by the robot device, A method for controlling a robot device, including. 〔Aspect 2〕 The method according to Aspect 1, further including taking the keyframe image while the robot device is trained to execute the task. 〔Aspect 3〕 The method according to Aspect 1, further including assigning a pixel descriptor to each pixel in the set of first pixels and the set of second pixels. 〔Aspect 4〕 The method according to Aspect 3, wherein the pixel descriptor has a set of values corresponding to pixel-level information and depth information. 〔Aspect 5〕 Judging one of the differences among distance, posture, or a combination between the keyframe image and the image, Adjusting the task based on the judged difference, The method according to Aspect 1, further including. 〔Aspect 6〕 The method according to Aspect 1, wherein the task has at least one of interacting with an object, traveling in an environment, or a combination thereof. 〔Aspect 7〕 The method according to Aspect 1, wherein the region corresponding to the set of first pixels is selected by a user. 〔Aspect 8〕 A memory, At least one processor, and the at least one processor is Taking an image corresponding to the current vision of a robot device, Identifying a keyframe image having a set of first pixels that matches a set of second pixels of the image, Executing a task corresponding to the keyframe image, A robot device configured as such. 〔Aspect 9〕 The robot device according to Aspect 8, wherein the at least one processor is further configured to take the keyframe image while the robot device is trained to execute the task. 〔Aspect 10〕 The robot device according to Aspect 8, wherein the at least one processor is further configured to assign a pixel descriptor to each pixel in the set of first pixels and the set of second pixels. [Aspect 11] The robot device according to aspect 10, wherein the pixel descriptor has a set of values corresponding to pixel level information and depth information. [Aspect 12] The at least one processor further determines one difference among distance, orientation, or a combination between the keyframe image and the image, and adjusts the task based on the determined difference, The robot device according to aspect 8, which is configured as described above. [Aspect 13] The task of the robot device according to aspect 8 includes at least one of interacting with an object, traveling in an environment, or a combination thereof. [Aspect 14] The area corresponding to the first set of pixels of the robot device according to aspect 8 is selected by a user. [Aspect 15] A non-transitory computer-readable medium storing program code for controlling a robot device, wherein the program code includes program code for capturing an image corresponding to the current vision of the robot device, program code for identifying a keyframe image having a first set of pixels that matches a second set of pixels of the image, and program code for executing a task corresponding to the keyframe image by the robot device, The non-transitory computer-readable medium includes the above. [Aspect 16] The program code of the non-transitory computer-readable medium according to aspect 15 further includes program code for capturing the keyframe image while the robot device is trained to execute the task. [Aspect 17] The program code of the non-transitory computer-readable medium according to aspect 15 further includes program code for assigning a pixel descriptor to each pixel in the first set of pixels and the second set of pixels. [Aspect 18] The pixel descriptor of the non-transitory computer-readable medium according to aspect 17 has a set of values corresponding to pixel level information and depth information. [Aspect 19] The program code further includes program code for determining one difference among distance, orientation, or a combination between the keyframe image and the image, and program code for adjusting the task based on the determined difference, The non-transitory computer-readable medium according to aspect 15 includes the above. [Aspect 20] The non-transitory computer-readable medium according to aspect 15, wherein the task has at least one of interacting with an object, traveling in an environment, or a combination thereof.

Claims

1. capturing an image corresponding to the current vision of the robot device; identifying the keyframe image based on the RGB (red-green-blue) values of each pixel in the first set of pixels of the keyframe image corresponding to the current vision of the robot device matching the RGB values of the corresponding pixels in the second set of pixels of the image, where the keyframe image was captured during a previous training period and the keyframe image is stored in a memory associated with the robot device; determining one or both of a first difference between a first pose of the robot device with respect to a first object in the keyframe image and a second pose of the robot device with respect to a second object in the image that corresponds to the first object, or a second difference between a first distance of the robot device with respect to the first object and a second distance of the robot device with respect to the second object; adjusting one or both of a speed or a position associated with the task and related to the keyframe image, based on one or both of the first difference or the second difference; performing the task based on adjusting one or both of the speed or the position of the end effector of the robot device via the end effector of the robot device; A method for controlling a robot device, comprising.

2. The method according to claim 1, wherein the robot device is trained to perform the task during the training period.

3. The method according to claim 1, wherein each pixel of the first set of pixels and the second set of pixels is associated with a pixel descriptor.

4. Each pixel descriptor has a set of values corresponding to pixel-level information and depth information, and the pixel-level information has the RGB values of the pixel associated with the pixel descriptor, the method according to claim 3.

5. The task has at least one of interacting with an object, traveling in an environment, or a combination thereof, the method according to claim 1.

6. The region corresponding to the set of the first pixels is selected by a user, the method according to claim 1.

7. A memory, At least one processor, and the at least one processor Takes a picture of an image corresponding to the current vision of the robot device, Identifies the keyframe image based on that the RGB (red-green-blue) value of each pixel of the set of the first pixels of the keyframe image corresponding to the current vision of the robot device matches the RGB value of the corresponding pixel of the set of the second pixels of the image, the keyframe image was taken during a previous training period, the keyframe image is stored in a memory associated with the robot device, Determines a first difference between a first pose of the robot device with respect to a first object of the keyframe image and a second pose of the robot device with respect to a second object of the image that is the second object corresponding to the first object, or a second difference between a first distance of the robot device with respect to the first object and a second distance of the robot device with respect to the second object, either or both of the first difference or the second difference, Adjusts either or both of the speed or position related to a task performed by the end effector of the robot device and associated with the keyframe image based on determining either or both of the first difference or the second difference, Executing the task based on adjusting one or both of the speed or position of the end effector via the end effector of the robot device A robot device configured as described above.

8. The robot device according to claim 7, wherein the robot device is trained to execute the task during the training period.

9. The robot device according to claim 7, wherein each pixel of the set of the first pixels and the set of the second pixels is associated with a pixel descriptor.

10. The robot device according to claim 9, wherein each pixel descriptor has a set of values corresponding to pixel-level information and depth information, and the pixel-level information has RGB values of the pixels associated with the pixel descriptor.

11. The robot device according to claim 7, wherein the task has at least one of interacting with an object, traveling in an environment, or a combination thereof.

12. The robot device according to claim 7, wherein the region corresponding to the set of the first pixels is selected by a user. For a computer for controlling a robot device, Taking an image corresponding to the current vision of the robot device, Identifying the keyframe image based on the RGB (red-green-blue) values of each pixel of the set of the first pixels of the keyframe image corresponding to the current vision of the robot device matching the RGB values of the corresponding pixels of the set of the second pixels of the image, wherein the keyframe image was taken during a previous training period and the keyframe image is stored in a memory associated with the robot device, The first difference between the first posture of the robot device with respect to the first object in the key frame image and the second posture of the robot device with respect to the second object that is the second object in the image and corresponds to the first object, or, determining one or both of the second differences between the first distance of the robot device with respect to the first object and the second distance of the robot device with respect to the second object, Based on determining one or both of the first difference or the second difference, adjusting one or both of the speed or position related to the task associated with the key frame image and related to the task, which is performed by the end effector of the robot device, Executing the task based on adjusting one or both of the speed or position of the end effector via the end effector of the robot device, A non-transitory computer-readable medium having program code recorded thereon for causing the above to be executed.

14. The non-transitory computer-readable medium according to claim 13, wherein the robot device is trained to execute the task during the training period.

15. Each pixel of the set of the first pixels and the set of the second pixels is associated with a pixel descriptor, the non-transitory computer-readable medium according to claim 13.

16. Each pixel descriptor has a set of values corresponding to pixel-level information and depth information, and the pixel-level information has RGB values of the pixels associated with the pixel descriptor, the non-transitory computer-readable medium according to claim 15.

17. The task has at least one of interacting with an object, traveling in an environment, or a combination thereof, the non-transitory computer-readable medium according to claim 13.

18. The method according to claim 1, wherein the robot device is trained to execute the task based on a human demonstration.

19. The robot device according to claim 7, wherein the robot device is trained to execute the task based on a human demonstration.

20. The non-transitory computer-readable medium according to claim 13, wherein the robot device is trained to execute the task based on a human demonstration.

Citation Information

Patent Citations

  • Data describing method and data processor

    JP2000287166A

  • Method of representing image group, descriptor of image group, searching method of image group, computer-readable storage medium, and computer system

    JP2010225172A

  • Robot device, control method of the robot device, computer program, and robot system

    JP2013022705A

  • Work robot

    WO2007138987A1