Visual Instruction and Repetition of a Mobile Operation System

The method allows a robotic device to adapt to changes in the environment by mapping task images to teaching images and updating parameterized motions, enhancing its ability to perform complex tasks with high accuracy and adaptability.

JP7693649B2Active Publication Date: 2025-06-17TOYOTA JIDOSHA KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022503980
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-13
Filing Date
2020-07-22
Publication Date
2025-06-17
Estimated Expiration
2040-07-22

AI Technical Summary

Technical Problem

Conventional robotic systems struggle to perform tasks when the starting point or direction of objects in the environment differs from the programmed or taught task, limiting their adaptability and effectiveness in diverse and dynamic environments.

Method used

A method for controlling a robotic device that involves placing it in a task environment, mapping a task image to a teaching image, defining a relative transformation between the two images, and updating parameterized motions to execute the task, allowing the robot to adapt to changes in the environment and perform tasks from different start positions or orientations.

Benefits of technology

This approach enables the robotic device to execute tasks with high accuracy and adaptability, even when the start position or orientation of objects changes, thereby improving its ability to perform complex tasks in diverse environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693649000001
    Figure 0007693649000001
  • Figure 0007693649000002
    Figure 0007693649000002
  • Figure 0007693649000003
    Figure 0007693649000003
Patent Text Reader

Abstract

A method for controlling a robotic device is presented. The method includes placing the robotic device in a task environment. The method also includes mapping a descriptor of a task image of a scene in the task environment to a teach image of a teach environment. The method further includes defining a relative transformation between the task image and the teach image based on the mapping. The method further includes updating parameters of a parameterized set of actions based on the relative transformation to perform a task corresponding to the teach image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the benefit of U.S. Provisional Application No. 62 / 877,792, filed on July 23, 2019, entitled "Key Frame Matcher", U.S. Provisional Application No. 62 / 877,791, filed on July 23, 2019, entitled "Visual Instruction and Repetition for Operations - Instruction VR", and U.S. Provisional Application No. 62 / 877,793, filed on July 23, 2019, entitled "Visualization", and claims the benefit of U.S. Patent Application No. 16 / 570,852, filed on September 13, 2019, entitled "Visual Instruction and Repetition of a Movement Operation System", the content of which is incorporated herein by reference.

[0002] Certain aspects of the present disclosure generally relate to robotic devices, and more particularly to systems and methods for teaching a robotic device, via virtual reality (VR), actions parameterized as repeatable operations.

Background Art

[0003] The tasks that people perform in a home or other environment are diverse. As robotic assistive technologies develop, robots are programmed to perform the diverse tasks that people perform in an environment such as a home. This makes it difficult to develop cost-effective, application-specific solutions. Further, the environment, objects, and tasks are not very consistent and are diverse. While some objects and tasks are similar, a robot may also encounter numerous unique objects and tasks.

[0004] Currently, a robot can be programmed and / or taught to perform a task. In conventional systems, a task is specific to a direction and a starting point. If the starting point and / or the direction / position of an object do not match the programmed or taught task, it is desirable to improve the robotic assistive system to perform the same task.

Summary of the Invention

[0005] In one aspect of the present disclosure, a method for controlling a robotic device is disclosed. The method includes placing the robotic device in a task environment. The method also includes mapping a descriptor of a task image of a scene in the task environment to a teaching image in a teaching environment. The method further includes defining a relative transformation between the task image and the teaching image based on the mapping. The method further includes updating a set of parameterized motions based on the relative transformation to execute a task corresponding to the teaching image.

[0006] In another aspect of the present disclosure, a non-transitory computer-readable medium storing non-transitory program code is disclosed. The program code is for controlling a robotic device. The program code, when executed by a processor, includes program code for placing the robotic device within a task environment. The program code also includes program code for mapping a descriptor of a task image of a scene in the task environment to a teaching image in a teaching environment. The program code further includes program code for defining a relative transformation between the task image and the teaching image based on the mapping. The program code further includes program code for updating parameters of a set of parameterized motions based on the relative transformation to execute a task corresponding to the teaching image.

[0007] Another aspect of the present disclosure relates to an apparatus for controlling a robotic device. The apparatus has a memory and one or more processors connected to the memory. The processor places the robotic device within a task environment. The processor also maps a descriptor of a task image of a scene in the task environment to a teaching image in a teaching environment. The processor further defines a relative transformation between the task image and the teaching image based on the mapping. The processor further updates parameters of a set of parameterized motions based on the relative transformation to execute a task corresponding to the teaching image.

[0008] The features and technical advantages of the present disclosure have been broadly outlined above so that the detailed description that follows may be better understood. Additional features and advantages of the present disclosure will be described hereinafter. It should be understood by those skilled in the art that the present disclosure can be readily used as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. It should also be recognized by those skilled in the art that such equivalent constructions do not depart from the teachings of the present disclosure as defined by the appended claims. The novel features believed to be characteristic of the present disclosure will be better understood from the following description when considered in conjunction with the accompanying drawings, along with its additional objects and advantages. However, it should be clearly understood that each drawing is provided for the purpose of illustration and description only and is not intended to limit the scope of the present disclosure.

Brief Description of the Drawings

[0009] The functions, nature, and advantages of the present disclosure will become more apparent from the detailed description that follows when considered in conjunction with the corresponding drawings in which like reference characters refer to the whole.

[0010]

Figure 1

Figure 2A

Figure 2B

Figure 3A

Figure 3B

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

DETAILED DESCRIPTION OF THE INVENTION

[0011] The following detailed description related to the accompanying drawings is intended to explain various configurations and is not intended to present a single configuration for implementing the concepts described herein. The detailed description includes specific details for the purpose of providing a complete understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts can be implemented without these specific details. In some cases, well-known structures and components are shown in block diagrams to avoid obscuring such concepts.

[0012] Based on the teachings, it should be understood by those skilled in the art that the scope of the present disclosure is intended to include any aspect of the present disclosure, whether implemented independently or in combination with other aspects of the present disclosure. For example, any number of aspects disclosed may be used to implement an apparatus or to carry out a method. In addition, the scope of the present disclosure is intended to include such an apparatus or method implemented using other structures and functions, or structures and functions, in addition to the various aspects disclosed in the present disclosure. It should be understood that any aspect of the present disclosure may be embodied by one or more elements of the claims.

[0013] As used herein, the term "exemplary" is used in the sense of "serving as an example, instance, or illustration". Any aspect of the present disclosure described as "exemplary" should not necessarily be understood as being preferred or advantageous compared to other aspects.

[0014] Although specific embodiments are described in this specification, numerous variations and substitutions to these embodiments are included within the scope of the present disclosure. While some benefits and advantages of the preferred embodiments are described, the scope of the present disclosure is not intended to be limited to specific benefits, uses, or purposes. Rather, the embodiments of the present disclosure are intended to be widely applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated by way of example in the description of the figures and the preferred embodiments. The detailed description and the drawings are for the purpose of describing the present disclosure only rather than limiting it, and the scope of the present disclosure is defined by the appended claims and their equivalents.

[0015] The world population is aging, and the proportion of people over 65 will increase substantially compared to those under 65 in the next decade. Robot assistance systems can help the elderly live longer and healthier lives. Naturally, robot assistance systems are not limited to assisting the elderly. Robot assistance systems can be assistance in various environments and / or for people of all ages.

[0016] Currently, robots can be programmed and / or taught to perform tasks. Tasks are specific to directions and starting points. Generally, a robot cannot perform the same task if the starting point and / or the direction / position of an object do not match the programmed or taught task. In the present disclosure, a robot within a robot assistance system may also be referred to as a robot device.

[0017] The robot is physically capable of movement operations. The operability of the robot is the ability to change the position of the end effector as a function of the joint configuration. In one configuration, the robot further has the function of automatically controlling the whole body and making a plan. As a result, a human operator can demonstrate seamless end effector motion in the task space in virtual reality (VR) with little or no concern for kinematic constraints or the robot's posture. The robot is equipped with one or more field-of-view RGB-D (red-green-blue and depth) sensors on a pan / tilt head, which can provide important context for a human operator in virtual reality to perform tasks. An RGB-D image is a combination of an RGB image and a corresponding depth image. A depth image is an image channel in which each pixel is associated with the distance between the image plane and the corresponding object on the RGB image.

[0018] Aspects of the present disclosure relate to a mobile operation hardware and software system (such as a robotic device) capable of autonomously performing human-level complex tasks in different environments after being taught a task by demonstration from a human operator within virtual reality. For example, a human operator can teach a robotic device to operate within an environment by operating the robotic device through a virtual reality platform within the environment.

[0019] In one aspect, the robotic device is located in an environment, and image data of the scenery is collected. Then the robotic device is controlled to perform a task (e.g., through a virtual reality interface). By restricting the field of view of a human operator in virtual reality to the field of view of the robotic device, it is ensured that the robotic device has sufficient information to perform the task alone during training.

[0020] A method of teaching an action / task to a robotic device may include parameterizing an operation performed by an operator through a virtual reality interface. For example, the virtual reality interface may include the use of paddles, hand-held controllers, paintbrush tools, wiping tools, and / or placement tools that the operator operates while wearing a headset that renders the VR environment. Thus, a human operator teaches a set of parameterized primitives (or operations) rather than directly teaching an operation in the task space. The parameterized primitives combine a collision-free motion plan and hybrid (position and force) Cartesian control to reduce the parameters to be taught and provide robustness during execution.

[0021] A parameterized operation is a task that is learned by dividing the task into a small number of separate chunks of operations. Each operation is defined by a set of parameters such as joint angle changes, rotation angles, or the like. The values of these parameters may be configured and updated based on the situation of the robot when performing the task. Parameterized operations may be learned and extracted from one learned task and combined with other tasks to form a larger task. A parameterized operation such as opening a door with a rotating handle may be implemented as performing the opening of any door handle (e.g., a door that requires a 30-degree rotation, or a door that requires a 60-degree rotation, or more rotations). For example, the rotation angle may be one parameter that defines a parameterized operation for opening a door with a rotating door handle.

[0022] The executed task is defined as a set of parameterized operations and is associated with the image data of the scene where the action was performed. The parameterized operations are linked to the scene using a robust learned dense visual keypoint embedding and a virtual reality-based masking of the relevant parts of the scene. In one configuration, the pixels of the test image are compared to the pixels of a reference image. The reference image may be referred to as a keyframe. The keyframe may be acquired during training. The use of the keyframe provides invariance to pose and image transformation. The vision system determines that the test image matches the keyframe when the number of matching pixels is greater than a threshold. In one configuration, the vision system compares pixel descriptors to identify the matching pixels between the test image and the keyframe. The pixel descriptor includes pixel-level information and depth information. The pixel-level information includes information such as the RGB values of the pixel and the context of the pixel within the image / surrounding pixels.

[0023] To execute the task, the robotic device may be located in the same or a similar environment (relative to its initial position during training). The robotic device may be located at different start positions and may optionally assume relatively different initial poses (e.g., adjusted to start positions with different joint angles). The robotic device may be tasked with performing the same task (e.g., a set of parameterized operations), such as picking up a bottle, opening a cabinet, and placing the bottle in the cabinet, without control by a human operator. For example, the robotic device may execute the same task by updating the parameters of the actions taught in a sequence controlled in virtual reality. The parameters may be updated based on the current pose and / or position of the robotic device compared to the pose and / or position used during training.

[0024] To update the parameters, the robot device takes an initial image of the scene and maps pixels and / or dense neural network descriptors from the new image to an in-training image called a keyframe. A keyframe is a snapshot of an image that has the depth information the robot device sees. Through mapping, a relative transformation between the new image and the in-training image (e.g., the keyframe) is defined.

[0025] The keyframe can be mapped to the new image by the relative transformation. The mapping may be performed by matching pixels and / or dense neural network descriptors of different images. The relative transformation may be defined by changes in the position of the robot on the x-axis, y-axis, z-axis, roll, pitch, and yaw. The relative transformation may be used to update the parameters of a parameterized motion from the taught parameters to the observed situation.

[0026] The relative transformation may be applied to a parameterized motion. By applying the relative transformation to a parameterized motion, the robot device can execute the same task as previously taught even if the start position and / or orientation changes. The robot system may continuously map pixels and / or dense neural network descriptors from the current scene to those from the keyframe to the parameterized motion so that continuous adjustment is made. For example, the relative transformation may be applied to a taught action defined by a set of parameterized motions such as pulling out and opening a drawer, opening a door, picking up a cup or bottle, or the like.

[0027] In some embodiments, the actions may be related to the overall scene and / or object specific. For example, for the action of picking up a bottle, it may be necessary to use keyframes related to the overall scene to travel to the location of the bottle, and once approaching the bottle, keyframes specific to the bottle may be analyzed independently of the environment. The traveling action is used to move the robot from one location to another. This enables the robot to identify the position of an object that can be located anywhere in the environment during the training of the "pick up" action and be able to perform tasks such as "pick up" regardless of the position of the bottle. The manipulating action can be used to move the parts of the robot (e.g., the torso and / or arms) to contact the desired object.

[0028] FIG. 1 shows an example of an operator 100 controlling a robot device 106 using a virtual reality platform during training according to an embodiment of the present disclosure. As shown in FIG. 1, the operator 100 includes a vision system 102 and a motion controller 104 (e.g., a gesture following system) for controlling the robot device 106. The vision system 102 may not only capture the vision of the operator 100 but also provide a feed video. The operator 100 may be located at a position remote from the position of the robot device 106. In this example, the robot device 106 is located in a kitchen 108, and the operator 100 is located at a location different from the kitchen 108, such as a robot control center 114.

[0029] The vision system 102 may provide a feed video of the position of the robot device 106. For example, the vision system 102 may provide a view of the kitchen 108 based on the front view point of the robot device 106. Other viewpoints such as a 360-degree (360°) scene may be provided. The viewpoint is provided using one or more vision sensors such as a video camera of the robot device 106. The vision system 102 is not limited to a headset as shown in FIG. 1. The vision system 102 may also be a monitor 110, an image projector, or any other device capable of displaying the feed video from the robot device 106.

[0030] One or more actions of the robot device 106 may be controlled via the motion controller 104. For example, the motion controller 104 captures the gestures of the operator 100, and the robot device 106 mimics the captured gestures. The operator 100 may control the movement, limb movements, and other actions of the robot device 106 via the motion controller 104. For example, the operator 100 may grasp the bottle 116 on the table 120, open the cabinet 118, and control the robot device 106 to place the bottle 116 into the cabinet 118. In this case, the bottle 116 is in an upright position or posture. The actions performed by the operator 100 are parameterized through the virtual reality interface. Each of the actions is defined by a set of parameters such as joint angle changes, rotation angles, or the like. The values of these parameters may be configured and updated based on the situation of the robot when performing the task. The parameterized actions may be learned and extracted from one learned task and combined with other tasks to form larger tasks.

[0031] In one aspect, in the virtual reality interface for training, a user wearing a headset and holding a controller may interactively control the operation of the robot. The environment drawn within the virtual reality headset is a virtual reality environment of the actual environment as seen from the robot device 106, such as the field of view shown in FIG. 1. As another field of view, it may include the operator's field of view as shown in FIGS. 2A and 2B below. In some aspects, the user interface provides a paintbrush tool to annotate or highlight an object to be operated on, such as a bottle on the kitchen countertop. For example, through the voxel map of the drawn environment that may be generated by the virtual reality generator, the operator / user can paint the segment of the voxel map occupied by the object to be interacted with. Other user tools include an eraser tool or a placement tool, and the operator can draw a box at the location where the action in the voxel map is to be performed.

[0032] Aspects of the present disclosure are not limited to capturing the gestures of the operator 100 via the motion controller 104. Other types of gesture capture systems are also conceivable. The operator 100 may control the robot device 106 via the wireless connection 112. Additionally, the robot device 106 may provide feedback such as a feed video to the operator 100 via the wireless connection 112.

[0033] FIG. 2A shows an example of an operator (not shown) controlling the robot device 200 in the dining environment 202 according to an aspect of the present disclosure. For clarity, FIG. 2A is a top view of the dining environment 202. As shown in FIG. 2A, the dining environment 202 includes a dining table 204, a sink 206, a drawer 208 containing a spoon 218, and a counter 210. The operator is located at a position away from the dining environment 202.

[0034] In the example of FIG. 2A, the robot device 200 was controlled to place a plate 212, a knife 214, and a fork 216 on the dining table 204. After placing the plate 212, the knife 214, and the fork 216 on the dining table 204, the operator may make a gesture towards the spoon 218. The gesture may include one or more of the following actions: an action 220 of the limb 222 towards the spoon 218, directing the field of view 224 (e.g., line of sight) towards the spoon 218, moving the robot device 200 towards the spoon, and / or other actions.

[0035] FIG. 2B shows an example of a display 250 provided to an operator according to an aspect of the present disclosure. The display 250 may be a visual system such as a headset, a monitor, or other types of displays. As shown in FIG. 2B, the display 250 includes a supply video 252 provided from a visual sensor of the robot device 200. For example, based on the field of view 224 of the robot device 200, the supply video 252 displays the sink 206, the counter 210, the drawer 208, and the spoon 218. In one configuration, a point cloud representation (not shown) may be overlaid and displayed on the supply video 252. The operator may guide the robot device 200 in an environment such as the dining environment 202 based on the supply video 252.

[0036] The display 250 may include an on-screen instruction area 254 for providing a notification to the operator. As shown in FIG. 2B, the on-screen instruction area 254 is separate from the supply video 252. Alternatively, the on-screen instruction area 254 may overlap the supply video 252.

[0037] In one configuration, the robot device 200 associated with the robot control system identifies a landscape of a task image in the vicinity of the robot device 200. For example, the robot device 200 identifies potential targets within the field of view 224 of the robot device 200. In this example, the sink 206, the counter 210, the drawer 208, and the spoon 218 are identified as potential targets. For example, the robot device may be tasked with performing a table setting task that includes carrying the spoon 218 from the drawer 208 to the dining table 204. Accordingly, the spoon 218 in the drawer 208 is considered a potential target for the task.

[0038] For example, the robot device 200 may be taught to grasp one or more spoons 218 and perform an action (e.g., place the spoon 218 on the table 204). Accordingly, the robot device 200 may open the hand attached to the limb (arm or leg) 222 to prepare to grasp one or more spoons 218. As another example, the robot device 200 may adjust a gesture or the current operation to improve the action. The robot device 200 may adjust the angle at which the limb 222 approaches to improve the operation of grasping the spoon 218. The adjustment of the gesture, operation, and / or limb is parameterized and stored for a specific task associated with a specific scenario.

[0039] FIG. 3A shows an example of a robot device 306 that performs a task from a different start position in the same environment as compared to the environment and start position shown in FIG. 1. The robot device 306 autonomously performs a human-level complex task in an actual home after being taught a task by a human operator in virtual reality by demonstration (e.g., one demonstration). For example, as shown in FIGS. 1 and 2, a human operator can teach the robot device 306 to operate in a teaching environment. The teaching environment corresponds to the training kitchen 108, and the task environment corresponds to the task kitchen 308A.

[0040] The training kitchen 108 in FIG. 1 is the same as the kitchen 308A in FIG. 3A, although the kitchen 308A need not be the same as the kitchen 108. For example, the task kitchen 308A may include a different refrigerator, a different table (e.g., table 320), a different oven, different cabinets, etc. located in a similar location as the training kitchen 108. For example, the parameters of the actions taught during a virtual reality controlled sequence are updated when the robotic device 306 is tasked (without control by a human operator) to perform tasks such as lifting a bottle (e.g., bottle 116 in FIG. 1 or bottle 316 in FIG. 3A), opening a cabinet (e.g., cabinet 118 in FIG. 1 or cabinet 318 in FIG. 3A), and placing the bottle inside the cabinet.

[0041] Since the initial position of the robotic device 306 is different from the initial position of the robotic device 106, the robotic device 306 is designated to update the set of parameterized actions. Additionally, due to the difference in the initial positions of the robotic device 306, the start image associated with the new task may be different from that of the task the robotic device was trained on.

[0042] To update the parameters, the robotic device 306 captures (e.g., using a visual or high-resolution camera) a new task image of the scenery from the initial position of the robotic device 306 within the task environment. In one aspect, the initial position of the robotic device 306 deviates from the start situation or position where the robotic device was taught to perform tasks using a virtual reality (VR) interface in a teaching environment (e.g., FIG. 1). For example, the deviation from the start situation or position includes different start positions and / or postures of the robotic device.

[0043] For example, when the robot device 306 is tasked to pick up the bottle 316, open the cabinet 118, and place the bottle 316 inside the cabinet 318, the robot device 306 updates the parameters based on the mapping of pixels and / or descriptors (e.g., dense neural network descriptors) from the new image to the training image. The mapping defines the relative transformation between the new image and the training image.

[0044] In the relative transformation, the mapping from the training image to the new image is performed by matching the pixels and / or dense neural network descriptors of different images. The relative transformation may be defined by changes in the position of the robot device 306 on the x-axis, y-axis, z-axis, roll, pitch, and yaw. The relative transformation is used to update the parameters of the parameterized actions from the taught parameters to the observed situation. For example, the parameterized actions corresponding to the driving actions and / or operations may be adjusted to compensate for changes in the start position and / or posture of the robot device.

[0045] FIG. 3B shows an example of the robot device 306 executing a task from a different start position in an environment that is similar but different from the environment and start position shown in FIG. 1. The kitchen 308B is different from the kitchen 108 in FIG. 1. For example, the table 120 where the bottle 116 is placed in FIG. 1 is in a different position from the table 320 where the bottle 316 is placed in the example of FIG. 3. Furthermore, the arrangement of the bottle 316 in FIG. 3B is different from the arrangement of the bottle 116 in FIG. 1. In addition, the start position of the robot device 306 in FIG. 3B is different from the start position of the robot device 106 in FIG. 1.

[0046] Since the initial position of the robot device 306 and the arrangement of the bottle 316 are different from the initial position of the robot device 106 and the arrangement of the bottle 116 during training (see FIG. 1), when the robot device 306 is tasked with picking up the bottle 316, opening the cabinet 318, and placing the bottle 316 inside the cabinet 318, the robot device 306 is specified to update a set of parameterized operations. For example, the robot device 306 may adjust the parameterized operations corresponding to the angle at which the limbs of the robot device (e.g., the limbs 222 in FIG. 2) approach in order to improve the grasping of the bottle 316.

[0047] The robot control system is not limited to performing actions on identified targets. Aspects of the present disclosure may also be used to operate autonomous or semi-autonomous vehicles such as cars. As shown in FIG. 4, an operator may control a vehicle 400 (e.g., an autonomous vehicle) via a user interface such as a remote operation in an environment such as a city 402. The operator may be located at a position away from the city 402 of the vehicle 400. As discussed herein, feed video may be provided to the operator via one or more sensors of the vehicle 400. The sensors may include cameras such as light detection and ranging (LiDAR) sensors, radio detection and ranging (RADAR) sensors, and / or other types of sensors.

[0048] As shown in FIG. 4, the operator controls the vehicle 400 to move along a first road 404 towards an intersection with a second road 406. To avoid a collision with a first building 408, the vehicle 400 needs to turn right 412 or left 414 at the intersection. Similar to the robot device, the operations performed by the operator are parameterized through a virtual reality interface.

[0049] As discussed, aspects of the present disclosure relate to a mobile manipulation hardware and software system capable of autonomously performing human-level tasks in a real-world environment after being taught a task by demonstration from a human within virtual reality. In one configuration, a mobile manipulation robot is used. The robot may include full-body task-space hybrid position / force control. Additionally, as discussed, the robot is taught parameterized primitives linked to a dense visual embedding representation of a robustly learned landscape. And a task graph of the taught motion may be generated.

[0050] Rather than programming or training a robot to recognize a set of fixed objects or perform predefined tasks, aspects of the present disclosure allow a robot to learn new objects and tasks from human demonstrations. The learned tasks may be autonomously executed by the robot in a naturally changing situation. The robot can be taught to associate a given set of motions to any landscape and object from a single example without using previous object models or maps. The vision system may be trained offline using existing supervised and unsupervised datasets, and the rest of the system may function without additional training data.

[0051] In contrast to conventional systems that directly teach motions in task space, aspects of the present disclosure teach a set of parameterized motions. These motions combine a collision-free motion plan and hybrid (position and force) Cartesian control of the end effector to minimize the taught parameters and provide robustness during execution.

[0052] In one configuration, an embedding for trained dense vision specialized for a task is computed. This pixel-wise embedding links the parameterized motions to the landscape. This link allows the system to handle a wide variety of environments with high robustness in exchange for generalization to new situations.

[0053] The operation of the task may be taught independently using visual input conditions and end conditions based on success. The operations may be interconnected within a dynamic task graph. Since the operations are connected, the robot may reuse actions to execute the task sequence.

[0054] The robot may have multiple degrees of freedom (DOF). For example, the robot may have 31 degrees of freedom (DOF) divided into five subsystems: a chassis, a torso, a left arm, a right arm, and a head. In one configuration, the chassis includes four drive-steerable wheels (e.g., a total of 8 degrees of freedom) that achieve "quasi-holonomic" mobility. The drive / steer actuator package may include various motors and gearheads. The torso may have 5 degrees of freedom (yaw, pitch, pitch, pitch, yaw). Each arm may have 7 degrees of freedom. The head may have 2 degrees of freedom for pan / tilt. Each arm may include a 1-degree-of-freedom gripper with underactuated fingers. Aspects of the present disclosure are not limited to the robot discussed above. Other configurations are possible. In one example, the robot may include a custom tool such as a sponge or a mop.

[0055] In one configuration, the robot is integrated with a force / torque sensor for measuring the interaction force with the environment. For example, the force / torque sensors may be arranged at the wrists of each arm. The head may be integrated with a perception sensor that provides a wide field of view and a VR context for humans and robots to perform tasks.

[0056] Aspects of the present disclosure provide several levels of abstraction for robot control. In one configuration, at the lowest control level, real-time coordinated control of all degrees of freedom of the robot is provided. Real-time control may include joint control and component control. Joint control implements low-level device communication and exposes device commands and states in a general form. In addition, joint control supports actuators, force sensors, and inertial measurement devices. Joint control may be configured at runtime to support different robots.

[0057] By component control, the robot is divided into components (such as the right arm, head, etc.), and by providing a set of parameterized controllers for each component, higher-level collaborative actions of the robot can be handled. By component control, controllers for joint position and velocity, joint admittance, camera vision, vehicle position and velocity, and posture, velocity, and admittance control in the hybrid task space may be provided.

[0058] Task space control of the end effector enables the abstraction of robot control in another dimension. This level of abstraction solves the problem of the robot posture for achieving the desired motion. The inverse kinematics (IK) of the whole body for hybrid Cartesian control is formed and solved as a quadratic program. There may be linear constraints on the components regarding joint position, velocity, acceleration, and gravitational torque.

[0059] The IK of the whole body may be used for the motion plan to reach the goal of the posture in Cartesian coordinates. In one configuration, occupied voxels of the environment are fitted with spheres or capsules. To avoid collisions between the robot and the world, collision constraints of the voxels are added to the quadratic program of the IK. In the quadratic program of the IK, sampling in Cartesian space may be performed as a steering function between nodes, and a rapidly-exploring random tree (RRT) may be used for the motion plan.

[0060] The plan in Cartesian space results in natural and direct motions. By using the quadratic program of the IK as the steering function, the reliability of the plan can be improved, and the same controller may be used for both the plan and execution to reduce the discrepancy between the two. Similarly, an RRT is used by combining a motion plan towards the goal of joint position and a joint position controller by component control acting as the steering function.

[0061] Parameterized operations are defined at the next level of abstraction. In one configuration, the parameterized operations are primitive actions that can be parameterized and combined to accomplish tasks. Operations include, but are not limited to, manipulation actions such as grasping, lifting, placing, pulling, retracting, wiping, direct control, locomotion actions such as moving joints, driving by speed commands, driving by position commands, path following while performing active obstacle avoidance, and preparatory actions such as visually stopping.

[0062] Each operation can have one or more different types of actions, such as the movement of one or more joints of a robot part or movement in Cartesian coordinates. Each action can use different control techniques such as position, speed, or admittance control, and can choose to use an operation plan to avoid external obstacles. The operations of the robot, whether or not using an operation plan, avoid self-collision and satisfy the operation control constraints.

[0063] Each operation is parameterized by different actions, or alternatively, the actions may have their own parameters. For example, a grasping operation may consist of four parameters: gripper angle, 6D approach, grasping, and (optional) gripper pose during lifting. In this example, these parameters define the following predefined sequence of actions: (1) Open the gripper to the desired gripper angle. (2) Plan and execute a collision-free path to the 6D approach pose. (3) Move the gripper to the 6D grasping pose and stop when contact is made. (4) Close the gripper, and (5) Move the gripper to the 6D lift pose.

[0064] The abstraction of the final level of control is the task. In one configuration, the task is defined as a sequence of actions that enable the robot to perform operations and navigate in a human environment. The task graph (see Figure 5) is a valid, periodic or aperiodic graph with different tasks as nodes, different movement situations as edges, and including anomaly detection and recovery from anomalies. The edge situations include the execution status of each action for handling different objects and environments, the inspection of the object in hand using a force / torque sensor, voice commands, and the matching with keyframes.

[0065] According to an aspect of the present disclosure, a perception pipeline for the robot to understand the surrounding environment is designed. The perception pipeline also enables the robot to acquire the ability to recognize which actions should be taken based on the taught tasks. In one configuration, a fused RGB-D image is created by projecting a plurality of depth images of a high-resolution color stereo pair onto one field-of-view image (e.g., the left image of a wide field of view). The system runs a set of deep neural networks to provide various pixel-level classifications and feature vectors (e.g., embeddings). Based on the visual features called from the taught sequence, the pixel-level classifications and feature vectors are accumulated into a temporary 3D voxel representation. The pixel-level classifications and feature vectors may be used to call the actions to be performed.

[0066] In one configuration, the category of the object is not defined. Additionally, or the model of the object or the environment is not assumed. Instead of explicitly detecting and segmenting the object and explicitly estimating the 6-degree-of-freedom object pose, a dense pixel-level embedding may be generated for more general tasks. The reference embedding from the taught sequence may be used to perform action classification or pose estimation.

[0067] The trained model may be a complete convolutional model. In one configuration, the pixels of the input image are each mapped to a point in an embedding space. The embedding space is given a metric that is implicitly defined by a loss function and a training procedure defined by the output of the model. The trained model may be used for various tasks.

[0068] In one configuration, if one annotated example is given, the trained model detects all objects in the semantic class. The objects in the semantic class may be detected by comparing the embeddings in the annotation with the embeddings in other regions. The model may be trained by a discriminative loss function.

[0069] The model may be trained to determine object instances. This model identifies and / or counts individual objects. The model may be trained to predict a vector (2D embedding) for each pixel. The vector may indicate the centroid of the object containing that pixel. At runtime, pixels that point to the same centroid may be grouped as segments of that scene. The execution at runtime may be performed in 3D.

[0070] The model may be trained for 3D correspondences. This model provides, for each pixel, an embedding that is invariant to views and lighting such that views of any 3D point in the scene are mapped to the same embedding. The model may be trained using a loss function.

[0071] The embeddings (and depth data) for pixels for each RGB-D frame are fused into a dynamic 3D voxel map. Each voxel accumulates the statistics of the positions, colors, and embeddings of the first and second orders. The expiration of dynamic objects is based on the back-projection of the voxel depth image. The voxel map is segmented using standard graph segmentation based on semantic and instance labels, as well as geometric approximations. The voxel map is dimensionally reduced to a 2.5D map with elevation and traversability classification statistics.

[0072] The 2.5D map is used for the motion of the vehicle body without collision, while the voxel map is used for the motion planning of the whole body without collision. For the inspection of collisions in 3D, voxels in the map may be grouped into capsules using a greedy method. The segmented objects may be used for motions to be attached to the hand when the object is grasped.

[0073] The robot may be trained by a one-shot learning approach to recognize features in a landscape (or of a specific operating object) that are highly related to the features recorded in a task taught in the past. When the task is demonstrated by the user, the features are saved in the form of keyframes throughout the task. The keyframe may be an RGB image including a multi-dimensional embedding with depth (if available) for each pixel.

[0074] The embedding functions as a feature descriptor that can establish pixel-wise correspondence relationships at runtime under the assumption that the current image is sufficiently similar to the reference image that existed at the time of teaching. Since depth exists in (almost) all pixels, the correspondence relationship can be used to solve the pose delta between the current image and the reference image. Euclidean constraints may be used to detect inliers, and the Levenberg-Marquardt least squares function is applied together with RANSAC to solve for the 6-degree-of-freedom pose.

[0075] The pose deltas serve as the role of corrections applicable to adapt the sequence of the taught actions to the current scenery. Since the embedding may be defined for each pixel, the keyframe may be as wide as to include all the pixels in the image, or as narrow as to use only the pixels within a mask defined by the user. As discussed, the user may define the mask by selectively annotating regions in the image as being related to a task or being on an object.

[0076] In addition to visual sensing, in one configuration, the robot collects and processes voice input. The voice provides a set of another embedding as an input for teaching the robot. As an example, the robot asks questions and obtains voice input by understanding the spoken language of the responses from humans. The voice responses may be understood using a custom keyword detection module.

[0077] The robot may utilize a fully convolutional keyword spotting model to understand custom wake words, a set of objects (e.g., "mug" or "bottle"), and a set of locations (e.g., "cabinet" or "refrigerator"). In one configuration, the model listens for wake words at an interval, such as 32 ms. Once the wake word is detected, the robot pays attention to whether object or location keywords are detected. During training, artificial noise is added to make the recognition more robust.

[0078] As discussed, to teach a task to a robot, an operator uses a set of VR modes. Each action may have a corresponding VR mode to set and command the parameters specific to that action. Each action mode may include a customized visualization according to the type of parameter to assist in setting each parameter. For example, when setting the parameters of the motion of pulling a door, the hinge axis is labeled and visualized as a line, and the posture candidates for pulling the gripper are restricted to an arc centered on the hinge. To assist the teaching process, several utility VR modes such as motion restoration, annotation of the environment by related objects, repositioning of the virtual robot, camera images, and the menu of the VR world are used.

[0079] During execution, the posture of the robot and the parts in the environment may be different from those used during training. Feature matching may be used to discover features in an environment similar to that taught. The pose delta may be established from the correspondence of the matched features. The actions taught by the user may be changed by the calculated pose delta. In one configuration, multiple keyframes are passed to the matching problem. Based on the number of correspondences, the best-matched keyframe is selected.

[0080] FIG. 5 is a diagram showing an example of a hardware implementation of a robot control system 500 according to an aspect of the present disclosure. The robot control system 500 may be a component of an autonomous or semi-autonomous system, such as a vehicle, a robot device 528, or other devices. In the example of FIG. 5, the robot control system 500 is a component of the robot device 528. The robot control system 500 may be used to control the actions of the robot device 528 based on updating the parameters of a set of parameterized actions according to relative transformations for performing tasks in a task environment.

[0081] The robot control system 500 may be implemented by a bus architecture generally represented as bus 530. Bus 530 may include any number of interconnecting buses and bridges depending on the specific application of the robot control system 500 and overall design constraints. Bus 530 connects various circuits such as one or more processors and / or hardware modules represented as processor 520, communication module 522, position module 518, sensor module 502, movement module 526, memory 524, task module 508, and computer-readable medium 514. Bus 530 may also connect various other circuits known to those skilled in the art, such as a timing source, peripherals, voltage regulators, power management circuits, and thus no further description will be given.

[0082] The robot control system 500 includes a transceiver 516 connected to processor 520, sensor module 502, task module 508, communication module 522, position module 518, movement module 526, memory 524, and computer-readable medium 514. Transceiver 516 is connected to antenna 534. Transceiver 516 communicates with various devices via a transmission medium. For example, transceiver 516 may receive instructions (e.g., to start a task) from an operator of the robot device 528 via communication. As discussed herein, the operator may be located at a position remote from the robot device 528. In some aspects, the task may also be initiated within the robot device 528, for example, via task module 508.

[0083] The robot control system 500 includes a processor 520 connected to a computer-readable medium 514. The processor 520 performs processing including the execution of software stored on the computer-readable medium 514 that provides functions according to the present disclosure. The software, when executed by the processor 520, causes the robot control system 500 to execute various functions described for specific devices such as the robot device 528 or modules 502, 508, 514, 516, 518, 520, 522, 524, 526. The computer-readable medium 514 may also be used to store data that is operated on by the processor 520 when the software is executed.

[0084] The sensor module 502 may be used to obtain measurement values via different sensors such as a first sensor 506 and a second sensor 504. The first sensor 506 may be a visual sensor such as a stereo camera or an RGB camera for taking 2D images. The second sensor 504 may be a ranging sensor such as a LiDAR sensor or a RADAR sensor. Of course, aspects of the present disclosure are not limited to the above sensors, and for example, other types of sensors such as temperature, sound waves, and / or lasers may also be considered as either of the sensors 504, 506. The measurement values by the first sensor 506 and the second sensor 504 may be processed by one or more of the processor 520, the sensor module 502, the communication module 522, the position module 518, the movement module 526, and the memory 524 in conjunction with the computer-readable medium 514 to implement the functions described herein. In one configuration, the data captured by the first sensor 506 and the second sensor 504 may be transmitted to the operator as a supplied video via the transceiver 516. The first sensor 506 and the second sensor 504 may be connected to the robot device 528 or may be in communication with the robot device 528.

[0085] The position module 518 may be used to determine the position of the robot device 528. For example, the position module 518 may use the Global Positioning System (GPS) to determine the position of the robot device 528. The communication module 522 may be used to facilitate communication via the transceiver 516. For example, the communication module 522 may provide communication capabilities via different wireless protocols such as WiFi, Long Term Evolution (LTE), 3G, etc. The communication module 522 may also be used to communicate with other components of the robot device 528 that are not modules of the robot control system 500.

[0086] The movement module 526 may be used to facilitate the movement of the robot device 528 and / or components of the robot device 528 (such as limbs, hands, etc.). For example, the movement module 526 may control the movement of the limbs 538 and / or the wheels 532. As another example, the movement module 526 may be in communication with a power source of the robot device 528 such as an engine or a battery. Of course, aspects of the present disclosure are not limited to providing movement via a propeller, and other types of components that provide movement such as treads, fins, and / or jet engines are also contemplated.

[0087] The robot control system 500 also includes a memory 524 for storing data related to the operation of the robot device 528 and the task module 508. The modules may be software modules executed within the processor 520, those resident / stored on the computer-readable medium 514 and / or the memory 524, one or more hardware modules connected to the processor 520, or combinations thereof.

[0088] The task module 508 may be communicable with the sensor module 502, transceiver 516, processor 520, communication module 522, location module 518, movement module 526, memory 524, and computer-readable medium 514. In one configuration, the task module 508 includes a parameterized operation module 510, an action module 512, and an object identification module 536. The object identification module 536 may identify an object near the robotic device 528. That is, based on the input received from sensors 504, 506 via the sensor module 502, the object identification module 536 identifies an object (e.g., a target). The object identification module 536 may be a trained object classifier (e.g., a convolutional neural network).

[0089] The identified object may be output to the parameterized operation module 510 to map pixels and / or dense neural network descriptors from the current scene to those from the keyframe, and as a result, adjustments or updates to the parameterized operations may be continuously made. For example, the adjustment may be based on the relative transformation between the task image and the teaching image (e.g., the keyframe) based on the mapping. The update of the parameters of the set of parameterized operations is based on the relative transformation. The action module 512 facilitates the execution of the cushion / task to the robotic device including the updated parameterized operations. The parameterized operations may be stored in the memory 524. For example, the updated parameterized operations are output to at least the movement module 526 for the robotic device 528 to execute the updated parameterized operations.

[0090] FIG. 6 shows an example of a graphical sequence 600 of the taught operations according to an aspect of the present disclosure. As shown in FIG. 6, the graphical sequence 600 includes a start node 602 and an end node 604. The graphical sequence 600 may branch or loop based on the sensed visual input, audio input, or other situations.

[0091] For example, as shown in FIG. 6, after the start node 602, the robot may execute the operation of "listen_for_object". In this example, the robot determines whether it has sensed visual or audio input corresponding to a cup or a bottle. In this example, different operation sequences are executed based on whether the sensed input corresponds to a cup or a bottle. The aspects of the present disclosure are not limited to the operations shown in FIG. 6.

[0092] FIG. 7 shows an example of software modules for a robot system according to an aspect of the present disclosure. The software modules of FIG. 7 may use one or more components of the hardware system of FIG. 5, such as a processor 520, a communication module 522, a position module 518, a sensor module 502, a movement module 526, a memory 524, a task module 508, and a computer-readable medium 514. The aspects of the present disclosure are not limited to the modules shown in FIG. 7.

[0093] As shown in FIG. 7, the robot may receive audio data 704 and / or image data / input 702. The image input 702 may be an RGB-D image. The audio network 706 may listen for wake words at certain intervals. The audio network 706 receives the raw audio data 704 to detect wake words and extract keywords from the raw audio data 704.

[0094] A neural network, such as a dense embedding network 708, receives the image data 702. The image data 702 may be received at certain intervals. The dense embedding network 708 processes the image input 702 and outputs an embedding 710 of the image input 702. The embedding 710 and the image data 702 may be combined to generate a voxel map 712. The embedding 710 may also be input to a keyframe matcher 712.

[0095] The keyframe matcher 712 compares the embedding 710 with a plurality of keyframes. When the embedding 710 corresponds to the embedding of a keyframe, the matching keyframe is identified. The embedding 710 may include a pixel descriptor, depth information, and other information.

[0096] The task module 714 may receive one or more task graphs 716. The task module 714 provides a response to a request from the keyframe matcher 712. The keyframe matcher 712 matches tasks to the matching keyframes. The tasks may be determined from the task graph 716.

[0097] The task module 714 may also send an operation request to the operation module 718. The operation module 718 provides an operation status to the task module 714. In addition, the operation module 718 may request information about the matching keyframe and the corresponding task from the keyframe matcher 712. The keyframe matcher 712 provides information about the matching keyframe and the corresponding task to the operation module 718. The operation module 718 may receive voxels from the voxel map 712.

[0098] In one configuration, the operation module 718 receives an operation plan from the operation planner 720 in response to an operation plan request. The operation module 718 also receives the part status from the part control module 722. The operation module 718 sends a part command to the part control module 722 in response to receiving the part status. Then, the part control module 722 receives the joint status from the joint control module 724. The part control module 722 sends a joint command to the joint control module 724 in response to receiving the joint status.

[0099] FIG. 8 shows a method 800 for controlling a robot device according to an aspect of the present disclosure. At block 802, when performing a task, the robot device is located within a task environment. The robot device is arranged to deviate from a start situation or position where the robot device is taught to perform a task in a teaching environment using a virtual reality (VR) interface. The task environment is similar to or the same as the teaching environment.

[0100] At block 804, when the robot device is tasked to execute a set of parameterized operations taught within a sequence controlled by virtual reality related to the task, pixels of a task image of the scenery within the task environment and / or a neural network descriptor are mapped to a teaching image of the teaching environment. At block 806, a relative transformation between the task image and the teaching image is defined based on the mapping. At block 808, the parameters of the set of parameterized operations are updated based on the relative transformation to perform the task in the task environment.

[0101] The various operations of the methods described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software components and / or modules including, but not limited to, a circuit, an application specific integrated circuit (ASIC), or a processor. When there are operations shown in the figures, these operations may have corresponding functional components assigned generally similar numbers.

[0102] As used herein, "determining" includes a wide variety of actions. For example, "determining" can include calculating, computing, processing, deriving, investigating, searching (e.g., searching within a table, database, or other structure), probing, etc. Additionally, "determining" can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Further, "determining" can include resolving, selecting, choosing, establishing, etc.

[0103] As used herein, the phrase "at least one of" refers to any combination of items from a list of items, including a single item. For example, "at least one of a, b, or c" is intended to include a, b, c, a-b, a-c, b-c, a-b-c.

[0104] The various illustrative logical blocks, modules, and circuits described in connection with the present disclosure may be implemented or executed by a processor configured according to the present disclosure, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array signal (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination of the above designed to perform the functions described herein. The processor may be a microprocessor, a controller, a microcontroller, or a state machine configured as described herein. The processor may also be implemented as a combination of computing devices, such as, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors combined with a DSP core, or other special configurations described herein.

[0105] The steps or algorithms of the methods described in connection with this disclosure may be embodied directly in hardware, in software modules executed by a processor, or in a combination of the two. The software modules may be present in a storage device, or a machine-readable medium, including any medium that can be used to carry or store the desired program code in the form of instructions or data structures and that is accessible by a computer, including Random Access Memory (RAM), Read Only Memory (ROM), Flash Memory, Erasable Programmable Read Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), registers, hard disk, removable disk, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other medium. The software modules may comprise a single instruction, or many instructions, and may be distributed over several different code segments, different programs, and multiple storage media. The storage media may be connected to the processor such that the processor can read information from, and write information to, the storage media. Alternatively, the storage media may be integral to the processor.

[0106] The methods disclosed herein include one or more steps or actions for implementing the disclosed methods. The steps and / or actions of the methods may be interchanged with each other without departing from the scope of the claims. In other words, the order and / or use of specific steps and / or actions may be changed without departing from the scope of the claims, unless a specific order of the steps and / or actions is specified.

[0107] The described functionality may be implemented by hardware, software, firmware, or any combination thereof. When implemented in hardware, an example of a hardware configuration may include a processing system within the device. The processing system may be implemented using a bus architecture. The bus may include any number of interconnecting buses and bridges depending on the specific application of the processing system and the overall design constraints. The bus may connect various circuits including a processor, a machine-readable medium, and a bus interface. The bus interface may be used to connect, among other things, a network adapter to the processing system via the bus. The network adapter may be used to implement signal processing functions. In certain aspects, a user interface (e.g., keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also connect various other circuits known to those of ordinary skill in the art, such as a timing source, peripherals, voltage control, power management circuits, etc., and thus will not be described further herein.

[0108] The processor may be responsible for processing including managing the bus and executing software stored on the machine-readable medium. Software shall be construed to mean instructions, data, or any combination thereof, regardless of the terminology used, such as software, firmware, middleware, microcode, hardware description language, or other names.

[0109] In a hardware implementation, the machine-readable medium may be part of a processing system separate from the processor. However, as will be readily understood by those skilled in the art, the machine-readable medium, or any portion thereof, may be external to the processing system. For example, the machine-readable medium may include a communication line, a carrier wave modulated by data, and / or a computer product detached from the device, all of which may be accessed by the processor via a bus interface. Alternatively, or in addition, the machine-readable medium, or a portion thereof, may be integrated into the processor as may be the case where a cache and / or a special register file may exist. Although the various components discussed have been described as having a particular location like local components, they may be configured in various ways like particular components configured as part of a distributed computing system.

[0110] The processing system may be composed of one or more microprocessors that provide processor functionality, and at least a portion of the machine-readable medium and external memory, all of which may be connected through a support circuit by an external bus architecture. Alternatively, the processing system may comprise one or more neuromorphic processors for implementing the neuron models and neural system models described herein. As another alternative, the processing system may be implemented by an application-specific integrated circuit (ASIC) having a processor, a bus interface, a user interface, a support circuit, and at least a portion of the machine-readable medium integrated on a single chip, or one or more Field Programmable Gate Arrays (FPGAs), Programmable Logic Devices (PLDs), controllers, state machines, gate logic, individual hardware components, or other suitable circuits, or any combination of circuits capable of performing the various functions described within this disclosure. Those skilled in the art will recognize how best to implement the described functionality of the processing system based on the particular application and the overall design constraints imposed on the system as a whole.

[0111] A machine-readable medium may comprise a number of software modules. The software modules may include a transmission module and a reception module. Each software module may reside within a single storage device or may be distributed across a plurality of storage devices. For example, a software module may be loaded from a hard drive to RAM when a triggering event occurs. During execution of the software module, the processor may load some instructions into a cache to increase access speed. One or more cache lines may then be loaded into a special-purpose register file for execution by the processor. By referring to the following functions of the software module, it will be understood that the functions are performed by the processor when the software module executes instructions. Further, it should be understood that aspects of the present disclosure improve the functions of a processor, computer, machine, or other system implementing such aspects.

[0112] If implemented in software, the functions may be stored or transmitted on a computer-readable medium as one or more instructions or code. The computer-readable medium includes both a computer storage device and a communication medium including any storage device that facilitates transfer of a computer program from one place to another. Additionally, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared (IR), radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as IR, radio, and microwave are included in the definition of the medium. As used herein, disk and disc include compact disc (CD), laser disc (registered trademark), optical disc, digital versatile disc (DVD), floppy (registered trademark) disc, and Blu-ray (registered trademark) disc, where disk typically magnetically reproduces data and disc optically reproduces data using a laser. Thus, in some aspects, the computer-readable medium may comprise a non-transitory computer-readable medium (e.g., a tangible medium). Additionally, in other aspects, the computer-readable medium may comprise a transitory computer-readable medium (e.g., a signal). The above combinations should be included within the scope of the computer-readable medium.

[0113] Accordingly, certain aspects may comprise a computer program product for performing the operations presented herein. For example, such a computer program product may comprise a computer-readable medium storing (and / or encrypting) instructions executable by one or more processors for performing the operations described herein. In certain aspects, the computer program product may include packaging materials.

[0114] Furthermore, it should be understood that the modules and / or other suitable means for carrying out the methods and techniques described herein may be downloadable and / or obtainable by the user terminal and / or the base station as necessary. For example, such devices can be connected to a server to facilitate the transfer of means for carrying out the methods described herein. Alternatively, the various methods described herein can be provided via a storage means in a form that enables the user terminal and / or the base station to obtain the various methods by connecting the storage means to the device or providing the storage means to the device. Furthermore, any other techniques for providing the methods and techniques described herein to the device can be used.

[0115] It should be understood that the claims are not limited to the exact configurations and components shown above. Various modifications, changes, and variations can be made to the arrangements, operations, and details of the methods and devices described above without departing from the scope of the claims. The invention disclosed in this specification includes the following aspects. 〔Aspect 1〕 Placing a robot device in a task environment, Mapping a descriptor of a task image of a scene in the task environment to a teaching image in a teaching environment, Defining a relative transformation between the task image and the teaching image based on the mapping, Updating parameters of a set of parameterized operations based on the relative transformation to execute a task corresponding to the teaching image, A method for controlling a robot device, including. 〔Aspect 2〕 The method according to Aspect 1, wherein the set of parameterized operations includes operations executed by a user while the robot device is trained using a virtual reality interface to execute the task. 〔Aspect 3〕 The method according to Aspect 1, further including mapping the descriptor from the current scene of the task image to the descriptor from the teaching image with an interval. 〔Aspect 4〕 The robot device is arranged deviating from the start situation or position used during training in the teaching environment, The task environment is similar to or the same as the teaching environment, The method according to Aspect 1. 〔Aspect 5〕 The method according to Aspect 4, wherein the deviation from the start situation or position includes different start positions and / or postures of the robot device. 〔Aspect 6〕 The method according to Aspect 4, wherein the deviation from the start situation or position includes different start positions and / or postures of an object that is the target for which the task is executed. 〔Aspect 7〕 The method according to Aspect 1, wherein the descriptor has a pixel descriptor or a neural network descriptor. 〔Aspect 8〕 A memory, Comprising at least one processor connected to the memory, and the at least one processor is Configured to place a robot device in a task environment, Map a descriptor of a task image of a scene in the task environment to a teaching image in a teaching environment, Define a relative transformation between the task image and the teaching image based on the mapping, Update parameters of a set of parameterized operations based on the relative transformation to execute a task corresponding to the teaching image, An apparatus for controlling a robot device, configured as such. 〔Aspect 9〕 The apparatus according to aspect 8, wherein the set of parameterized operations includes operations performed by a user while the robotic device is trained using a virtual reality interface to perform the task. 〔Aspect 10〕 The apparatus according to aspect 8, wherein the at least one processor is further configured to map the descriptor from the current scene of the task image to the descriptor from the teaching image at intervals. 〔Aspect 11〕 The robotic device is arranged deviating from the start situation or position used during training in the teaching environment, and the task environment is similar to or the same as the teaching environment. The apparatus according to aspect 8. 〔Aspect 12〕 The apparatus according to aspect 11, wherein the deviation from the start situation or position includes different start positions and / or postures of the robotic device. 〔Aspect 13〕 The apparatus according to aspect 11, wherein the deviation from the start situation or position includes different start positions and / or postures of the object which is the target for which the task is performed. 〔Aspect 14〕 The apparatus according to aspect 8, wherein the descriptor has a pixel descriptor or a neural network descriptor. 〔Aspect 15〕 A non-transitory computer-readable medium recording program code for controlling a robotic device, wherein the program code is executed by a processor, program code for placing the robotic device in a task environment, program code for mapping a descriptor of a task image of a scene in the task environment to a teaching image of a teaching environment, program code for defining a relative transformation between the task image and the teaching image based on the mapping, and program code for updating parameters of a set of parameterized operations based on the relative transformation to perform a task corresponding to the teaching image. The non-transitory computer-readable medium comprising 。 〔Aspect 16〕 The non-transitory computer-readable medium according to aspect 15, wherein the set of parameterized operations includes operations performed by a user while the robotic device is trained using a virtual reality interface to perform the task. 〔Aspect 17〕 The non-transitory computer-readable medium according to aspect 15, wherein the program code further includes program code for mapping the descriptor from the current scenery in the task image to the descriptor from the teaching image at intervals. 〔Aspect 18〕 The robot device is arranged deviating from the start situation or position used during training in the teaching environment, The task environment is similar to or the same as the teaching environment. The non-transitory computer-readable medium according to aspect 15. 〔Aspect 19〕 The non-transitory computer-readable medium according to aspect 18, wherein the deviation from the start situation or position includes different start positions and / or postures of the robot device. 〔Aspect 20〕 The non-transitory computer-readable medium according to aspect 18, wherein the deviation from the start situation or position includes different start positions and / or postures of the object on which the task is to be executed.

Claims

1. A method for controlling a robot device, comprising: Placing the robot device in a task environment; Mapping a plurality of task image pixel descriptors associated with pixels in a task image of a scene in the task environment to a plurality of teaching image pixel descriptors associated with pixels in a teaching image of a teaching environment, wherein each task image pixel descriptor of the plurality of task image pixel descriptors includes a first pixel value associated with a pixel of the pixel in the task image, and each teaching image pixel descriptor of the plurality of teaching image pixel descriptors includes a second pixel value associated with a pixel of the pixel in the teaching image, and the first pixel value associated with each task image pixel descriptor has the same value as the second pixel value associated with the teaching image pixel descriptor mapped to each task image pixel descriptor; Defining a relative transformation between the task image and the teaching image based on the mapping of the plurality of task image pixel descriptors, the relative transformation indicating changes in the X-axis, Y-axis, Z-axis, roll, pitch, and yaw between the task image and the teaching image; Updating parameters of a set of operations that are joint angle changes of the parameterized robot device based on the relative transformation to execute a task corresponding to the teaching image; The method comprising the above steps.

2. The method according to claim 1, wherein the set of parameterized operations includes operations performed by a user while the robot device is trained using a virtual reality interface to perform the task.

3. The method according to claim 1, wherein the plurality of task image pixel descriptors are continuously mapped to the plurality of teaching image pixel descriptors.

4. The method according to claim 1, wherein a start position and / or an initial posture of the robot device in the task environment are different from a start position and / or an initial posture of the robot device during training in the teaching environment.

5. The method according to claim 4, wherein a start position and / or a posture of an object that is a target for which the task is executed in the task environment are different from a start position and / or a posture of an object that is a target for which the task is executed during training in the teaching environment.

6. An apparatus for controlling a robot device disposed in a task environment, a memory, and at least one processor connected to the memory, the at least one processor maps a plurality of task image pixel descriptors associated with pixels in a task image of a landscape in a task environment to a plurality of teaching image pixel descriptors associated with pixels in a teaching image of a teaching environment, wherein each task image pixel descriptor of the plurality of task image pixel descriptors includes a first pixel value associated with a pixel of the pixel in the task image, each teaching image pixel descriptor of the plurality of teaching image pixel descriptors includes a second pixel value associated with a pixel of the pixel in the teaching image, and the first pixel value associated with each task image pixel descriptor has the same value as the second pixel value associated with the teaching image pixel descriptor mapped to each task image pixel descriptor, defines a relative transformation between the task image and the teaching image based on the mapping of the plurality of task image pixel descriptors, the relative transformation indicating changes in the X-axis, Y-axis, Z-axis, roll, pitch, and yaw between the task image and the teaching image, updates parameters of a set of operations that are changes in joint angles of the parameterized robot device based on the relative transformation to execute a task corresponding to the teaching image. An apparatus configured as such.

7. The apparatus according to claim 6, wherein the set of parameterized operations includes operations performed by a user while the robotic device is trained using a virtual reality interface to perform the task.

8. The apparatus according to claim 6, wherein the plurality of task image pixel descriptors are continuously mapped to the plurality of teaching image pixel descriptors.

9. The apparatus according to claim 6, wherein a start position and / or an initial posture of the robotic device in the task environment is different from a start position and / or an initial posture of the robotic device during training in the teaching environment.

10. The apparatus according to claim 9, wherein a start position and / or a posture of an object which is a target for the task to be performed in the task environment is different from a start position and / or a posture of an object which is a target for the task to be performed during training in the teaching environment.

11. A computer for controlling a robotic device disposed in a task environment, mapping a plurality of task image pixel descriptors associated with pixels in a task image of a landscape in the task environment to a plurality of teaching image pixel descriptors associated with pixels in a teaching image of a teaching environment, each task image pixel descriptor of the plurality of task image pixel descriptors includes a first pixel value associated with the pixel of the pixel in the task image, each teaching image pixel descriptor of the plurality of teaching image pixel descriptors includes a second pixel value associated with the pixel of the pixel in the teaching image, and the first pixel value associated with each task image pixel descriptor has the same value as the second pixel value associated with the teaching image pixel descriptor mapped to each task image pixel descriptor, and Based on the mapping of a plurality of task image pixel descriptors, a relative transformation between a task image and a teaching image, defining a relative transformation indicating changes in the X-axis, Y-axis, Z-axis, roll, pitch, and yaw between the task image and the teaching image; Updating the parameters of a set of operations that are the joint angle changes of the parameterized robot device based on the relative transformation in order to execute the task corresponding to the teaching image; A non-transitory computer-readable medium recording program code for causing the execution.

12. The non-transitory computer-readable medium according to claim 11, wherein the set of parameterized operations includes operations performed by a user while the robot device is trained using a virtual reality interface to execute the task.

13. The non-transitory computer-readable medium according to claim 11, wherein the plurality of task image pixel descriptors are continuously mapped to the plurality of teaching image pixel descriptors.

14. The non-transitory computer-readable medium according to claim 11, wherein the start position and / or initial posture of the robot device in the task environment is different from the start position and / or initial posture of the robot device during training in the teaching environment.

15. The non-transitory computer-readable medium according to claim 14, wherein the start position and / or posture of the object that the task is to be performed on in the task environment is different from the start position and / or posture of the object that the task is to be performed on during training in the teaching environment.

16. The method according to claim 1, wherein the task environment and the teaching environment are associated with the same environment.

17. The apparatus according to claim 6, wherein the task environment and the teaching environment are associated with the same environment.

Citation Information

Patent Citations

  • Data describing method and data processor

    JP2000287166A

  • Method of representing image group, descriptor of image group, searching method of image group, computer-readable storage medium, and computer system

    JP2010225172A

  • Robot control system using virtual reality input

    JP2017519644A

  • Work robot

    WO2007138987A1