Shared dense network with robot task-specific heads

By combining a shared dense network with a task-specific head, the problems of wasted computing resources and low training efficiency in existing robot systems are solved, achieving efficient utilization of computing resources and reducing human-computer interaction, thus improving the flexibility and efficiency of robot vision tasks.

CN114746906BActive Publication Date: 2025-11-11X DEVELOPMENT LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080081147.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-17
Filing Date
2020-11-19
Publication Date
2025-11-11
Estimated Expiration
2040-11-19

AI Technical Summary

Technical Problem

Existing robotic systems require training a separate network for each of the multiple vision tasks, which leads to a waste of computing resources and low training efficiency, making it difficult to efficiently utilize computing resources and reduce reliance on human-computer interaction.

Method used

By combining a shared dense network with a task-specific head, feature values ​​are extracted through training the shared dense network, and task-specific heads are added as needed to complete different robot vision tasks, thereby reducing redundant calculations and improving training efficiency.

Benefits of technology

It enables efficient use of computing resources in various robot vision tasks, reduces the need for human-computer interaction, improves training and operation efficiency, and enhances the system's flexibility and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114746906B_ABST
    Figure CN114746906B_ABST
Patent Text Reader

Abstract

A method includes receiving image data representing the environment of a robotic device from a camera on the robotic device. The method further includes applying a trained dense network to the image data to generate a set of feature values, wherein the trained dense network has been trained to perform a first robotic vision task. The method further includes applying a trained task-specific head to the set of feature values ​​to generate a task-specific output to perform a second robotic vision task, wherein the trained task-specific head has been trained to perform the second robotic vision task based on feature values ​​previously generated by the trained dense network, wherein the second robotic vision task differs from the first robotic vision task. The method also includes controlling the robotic device to operate in the environment based on the task-specific output generated to perform the second robotic vision task.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Patent Application No. 16 / 717,498, filed December 17, 2019, the entire contents of which are incorporated herein by reference. Technical Field Background Technology

[0004] With technological advancements, various types of robotic devices are being created to perform a wide range of functions that can assist users. These devices can be used in applications involving material handling, transportation, welding, assembly, and distribution. Over time, the operation of these robotic systems is becoming increasingly intelligent, efficient, and intuitive. As robotic systems become more prevalent in many aspects of modern life, there is a growing desire for their efficiency. Therefore, the need for efficient robotic systems has helped to open up new areas of innovation in actuators, motion, sensing technologies, and component design and assembly. Summary of the Invention

[0005] The example implementations involve shared dense networks, such as Feature Pyramid Networks (FPNs), which are combined with task-specific heads to accomplish different robot vision tasks.

[0006] In one embodiment, a method includes receiving image data representing the environment of a robotic device from a camera on the robotic device. The method further includes applying a trained dense network to the image data to generate a set of feature values, wherein the trained dense network has been trained to perform a first robotic vision task. The method further includes applying a trained task-specific head to the set of feature values ​​to generate a task-specific output to perform a second robotic vision task, wherein the trained task-specific head has been trained to perform the second robotic vision task based on feature values ​​previously generated by the trained dense network, wherein the second robotic vision task differs from the first robotic vision task. The method also includes controlling the robotic device to operate in the environment based on the task-specific output generated to perform the second robotic vision task.

[0007] In another embodiment, the robotic device includes a camera and a control system configured to receive image data representing the environment of the robotic device from the camera. The control system may be further configured to apply a trained dense network to the image data to generate a set of feature values, wherein the trained dense network has been trained to perform a first robotic vision task. The control system may also be further configured to apply a trained task-specific head to the set of feature values ​​to generate a task-specific output to perform a second robotic vision task, wherein the trained task-specific head has been trained to perform the second robotic vision task based on feature values ​​previously generated by the trained dense network, wherein the second robotic vision task differs from the first robotic vision task. The control system may also be configured to control the robotic device to operate in the environment based on the task-specific output generated to perform the second robotic vision task.

[0008] In yet another embodiment, a non-transitory computer-readable medium is provided, comprising programming instructions executable by at least one processor to cause the at least one processor to perform functions. The functions include receiving image data representing the environment of the robotic device from a camera on the robotic device. The functions further include applying a trained dense network to the image data to generate a set of feature values, wherein the trained dense network has been trained to perform a first robotic vision task. The functions further include applying a trained task-specific head to the set of feature values ​​to generate a task-specific output to perform a second robotic vision task, wherein the trained task-specific head has been trained to perform the second robotic vision task based on feature values ​​previously generated by the trained dense network, wherein the second robotic vision task differs from the first robotic vision task. The functions also include controlling the robotic device to operate in the environment based on the task-specific output generated to perform the second robotic vision task.

[0009] In another embodiment, a system is provided that includes means for receiving image data representing the environment of a robotic device from a camera on the robotic device. The system further includes means for applying a trained dense network to the image data to generate a set of feature values, wherein the trained dense network has been trained to perform a first robot vision task. The system further includes means for applying a trained task-specific head to the set of feature values ​​to generate a task-specific output to perform a second robot vision task, wherein the trained task-specific head has been trained to perform the second robot vision task based on feature values ​​previously generated by the trained dense network, wherein the second robot vision task differs from the first robot vision task. The system also includes means for controlling the robotic device to operate in the environment based on the task-specific output generated to perform the second robot vision task.

[0010] The foregoing description of the invention is merely illustrative and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent from the accompanying drawings and the following detailed description. Attached Figure Description

[0011] Figure 1 The illustration shows the configuration of a robot system according to an example embodiment.

[0012] Figure 2 The illustration depicts a mobile robot according to an example embodiment.

[0013] Figure 3 An exploded view of a mobile robot according to an example embodiment is shown.

[0014] Figure 4 The illustration shows a robotic arm according to an example embodiment.

[0015] Figure 5 This is a side view of image data representing the environment captured by a robot according to an example embodiment.

[0016] Figure 6A The illustration shows the training of a feature pyramid network according to an example embodiment.

[0017] Figure 6B The illustration shows the training of a robot task head according to an example embodiment.

[0018] Figure 6C The illustration shows the runtime application of multiple robot task heads according to an example embodiment.

[0019] Figure 7 This is a system design diagram based on an example embodiment.

[0020] Figure 8 This is a block diagram of a method according to an example embodiment. Detailed Implementation

[0021] This document describes example methods, devices, and systems. It should be understood that the terms "example" and "exemplary" as used herein mean "serving as an example, instance, or illustration." Any embodiment or feature described herein as "example" or "exemplary" is not necessarily to be construed as more preferred or advantageous than other embodiments or features unless so indicated. Other embodiments may be used, and other changes may be made, without departing from the scope of the subject matter set forth herein.

[0022] Therefore, the exemplary embodiments described herein are not intended to be limiting. It will be readily understood that the aspects of this disclosure as generally described herein and shown in the figures can be arranged, substituted, combined, separated, and designed in a variety of different configurations.

[0023] Throughout the description, the articles “(a)” or “(an)” are used to introduce elements of the example embodiments. Unless otherwise stated, or unless the context clearly specifies otherwise, any reference to “a” or “an” means “at least one”, and any reference to “the / said” means “the / said at least one”. The conjunction “or” is used in the list of at least two terms described to indicate any listed term or any combination of listed terms.

[0024] The use of ordinal numbers such as “first,” “second,” and “third” is to distinguish individual elements, not to indicate a specific order of these elements. For the purposes of this description, the terms “multiple” and “more than one” mean “two or more” or “more than one.”

[0025] Furthermore, unless the context otherwise requires, the features shown in each figure can be used in combination with each other. Therefore, Figure 1 The figures should generally be considered as constituent aspects of one or more overall embodiments, but it should be understood that not all features shown are necessary for every embodiment. In the figures, similar symbols generally identify similar parts unless the context otherwise requires. Furthermore, unless otherwise stated, the figures are not drawn to scale and are for illustrative purposes only. Moreover, the figures are representative only and do not show all parts. For example, additional structural or constraint parts may not be shown.

[0026] Furthermore, any enumeration of elements, frames, or steps in this specification or claims is for clarity only. Therefore, such enumeration should not be construed as requiring or implying that these elements, frames, or steps follow a particular arrangement or are performed in a particular order.

[0027] I. Overview

[0028] Robots can process image data to determine how to operate within an environment. Robots can perform a variety of different tasks, and each task may require different information about the environment. As one example, a robot might need to identify areas in the environment that it needs to clean (e.g., by emptying empty plates). As another example, a robot might need to identify different types of disposable objects in order to sort them into the appropriate trash can or recycling bin. As yet another example, a robot might need to identify doorknobs to enter or leave an area. Different robotic tasks may be associated with different robot vision tasks, each of which involves processing image data to obtain the specific information needed to complete the corresponding robotic task.

[0029] For many such robot vision tasks, machine learning can be used to train and control the robot. To use machine learning models (e.g., neural networks) efficiently, the robot vision stack can be optimized. More specifically, robot vision systems can be designed to allow automated training and deployment while minimizing the need for human-in-the-loop and the latency between data collection and result review. Additionally, robot vision systems can be designed to optimize performance quality while reducing the computational resources used by the robot during runtime. Furthermore, robot vision systems can be designed to reduce complexity while also making the system more robust, easier to maintain, and more responsive to changing needs.

[0030] With this goal in mind, the example robot vision system described herein involves a single shared dense network (e.g., FPN) combined with task-specific heads to accomplish different robot vision tasks. Sharing a single dense network allows for heavy computation to be performed only once for each image, while each head allows for task-specific optimization. Instead of training a separate network for each task, a dense network (e.g., hundreds of layers) is trained to extract features from many images captured by the robot or similar robots. Individual add-ons can then be added as needed for task-specific applications. The dense network may only need to be trained infrequently (e.g., monthly, quarterly, or yearly). New task-specific heads may require only a few layers and are therefore able to be trained quickly.

[0031] In some examples, a shared dense network can initially be trained to produce task-specific outputs based on image data to complete a first robot vision task, which may also be called a seed task. For example, the first robot vision task might involve identifying areas in the environment that should be handled by the robot. As part of this training process, the shared dense network can be trained to produce a rich set of feature values ​​after training on many images. These feature values ​​can then be used to facilitate the training of task-specific heads to perform other tasks within the same general space as the robot's perception task. For example, a task-specific head could be trained to recognize aspects of the environment that can be manipulated by the robot in different ways, such as recognizing objects that can be grasped by the robot's gripper. As another example, perception tasks logically more distant from the initial seed task can also benefit from prior training. For example, different task-specific heads can be trained to provide outputs to assist in tasks involving the robot in different ways, such as recognizing objects currently in the robot's gripper.

[0032] Therefore, the example framework described in this paper can be extended to leverage prior training to accomplish a variety of different robot perception tasks, while minimizing the additional training required and the computational resources needed for the robot to complete the task.

[0033] II. Example Robot System

[0034] Figure 1 An example configuration of a robotic system that can be used in conjunction with the embodiments described herein is illustrated. Robotic system 100 can be configured to operate autonomously, semi-autonomously, or using instructions provided by one or more users. Robotic system 100 can be implemented in various forms, such as a robotic arm, an industrial robot, or some other arrangement. Some example embodiments involve robotic system 100 designed to be manufactured at large scale and at low cost and designed to support a variety of tasks. Robotic system 100 can be designed to operate in the presence of a human. Robotic system 100 can also be optimized for machine learning. Throughout the description, robotic system 100 may also be referred to as a robot, robotic device, or mobile robot, among other names.

[0035] like Figure 1 As shown, robot system 100 may include one or more processors 102, data storage devices 104, and one or more controllers 108, which together may be part of control system 118. Robot system 100 may also include one or more sensors 112, one or more power sources 114, mechanical components 110, and electrical components 116. Nevertheless, robot system 100 is shown for illustrative purposes and may include more or fewer components. The various components of robot system 100 can be connected in any way, including wired or wireless connections. Furthermore, in some examples, the components of robot system 100 may be distributed across multiple physical entities rather than a single physical entity. Other illustrative examples of robot system 100 may also exist.

[0036] One or more processors 102 may operate as one or more general-purpose hardware processors or special-purpose hardware processors (e.g., digital signal processors, application-specific integrated circuits, etc.). One or more processors 102 may be configured to execute computer-readable program instructions 106 and manipulate data 107, both of which are stored in data storage device 104. One or more processors 102 may also interact directly or indirectly with other components of the robot system 100 (e.g., one or more sensors 112, one or more power sources 114, mechanical components 110, or electrical components 116).

[0037] Data storage device 104 may be one or more types of hardware memory. For example, data storage device 104 may include one or more computer-readable storage media or in the form thereof that can be read or accessed by processor(s)102. The one or more computer-readable storage media may include volatile or non-volatile storage components, such as optical, magnetic, organic, or other types of memory or storage devices, which may be integrally or partially integrated with processor(s)102. In some embodiments, data storage device 104 may be a single physical device. In other embodiments, two or more physical devices may be used to implement data storage device 104, which may communicate with each other via wired or wireless communication. As previously described, data storage device 104 may include computer-readable program instructions 106 and data 107. Data 107 may be any type of data, such as configuration data, sensor data, or diagnostic data, and other possibilities.

[0038] The controller 108 may include one or more circuits, digital logic units, computer chips, or microprocessors configured (and may in other tasks) to interface with any combination of mechanical components 110, one or more sensors 112, one or more power sources 114, electrical components 116, control system 118, or the user of the robot system 100. In some embodiments, the controller 108 may be a specially designed embedded device for performing specific operations in conjunction with one or more subsystems of the robot system 100.

[0039] The control system 118 can monitor and physically change the operating conditions of the robot system 100. In doing so, the control system 118 can act as a link between various parts of the robot system 100, such as between mechanical components 110 or electrical components 116. In some cases, the control system 118 can act as an interface between the robot system 100 and another computing device. Furthermore, the control system 118 can act as an interface between the robot system 100 and a user. In some cases, the control system 118 may include various components for communicating with the robot system 100, including joysticks, buttons, or ports. The example interfaces and communications mentioned above can be implemented via wired or wireless connections or both. The control system 118 can also perform other operations for the robot system 100.

[0040] During operation, the control system 118 can communicate with other systems of the robot system 100 via wired or wireless connections, and can be further configured to communicate with one or more users of the robot. As one possible illustration, the control system 118 can receive input (e.g., from a user or from another robot) instructing the robot system 100 to perform a requested task (e.g., picking up an object and moving it from one location to another). Based on this input, the control system 118 can perform actions to cause the robot system 100 to make a series of movements to perform the requested task. As another illustration, the control system can receive input instructing the robot system 100 to move to a requested location. In response, the control system 118 (with the assistance of other components or systems) can determine direction and speed to move the robot system 100 through the environment en route to the requested location.

[0041] The operation of the control system 118 may be performed by one or more processors 102. Alternatively, these operations may be performed by one or more controllers 108 or a combination of one or more processors 102 and one or more controllers 108. In some embodiments, the control system 118 may reside partially or entirely on a device other than the robot system 100, and thus may at least partially remotely control the robot system 100.

[0042] Mechanical component 110 represents the hardware of robot system 100 that enables robot system 100 to perform physical operations. As some examples, robot system 100 may include one or more physical components, such as an arm, end effector, head, neck, torso, base, and wheels. The physical components or other parts of robot system 100 may further include actuators arranged to move the physical components relative to each other. Robot system 100 may also include one or more structured bodies for housing control system 118 or other components, and may further include other types of mechanical components. The specific mechanical component 110 used in a given robot may vary based on the robot's design and may also be based on the operations or tasks the robot can be configured to perform.

[0043] In some examples, mechanical component 110 may include one or more removable components. Robotic system 100 may be configured to add or remove such removable components, which may involve assistance from a user or another robot. For example, robotic system 100 may be configured with removable end effectors or fingers that can be replaced or changed as needed or desired. In some embodiments, robotic system 100 may include one or more removable or replaceable battery cells, control systems, power systems, shock absorbers, or sensors. Other types of removable components may be included in some embodiments.

[0044] Robot system 100 may include one or more sensors 112 arranged to sense various aspects of robot system 100. The sensors 112 may include one or more force sensors, torque sensors, velocity sensors, acceleration sensors, position sensors, proximity sensors, motion sensors, positioning sensors, load sensors, temperature sensors, touch sensors, depth sensors, ultrasonic range sensors, infrared sensors, object sensors, or cameras, etc. In some examples, robot system 100 may be configured to receive sensor data from sensors physically separated from the robot (e.g., sensors located on other robots or located within the environment in which the robot is operating).

[0045] One or more sensors 112 can provide sensor data to one or more processors 102 (via data 107) to allow the robot system 100 to interact with its environment and to allow monitoring of the operation of the robot system 100. The sensor data can be used to evaluate various factors for starting, moving, and deactivating mechanical components 110 and electrical components 116 via the control system 118. For example, one or more sensors 112 can capture data corresponding to the location of environmental terrain or nearby objects, which can aid in environmental identification and navigation.

[0046] In some examples, sensor(s)112 may include a RADAR (e.g., for long-range object detection, distance determination, or velocity determination), a LIDAR (e.g., for short-range object detection, distance determination, or velocity determination), or a SONAR (e.g., for underwater object detection, distance determination, or velocity determination). (For example, for motion capture), one or more cameras (e.g., stereo cameras for 3D vision), a Global Positioning System (GPS) transceiver, or other sensors for capturing information about the environment in which the robotic system 100 is operating. The sensors 112 can monitor the environment in real time and detect obstacles, terrain elements, weather conditions, temperature, or other aspects of the environment. In another example, the sensors 112 can capture data corresponding to one or more characteristics of a target or identified object (e.g., the object's size, shape, outline, structure, or orientation).

[0047] Furthermore, the robot system 100 may include one or more sensors 112 configured to receive information indicating the state of the robot system 100, including one or more sensors 112 capable of monitoring the state of various components of the robot system 100. The one or more sensors 112 may measure the activity of the robot system 100 and receive information based on the operation of various characteristics of the robot system 100 (e.g., the operation of an extendable arm, end effector, or other mechanical or electrical characteristics of the robot system 100). The data provided by the one or more sensors 112 enables the control system 118 to determine operational errors and monitor the overall operation of the components of the robot system 100.

[0048] As an example, robot system 100 may use force / torque sensors to measure loads on various components of robot system 100. In some embodiments, robot system 100 may include one or more force / torque sensors on an arm or end effector to measure loads on actuators that move one or more components of the arm or end effector. In some examples, robot system 100 may include force / torque sensors at or near the wrist or end effector, but not at or near other joints of the robot arm. In further examples, robot system 100 may use one or more position sensors to sense the position of actuators in the robot system. For example, such position sensors may sense the extension, retraction, positioning, or rotation of actuators on the arm or end effector.

[0049] As another example, sensor(s)112 may include one or more velocity or acceleration sensors. For example, sensor(s)112 may include an inertial measurement unit (IMU). The IMU can sense velocity and acceleration relative to the gravity vector in the world coordinate system. The velocity and acceleration sensed by the IMU can then be converted into the velocity and acceleration of the robot system 100 based on the IMU's position in the robot system 100 and the kinematics of the robot system 100.

[0050] Robotic system 100 may include other types of sensors not explicitly discussed herein. Additionally or alternatively, the robotic system may use specific sensors for purposes not listed herein.

[0051] The robot system 100 may also include one or more power sources 114 configured to supply power to various components of the robot system 100. In other possible power systems, the robot system 100 may include hydraulic systems, electrical systems, batteries, or other types of power systems. As an example, the robot system 100 may include one or more batteries configured to supply power to its components. Some of the mechanical components 110 or electrical components 116 may be connected to different power sources, may be powered by the same power source, or may be powered by multiple power sources.

[0052] Any type of power source can be used to power the robot system 100, such as electricity or a gasoline engine. Additionally or alternatively, the robot system 100 may include a hydraulic system configured to power mechanical components 110 using fluid power. For example, components of the robot system 100 may operate based on hydraulic fluid transmitted through the hydraulic system to various hydraulic motors and cylinders. The hydraulic system may transmit hydraulic power via pressurized hydraulic fluid through pipes, flexible hoses, or other linkages between components of the robot system 100. One or more power sources 114 may be charged using various types of charging, such as wired connection to an external power source, wireless charging, combustion, or other examples.

[0053] Electrical component 116 may include various mechanisms capable of handling, transmitting, or providing electrical charge or signals. In possible examples, electrical component 116 may include wires, circuits, or wireless communication transmitters and receivers to enable operation of robot system 100. Electrical component 116 may interact with mechanical component 110 to enable robot system 100 to perform various operations. For example, electrical component 116 may be configured to provide power from power source(s) 114 to various mechanical components 110. Furthermore, robot system 100 may include an electric motor. Other examples of electrical component 116 may also be present.

[0054] Robot system 100 may include a body that can be attached to or house attachments and components of the robot system. Therefore, the structure of the body can vary in examples and can further depend on the specific operations that a given robot may have been designed to perform. For example, a robot developed to carry heavy loads may have a wide body capable of holding the load. Similarly, a robot designed to operate in confined spaces may have a relatively tall and narrow body. Furthermore, various types of materials, such as metal or plastic, can be used to develop the body or other components. In other examples, the robot may have a body with different structures or made of various types of materials.

[0055] The body or other components may include or carry one or more sensors 112. These sensors may be located at different sites on the robot system 100, such as the body, head, neck, base, torso, arm, or end effector, among other examples.

[0056] Robot system 100 can be configured to carry a load, such as a type of cargo to be transported. In some examples, the load may be placed by robot system 100 into a box or other container attached to robot system 100. The load may also represent an external battery or other type of power source (e.g., solar panel) that robot system 100 can utilize. Carrying a load indicates one example use for which robot system 100 can be configured, but robot system 100 may also be configured to perform other operations.

[0057] As described above, robot system 100 may include various types of attachments, wheels, end effectors, gripping devices, and the like. In some examples, robot system 100 may include a movable base with wheels, pedals, or some other form of displacement. Additionally, robot system 100 may include a robot arm or some other form of robot manipulator. In the case of a movable base, the base can be considered one of the mechanical components 110 and may include wheels powered by one or more actuators, which also allows movement of the robot arm in addition to the rest of the body.

[0058] Figure 2 The illustration depicts a mobile robot according to an example embodiment. Figure 3 An exploded view of a mobile robot according to an example embodiment is illustrated. More specifically, robot 200 may include a mobile base 202, a midsection 204, an arm 206, an end-of-arm system (EOAS) 208, a mast 210, a sensing housing 212, and a sensing kit 214. Robot 200 may also include a computing box 216 stored within the mobile base 202.

[0059] The movable base 202 includes two drive wheels positioned at the front end of the robot 200 to provide displacement for the robot 200. The movable base 202 also includes additional casters (not shown) to facilitate movement of the movable base 202 on the ground. The movable base 202 may have a modular architecture that allows easy removal of the computing box 216. The computing box 216 can serve as a removable control system for the robot 200 (rather than a mechanically integrated control system). After removing the outer housing, the computing box 216 can be easily removed and / or replaced. The movable base 202 may also be designed to allow for additional modularity. For example, the movable base 202 may also be designed to allow for easy removal and / or replacement of the power system, battery, and / or external shock absorbers.

[0060] The mid-section 204 can be attached to the front end of the movable base 202. The mid-section 204 includes a mounting post fixed to the movable base 202. The mid-section 204 also includes a rotational joint for the arm 206. More specifically, the mid-section 204 includes the first two degrees of freedom of the arm 206 (shoulder yaw J0 joint and shoulder pitch J1 joint). The mounting post and the shoulder yaw J0 joint can form part of a stacking tower at the front of the movable base 202. The mounting post and the shoulder yaw J0 joint can be coaxial. The length of the mounting post of the mid-section 204 can be selected to provide sufficient height for the arm 206 to perform manipulation tasks at commonly encountered height levels (e.g., coffee table level and countertop level). The length of the mounting post of the mid-section 204 can also allow the shoulder pitch J1 joint to rotate the arm 206 on the movable base 202 without contacting the movable base 202.

[0061] When connected to segment 204, arm 206 can be a 7DOF robotic arm. As described above, the first two DOFs of arm 206 can be included in segment 204. The remaining five DOFs can be included in separate segments of arm 206, such as... Figure 2 and Figure 3 As shown. Arm 206 can be made of a single plastic linkage structure. The arm 206 can accommodate a separate actuator module, a local motor driver, and through-hole cable wiring.

[0062] EOAS208 can be an end effector at the end of arm 206. EOAS208 allows robot 200 to manipulate objects in the environment. Figure 2 and Figure 3 As shown, EOAS208 can be a gripper, such as an underactuated clamping gripper. The gripper may include one or more contact sensors (e.g., force / torque sensors) and / or non-contact sensors (e.g., one or more cameras) to facilitate object detection and gripper control. EOAS208 may also be a different type of gripper (e.g., a suction gripper) or a different type of tool (e.g., a drill bit or brush). EOAS208 may also be replaceable or include replaceable components, such as gripper fingers.

[0063] Mast 210 may be a relatively long and narrow component between the shoulder yaw J0 joint of arm 206 and sensing housing 212. Mast 210 may be part of a stacking tower in front of movable base 202. Mast 210 may be fixed relative to movable base 202. Mast 210 may be coaxial with midsection 204. The length of mast 210 may facilitate sensing of objects manipulated by sensing kit 214 by EOAS 208. Mast 210 may be of such length that the highest point of the biceps of arm 206 is approximately aligned with the top of mast 210 when the shoulder pitch J1 joint rotates vertically upward. Thus, the length of mast 210 may be sufficient to prevent collision between sensing housing 212 and arm 206 when the shoulder pitch J1 joint rotates vertically upward.

[0064] like Figure 2 and Figure 3 As shown, mast 210 may include a 3D lidar sensor configured to collect depth information about the environment. The 3D lidar sensor may be coupled to a detached portion of mast 210 and fixed at a downward angle. The lidar position can be optimized for positioning, navigation, and forward cliff detection.

[0065] The sensing housing 212 may include at least one sensor that constitutes the sensing kit 214. The sensing housing 212 may be connected to translation / tilt controls to allow reorientation of the sensing housing 212 (e.g., to view an object manipulated by the EOAS 208). The sensing housing 212 may be part of a stacked tower fixed to a movable base 202. The rear of the sensing housing 212 may be coaxial with the mast 210.

[0066] The perception kit 214 may include a sensor suite configured to collect sensor data representing the environment of the robot 200. The perception kit 214 may include an infrared (IR)-assisted stereo depth sensor. The perception kit 214 may additionally include a wide-angle red-green-blue (RGB) camera for human-machine interaction and contextual information. The perception kit 214 may additionally include a high-resolution RGB camera for object classification. A facial halo may also be included around the perception kit 214 to improve human-machine interaction and scene lighting. In some examples, the perception kit 214 may also include a projector configured to project images and / or video into the environment.

[0067] Figure 4The illustration depicts a robotic arm according to an example embodiment. The robotic arm includes seven points of force (DOF): shoulder yaw J0 joint, shoulder pitch J1 joint, biceps roll J2 joint, elbow pitch J3 joint, forearm roll J4 joint, wrist pitch J5 joint, and wrist roll J6 joint. Each of these joints can be coupled to one or more actuators. The actuators coupled to the joints are operable to cause movement of links along the kinetic chain (and any end effectors attached to the robotic arm).

[0068] The shoulder yaw J0 joint allows the robotic arm to rotate towards the front and back of the robot. One useful application of this motion is allowing the robot to pick up an object in front of it and quickly place that object onto the rear portion of the robot (as well as reverse movement). Another useful application of this motion is for rapidly moving the robotic arm from a retracted configuration at the rear of the robot to an active position at the front (as well as reverse movement).

[0069] The shoulder pitch J1 joint allows the robot to raise its arm (e.g., so that the biceps are at the level of the robot's sensory kit) and lower its arm (e.g., so that the biceps are just above the movable base). This movement facilitates efficient manipulation of the robot at different target height levels in the environment (e.g., top gripping and side gripping). For example, the shoulder pitch J1 joint can rotate to a vertically upward position to allow the robot to easily manipulate objects on a table in the environment. The shoulder pitch J1 joint can rotate to a vertically downward position to allow the robot to easily manipulate objects on the ground in the environment.

[0070] The biceps rolling J2 joint allows the robot to rotate the biceps to move the elbow and forearm relative to the biceps. This movement can be particularly beneficial for clearer viewing of the EOAS (Electrical Orifice Assignment) via the robot's sensing suite. By rotating the biceps rolling J2 joint, the robot can kick out the elbow and forearm to improve its line of sight to the object held in the robot gripper.

[0071] Moving down the kinetic chain, alternating pitch and roll joints (shoulder pitch J1, biceps roll J2, elbow pitch J3, forearm roll J4, wrist pitch J5, and wrist roll J6) are provided to improve the maneuverability of the robotic arm. The axes of the wrist pitch J5, wrist roll J6, and forearm roll J4 joints intersect to reduce arm movement and thus reorient the object. A wrist roll J6 point is provided instead of the two pitch joints in the wrist to improve object rotation.

[0072] In some examples, such as Figure 4The robotic arm shown may be capable of operating in a teach mode. Specifically, a teach mode can be an operational mode of the robotic arm that allows the user to physically interact with the arm and guide it to perform and record various movements. In teach mode, external forces are applied to the robotic arm (e.g., by the user) based on teach inputs designed to teach the robot how to perform a specific task. The robotic arm can thus acquire data about how to perform the specific task based on instructions and guidance from the user. This data may relate to multiple configurations of mechanical components, joint position data, velocity data, acceleration data, torque data, force data, and power data, etc.

[0073] During teach mode, in some examples, the user can grasp the EOAS or wrist, or in others, any part of the robotic arm, and provide external force by physically moving the arm. Specifically, the user can guide the robotic arm to grasp an object and then move it from a first point to a second. While the user guides the robotic arm during teach mode, the robot can acquire and record data related to the movement, allowing the robotic arm to be configured to perform tasks independently at future times during independent operation (e.g., when the robotic arm operates independently outside of teach mode). In some examples, external force can also be applied by other entities in the physical workspace, such as other objects, machines, or robotic systems.

[0074] Figure 5 This is a side view of a robot capturing image data representing its environment, according to an example embodiment. More specifically, a robot 502, including a robot base 504, a robot arm 506, and a sensing housing 508, can operate in the environment. To control the robot 502, a robot control system can receive image data captured by one or more sensors in the sensing housing 508. To interpret the surrounding environment so that the robot 502 can perform tasks within it, the robot control system can apply one or more machine learning models to the captured image data. In some examples, the image data can be red-green-blue depth (RGBD) data, which includes both color and depth information. RGBD data can be generated using one or more separate sensors. In further examples, one or more machine learning models can be applied to red-green-blue (RGB) color images. In still other examples, one or more machine learning models can be applied to three-dimensional grayscale depth data. One or more machine learning models can also, or alternatively, be applied to other types of sensor data representing the environment.

[0075] In some examples, robot 502 can be controlled to interact with one or more features (e.g., objects or surfaces) in the surrounding environment. For instance, robot 502 can be controlled to interact with the environment using the end effector of robot arm 506, such as a gripper. (Reference) Figure 5 The environment includes a table 510, which contains a mug 512 and liquid spillage 514. Robot 502 can capture image data including the table 510 and its contents. Robot 502 can then attempt to process the image data to perform one or more robotic tasks. Different robotic tasks may require different information about features in the surrounding environment, such as objects or surfaces.

[0076] As an example, robot 502 can perform the task of identifying tidy areas in the environment. A machine learning model can be used to input image data as robot 502 and output tidy areas. For example, the model can identify liquid spill 514 as a tidy area. If mug 512 is empty, the model can also identify mug 512 as a tidy part of the environment that should be cleaned up. Given that the areas robot 502 should tidy may depend on user preferences, the machine learning model can be used to help robot 502 learn which parts of the environment to tidy up. User feedback (whether in a simulation or in the physical world) can be used to help train the model.

[0077] As another example, robot 502 could alternatively perform the task of tidying up objects in the environment (e.g., mug 512). To perform this task, robot 502 might first need to identify graspable areas in the environment, such as the handle of mug 512. Which objects or parts of objects the gripper of robot 502 can grasp might not be immediately obvious. Furthermore, there is a risk of damaging the objects if they are grasped in a suboptimal manner. Therefore, it might be advantageous to apply a machine learning model that processes image data of the environment and identifies the graspable areas of robot 502. In this example, identifying the graspable areas of objects in the environment can be considered a robot vision task corresponding to the robot's task of picking up and tidying up objects in the environment.

[0078] As a further example, robot 502 could alternatively perform the task of identifying wipeable areas in the environment to be wiped clean with different end effectors (e.g., a duster instead of a gripper). Which surfaces robot 502 can or should wipe clean may not be very obvious. For example, which surfaces are relatively permanent and therefore likely to collect dust may not be apparent. Therefore, applying machine learning models to identify wipeable surfaces in the environment can be advantageous in helping to control robot 502 to perform tasks involving dusting or otherwise cleaning surfaces.

[0079] Illustrative robotic tasks such as tidying up objects and wiping surfaces may each require an understanding of the environment, including the location and type of features (e.g., objects and surfaces). To master providing useful task-specific outputs to control robot 502, it may be necessary to train a machine learning model on numerous images captured by robot 502 or similar robots. During this training process, the model can be trained to determine a robust set of features from the input images. These features can then be used to generate task-specific outputs. In the example described herein, these features, along with the training work that allows the determination of such features, can be leveraged to allow robot 502 to efficiently obtain task-specific outputs for other robotic vision tasks. For example, if the model is trained to allow robot 502 to accurately perform tidying operations in its environment, the trained model can be combined with a task-specific head to similarly produce task-specific outputs, thereby assisting robot 502 in performing other tasks, such as tidying up objects, wiping surfaces, and other robotic tasks involving logically different perception tasks.

[0080] Figures 6A-6C Together, they illustrate the process of training and applying an FPN with a task-specific robot head. In the example shown, the shared dense network takes the form of an FPN, which can be particularly advantageous for processing features of different sizes at different resolutions. In other examples, different types of dense networks, such as convolutional neural networks (CNNs), can be used. A shared network can be called a dense network to indicate that the network has more layers than the task-specific head. In some examples, such as the one shown, the dense shared network can be configured to generate feature values ​​in the form of one or more feature maps corresponding to the captured image data. In other examples, the feature values ​​can take different forms, such as values ​​associated with nodes in the hidden layers of a neural network.

[0081] refer to Figure 6A As shown in information flow 600, the FPN 602 can initially be trained using a seed robot task. To train the FPN 602, training data can be accumulated, which includes image data of the environment and corresponding task-specific outputs for the seed task. For example, in some examples, the seed task may involve identifying manageable regions of the environment, in which case the task-specific output can indicate whether a region captured in a particular image is manageable. Other types of robot perception tasks can also be used, or alternatively, as seed tasks. In general, robot perception tasks available for their diverse image data and corresponding desired outputs are likely to be effective for the initial training of the FPN 602 to develop a rich feature set, which can then be used by other task-specific heads.

[0082] Figure 6AThe illustration shows the training of the FPN 602 and the robot seed task head 622 as part of information flow 600 using training image 604 and corresponding training output 624. Training image 604 can be captured or determined using one or more sensors on one or more robots (e.g., RGB cameras and / or stereo cameras). In some examples, training image 604 may be an RGBD image. Training output 624 can be determined based on labels provided by a human user in the physical world and / or simulation. In other examples, training output 624 may be determined by the robot experimentally. Some or all of the training data may also be provided by an external source (e.g., using a publicly accessible image dataset).

[0083] FPN 602 is a feature extractor designed to detect features at different scales, which can be useful for a variety of robotic tasks involving robot interaction with its environment. FPN 602 can be configured to compute convolutional feature maps at different resolutions for a single input image. The fully convolutional nature of FPN 602 allows the network to capture images of arbitrary sizes and output scaled feature maps at multiple levels of the feature pyramid. Higher-level feature maps contain grid cells covering larger areas of the image and are therefore better suited for detecting larger features. Furthermore, grid cells from lower-level feature maps may be more effective for detecting smaller features.

[0084] FPN 602 can be configured to generate corresponding feature maps 612, 614, and 616 for each training image 604. Any number of feature maps of different resolutions can be generated. Furthermore, any or all feature maps 612, 614, and 616 can be used individually or in combination to allow the robot seed task head 622 to make predictions about the environment. The robot seed task head 622 can contain additional layers that allow the generation of task-specific outputs for one or more input feature maps. In some examples, the robot seed task head 622 can be a CNN with fewer convolutional layers than FPN 602. During training of FPN 602, both FPN 602 and robot seed task head 622 can be trained simultaneously to map training image 604 to corresponding outputs of training output 624. Figure 6A During the training phase shown, FPN 602 can be trained to generate rich feature maps 612, 614 and 616, which can be used by robot seed task head 622 and other robot task heads that require relatively additional work.

[0085] Figure 6BThe illustration shows the training of a specific robot task head after FPN 602 has been trained. In some examples, workflow 630 can be executed on demand to quickly train a new robot task head 652, enabling the robot to quickly learn how to perform new tasks in the environment. Additional training data, in the form of training image 634 and training output 654, can be used to train robot task head 652. Training output 654 can be associated with a robot task different from the seed task (e.g., identifying graspable regions of the environment instead of tidyable regions). During workflow 630, FPN 602 can be fixed such that training image 634 and training output 654 do not change how FPN 602 generates feature maps 642, 644, and 646. Therefore, only the layers of robot task head 652 can be adjusted to better map training image 634 to training output 654 based on feature maps 642, 644, and 646 generated by FPN 602. Because FPN 602 may have been trained to generate rich feature maps, the robot task head 652 may require relatively few layers to produce accurate results, especially when the robot task head 652 is associated with a robot vision task similar to the initial seed task used to train FPN 602. Figure 6B The result of the workflow 630 shown can be that the robot task head 652 is effectively trained to operate on the feature maps from the FPN 602, thereby producing task-specific outputs that are useful for the robot to perform a different task than the one originally used to train the FPN 602.

[0086] Figure 6C The illustration depicts the runtime application of multiple robot task heads according to an example embodiment. More specifically, after FPN 602 and one or more task-specific heads have been trained, FPN 602 and the one or more task-specific heads can be applied by the robot at runtime to generate task-specific outputs that can be used to control the robot. In particular, the robot can capture an input image 664 and feed the input image 664 to FPN 602 to generate feature maps 672, 674, and 676. Feature maps 672, 674, and 676 can be input into multiple different robot task heads, including robot task head 652 that produces task-specific output 684 and robot task head 692 that produces task-specific output 694. Notably, for two robot tasks associated with robot task head 652 and robot task head 692, computational benefits can be obtained by using separate networks because FPN 602 only needs to be applied to the input image 664 by the robot at runtime. Then, robot task heads 652 and 692, each containing relatively few layers, can be applied to feature maps 672, 674 and 676 generated by FPN 602, respectively.

[0087] In further examples, the robot can use more than two robot task heads simultaneously. In some examples, the active (and therefore applied to) specific robot task heads from shared dense networks can be dynamically adjusted by the robot control system and / or by a separate user control system. In further examples, a programming interface can be provided to developers that allows for the dynamic activation and deactivation of specific robot task heads to conserve resources. Although Figure 6C It is not explicitly stated, but one or more robot task heads can be dynamically retrained, for example, based on the effectiveness of the robot in performing the desired task.

[0088] Figure 7 This is a system design diagram based on an example embodiment. More specifically, the system layout 700 of the software modules can be deployed on a robotic device to allow for the efficient generation of different types of task-specific outputs. FPN 702 can be a shared dense network that is trained relatively infrequently (e.g., weekly, monthly, or quarterly). In some examples, FPN 702 can be periodically replaced with different models (e.g., more complex versions) with only modifications to one or more initial layers of the different models for retraining. In further examples, different types of shared dense networks may also be used.

[0089] System arrangement 700 includes several individual robot task-specific heads, including detection 712, segmentation 714, classification 716, embedding 718, handheld 720, additional task #1 722, and additional task #2 724. Arrangement 700 is provided for illustrative purposes. In other examples, more or fewer task-specific heads may be used in robot vision systems with shared dense networks, including different combinations of task-specific heads. Locking the FPN 702 allows for frequent training (e.g., daily retraining) of the task-specific heads without having to coordinate with other task-specific heads. Furthermore, the task-specific heads can remain relatively small and execute quickly, so that the FPN 702 performs most of the heavy lifting only once per input image.

[0090] In further runtime enhancements, all layers of the FPN 702 can be retained on the graphics processing unit (GPU) of a computing device such as a robot control system. This arrangement avoids almost all the duplication required from the GPU to the central processing unit (CPU). In contrast, such as Figure 7 The task-specific header layer shown can be retained on the CPU.

[0091] Return to reference Figure 7Detection 712 refers to a task-specific head that generates bounding boxes around features that may be relevant to the robotic task. Segmentation 714 refers to a pixel-level instance mask for each bounding box associated with a feature. Notably, in addition to the output of FPN 702, Segmentation 714 can also take the output of Detection 712 as input. Other types of hierarchical arrangements where the task-specific head takes the output from different task-specific heads as input can also be used in various examples. Segmentation 714 can be done at the instance level (e.g., all pixels associated with a particular wall) or at the semantic level (e.g., all pixels associated with any wall in the environment). Classification 716 refers to associating features with a specific category among several possible categories for which the system was trained. Embedding 718 refers to category-related or instance-related information that can be used to help identify or interpret features.

[0092] Additionally, 720 in the head refers to the head that identifies objects within the gripper of the robot. In further examples, the task-specific head may alternatively be configured to identify any object partially occluded by a portion of the robot. In still other examples, the task-specific head may alternatively be configured to identify objects partially occluded by a specific portion of the robot that is different from the robot's gripper.

[0093] Arrangement 700 may additionally include additional task-specific heads 722 and 724. In some examples, the additional task may involve identifying a specific type of object. One or more of the object types may be objects that the robot can manipulate to enter or leave an area. The robot can manipulate objects to open or close doors. For example, task-specific head 722 may involve identifying door handles, while task-specific head 724 may involve identifying elevator buttons. In such examples, each of task-specific heads 722 and 724 may, for example, obtain output from detection 712 in addition to FPN 702. Using separate robot task heads to identify relatively small objects that are particularly important to the robot's operation can be advantageous when a more general object detector may not be skilled enough to handle a wide variety of objects of different sizes and types.

[0094] In another example, the additional task-specific heads 722 and 724 can each refer to different ways in which the robot can manipulate its environment. For example, task-specific head 722 could refer to a graspable area of ​​the environment for the robot to identify, while task-specific head 724 could refer to a wipeable area of ​​the environment for the robot to identify. In such an example, each task-specific head 722 and 724 can be associated with a task involving a different end effector of the robot. For example, task-specific head 722 might relate to a gripper, while task-specific head 724 might relate to a dust collector. In yet another example, task-specific heads 722 and 724 can be associated with other ways in which the robot can manipulate its environment, such as by identifying pushable and pullable areas, respectively.

[0095] In some examples, the set of task-specific heads associated with the deployment 700 can be dynamically adjusted. For example, individual task-specific heads can be activated or deactivated, which can allow for resource conservation. In further examples, the individual task-specific heads can also, or alternatively, be dynamically trained by the robot.

[0096] Figure 8 This is a block diagram of a method according to an example embodiment. In some examples, Figure 8 Method 800 can be executed by a control system (e.g., control system 118 of robot system 100). In a further example, method 500 can be executed by one or more processors, such as processor(s) 102, to execute program instructions, such as program instructions 106, stored in a data storage device, such as data storage device 104. Execution of method 800 can involve robotic devices, such as those relating to… Figure 1-4 The illustrated and described robotic device. Other robotic devices may also be used to perform method 800. In further examples, some or all of the blocks of method 800 may be performed by a control system remote from the robotic device. In still other examples, different blocks of method 800 may be performed by different control systems located on and / or remote from the robotic device. In yet another example, different blocks of method 800 may be performed by separate robotic devices.

[0097] In box 810, method 800 includes receiving image data representing the environment of the robotic device from a camera on the robotic device. In some examples, the image data includes RGBD data. Image data can be received from one or more sensors. In further examples, the image data can be processed before being used as input to a machine learning model (e.g., by fusing sensor data from multiple different sensing sensors of different types).

[0098] In box 820, method 800 further includes applying a trained dense network to image data to generate a set of feature values. The trained dense network may have been trained to perform a first robot vision task. In some examples, the trained dense network may be an FPN. In further examples, the trained dense network has been trained using image data from one or more other robotic devices having the same or similar cameras as the robotic devices. In some examples, the first robot vision task may involve determining whether an area of ​​the environment is maneuverable by the robot. Other types of robot vision tasks may also be used, or alternatively, as seed tasks.

[0099] In box 830, method 800 further includes applying a trained task-specific head to the set of feature values ​​to generate a task-specific output, thereby completing a second robot vision task. The trained task-specific head may have been trained to perform the second robot vision task based on feature values ​​previously generated from a trained dense network. The second robot vision task may differ from the first robot vision task. In some examples, the trained dense network has more network layers than the trained task-specific head. In some examples, the trained task-specific head is applied to both the set of feature values ​​and different task-specific outputs from different task-specific heads.

[0100] In box 840, method 800 further relates to controlling a robotic device to operate in an environment based on task-specific outputs generated to complete a second robotic vision task. For example, the robotic device may be controlled to pick up or otherwise manipulate objects, enter or leave an area, or otherwise interact with the surrounding environment of the robotic device.

[0101] In another example, a first robot vision task involves determining whether a first type of robot manipulation is operable in the environment, and a second robot vision task involves determining whether a second type of robot manipulation is operable in the environment. In such an example, the first type of robot manipulation may involve a first robot manipulator, and the second type of robot manipulation may involve a second robot manipulator.

[0102] In a further example, the trained task-specific head is one of at least three trained task-specific heads corresponding to the respective functions of detection, segmentation, and classification. In another example, the second robot vision task for the trained task-specific head involves determining whether an object is partially occluded by a portion of the robotic device. In a further example, the second robot vision task for the trained task-specific head involves determining whether an object is in the gripper of the robotic device.

[0103] In another example, a trained task-specific head is one of a plurality of trained task-specific heads that recognize a plurality of corresponding object types. The plurality of corresponding object types may include at least one object type that is maneuverable by the robot to enable the robotic device to enter or leave an area of ​​the environment. More specifically, the at least one object type may be maneuverable by the robot to open or close a door in the environment. For example, one object type associated with one task-specific head may be a door handle, while another object type associated with another task-specific head may be an elevator button.

[0104] In some examples, the control system of the robotic device can be configured to periodically adjust which of a plurality of task-specific heads is active. In other examples, the layers of the trained dense network are processed by the GPU of the robotic device, while the layers of the trained task-specific heads are processed by the CPU of the robotic device. The examples described herein may further involve periodically retraining the trained task-specific heads without altering the trained dense network.

[0105] III. Conclusion

[0106] This disclosure is not limiting in its description of the specific embodiments described herein, which are intended to be illustrative of various aspects. Many modifications and variations can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatus within the scope of this disclosure, other than those listed herein, will be apparent to those skilled in the art based on the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims.

[0107] The above detailed description, with reference to the accompanying drawings, illustrates various features and functions of the disclosed systems, devices, and methods. In the drawings, similar symbols generally identify similar parts unless the context otherwise requires. The exemplary embodiments described herein and in the figures are not intended to be limiting. Other embodiments may be used, and other changes may be made, without departing from the spirit or scope of the subject matter set forth herein. It will be readily understood that, as generally described herein and shown in the figures, aspects of this disclosure can be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are explicitly contemplated herein.

[0108] A box representing information processing may correspond to a circuit that can be configured to perform a specific logical function of the method or technique described herein. Alternatively or additionally, a box representing information processing may correspond to a module, segment, or portion of program code (including associated data). The program code may include one or more processor-executable instructions for implementing a specific logical function or action in the method or technique. The program code or associated data may be stored on any type of computer-readable medium, such as storage devices including disks or hard disk drives, or other storage media.

[0109] Computer-readable media can also include non-transitory computer-readable media, such as computer-readable media that store data for a short period of time, such as register memory, processor cache, and random access memory (RAM). Computer-readable media can also include non-transitory computer-readable media that store program code or data for a longer period of time, such as secondary or perpetual long-term storage, such as read-only memory (ROM), optical discs or magnetic disks, and optical disc read-only memory (CD-ROM), for example. Computer-readable media can also be any other volatile or non-volatile storage system. For example, a computer-readable medium can be considered a computer-readable storage medium or a tangible storage device.

[0110] Furthermore, a frame representing one or more information transfers can correspond to information transfers between software modules or hardware modules within the same physical device. However, other information transfers can occur between software modules or hardware modules in different physical devices.

[0111] The specific arrangements shown in the figures should not be considered limiting. It should be understood that other embodiments may include more or fewer of each element shown in the given figures. Furthermore, some illustrated elements may be combined or omitted. Additionally, example embodiments may include elements not shown in the figures.

[0112] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The aspects and embodiments disclosed herein are for illustrative purposes and are not intended to be limiting; the true scope is indicated by the appended claims.

[0113] Attached Figure

[0114] Figure 1

[0115] 100 Robot Systems

[0116] 110 Mechanical parts

[0117] 114 (one or more) power sources

[0118] 116 Electrical components

[0119] 112 (one or more) sensors

[0120] 118 Control System

[0121] 102 (one or more) processors

[0122] 104 Data storage devices

[0123] 106 Program Instructions

[0124] 107 data

[0125] 108 (one or more) controllers

[0126] Figure 6A

[0127] 604 training images

[0128] 602 Feature Pyramid Network (FPN)

[0129] 612 Feature Map

[0130] 614 Feature Map

[0131] 616 Feature Map

[0132] 622 Robot Seed Mission Head

[0133] 624 Training Output

[0134] Training

[0135] Figure 6B

[0136] 634 training images

[0137] 602 Feature Pyramid Network (FPN)

[0138] 642 Feature Map

[0139] 644 Feature Map

[0140] 646 Feature Map

[0141] 652 Robot Task Head

[0142] 654 Training Output

[0143] Fixed

[0144] Training

[0145] Figure 6C

[0146] 664 Input Images

[0147] 602 Feature Pyramid Network (FPN)

[0148] 672 Feature Map

[0149] 674 Feature Map

[0150] 676 Feature Map

[0151] 652 Robot Task Head

[0152] 692 Robot Task Head

[0153] 684 Task-Specific Output

[0154] 694 Task-Specific Output

[0155] Fixed

[0156] Figure 7

[0157] Feature Pyramid Network 702

[0158] Trained Rarely

[0159] Trained frequently

[0160] Detection 712

[0161] Split 714

[0162] Category 716

[0163] Embedded 718

[0164] 720 in hand

[0165] Additional Task #1 722

[0166] Additional Task #2 724

[0167] Figure 8

[0168] 810 receives image data representing the environment of the robot from a camera on the robot device.

[0169] The 820 applies a trained dense network to image data to generate a set of feature values, where the trained dense network has been trained to complete the first robot vision task.

[0170] 830 applies a trained task-specific head to the set of feature values ​​to generate a task-specific output to complete a second robot vision task, wherein the trained task-specific head has been trained to perform the second robot vision task based on feature values ​​previously generated by a trained dense network, and wherein the second robot vision task is different from the first robot vision task.

[0171] 840 controls the robot device to operate in the environment based on task-specific outputs generated to complete the second robot vision task.

Claims

1. A method for a robotic device, comprising: Receive image data representing the environment of the robot from the camera on the robot device; A trained dense network is applied to the image data to generate a set of feature values, wherein the trained dense network and a first trained task-specific head have been trained simultaneously and in combination to perform a first robot vision task. A second trained task-specific head is applied to the set of feature values ​​to generate a task-specific output to complete a second robot vision task. The second trained task-specific head has been trained to complete the second robot vision task based on feature values ​​previously generated by the trained dense network. The second trained task-specific head is trained to complete the second robot vision task after the trained dense network and the first robot vision task have been trained simultaneously and in combination to complete the first robot vision task. Each of the first and second robot vision tasks involves processing image data to obtain specific information required to complete different corresponding robot tasks, where each robot task involves the robot device interacting with objects in the environment and changing the state of those objects. as well as The robot device is controlled to operate in the environment based on the task-specific output generated to complete the second robot vision task.

2. The method according to claim 1, wherein the trained dense network is a Feature Pyramid Network (FPN).

3. The method of claim 1, further comprising periodically retraining the first or second trained task-specific head without altering the trained dense network.

4. The method of claim 1, wherein the trained dense network has more network layers than each of the first trained task-specific head and the second trained task-specific head.

5. The method of claim 1, wherein the trained dense network has been trained using image data from one or more other robotic devices having the same or similar cameras as the robotic device.

6. The method of claim 1, wherein the first robot vision task involves determining whether a region is maneuverable by the robot.

7. The method of claim 1, wherein the first robot vision task relates to determining whether a first type of robot manipulation is operable in the environment, and the second robot vision task relates to determining whether a second type of robot manipulation is operable in the environment.

8. The method of claim 7, wherein the first type of robot manipulation involves a first robot manipulator, and the second type of robot manipulation involves a second robot manipulator.

9. The method of claim 1, wherein the first or second trained task-specific head is one of at least three trained task-specific heads corresponding to the respective functions of detection, segmentation, and classification.

10. The method of claim 1, wherein the second robot vision task for the second trained task-specific head involves determining whether an object is partially occluded by a portion of the robotic device.

11. The method of claim 1, wherein the second robot vision task for the second trained task-specific head involves determining whether an object is in the gripper of the robot device.

12. The method of claim 1, wherein the first or second trained task-specific head is one of a plurality of trained task-specific heads corresponding to the identification of a plurality of corresponding object types.

13. The method of claim 12, wherein the plurality of corresponding object types includes at least one object type that is robot-manipulable, enabling the robotic device to enter or leave an area of ​​the environment.

14. The method of claim 13, wherein the at least one object type is a door that a robot can manipulate to open or close in the environment.

15. The method of claim 1, wherein the control system of the robotic device comprises a plurality of task-specific heads, and wherein the method further comprises periodically adjusting which of the plurality of task-specific heads is active.

16. The method of claim 1, wherein the second trained task-specific head is applied to both the set of feature values ​​and different task-specific outputs from different task-specific heads.

17. The method of claim 1, wherein the image data comprises red-green-blue depth (RGBD) data.

18. The method of claim 1, wherein the layers of the trained dense network are processed by the graphics processing unit (GPU) of the robotic device, and wherein the layers of the first trained task-specific head or the second trained task-specific head are processed by the central processing unit (CPU) of the robotic device.

19. A robotic device, comprising: camera; and The control system is configured as follows: Receive image data representing the environment of the robot from the camera on the robot device; A trained dense network is applied to the image data to generate a set of feature values, wherein the trained dense network and a first trained task-specific head have been trained simultaneously and in combination to perform a first robot vision task. A second trained task-specific head is applied to the set of feature values ​​to generate a task-specific output to complete a second robot vision task, wherein the second trained task-specific head has been trained to complete the second robot vision task based on feature values ​​previously generated by the trained dense network, wherein the second trained task-specific head is trained to complete the second robot vision task after the trained dense network and the first robot vision task are trained simultaneously and in combination to complete the first robot vision task, wherein each of the first robot vision task and the second robot vision task involves processing image data to obtain specific information required to complete different corresponding robot tasks, wherein each robot task involves the robot device interacting with objects in the environment and changing the state of the objects; as well as The robot device is controlled to operate in the environment based on task-specific outputs generated to complete the second robot vision task.

20. A non-transitory computer-readable medium comprising program instructions executable by at least one processor to cause the at least one processor to perform operations, said operations including: Receive image data representing the environment of the robot from the camera on the robot device; A trained dense network is applied to the image data to generate a set of feature values, wherein the trained dense network and a first trained task-specific head have been trained simultaneously and in combination to perform a first robot vision task. A second trained task-specific head is applied to the set of feature values ​​to generate a task-specific output to complete a second robot vision task, wherein the second trained task-specific head has been trained to complete the second robot vision task based on feature values ​​previously generated by the trained dense network, wherein the second trained task-specific head is trained to complete the second robot vision task after the trained dense network and the first robot vision task are trained simultaneously and in combination to complete the first robot vision task, wherein each of the first robot vision task and the second robot vision task involves processing image data to obtain specific information required to complete different corresponding robot tasks, wherein each robot task involves the robot device interacting with objects in the environment and changing the state of the objects; as well as The robot device is controlled to operate in the environment based on task-specific outputs generated to complete the second robot vision task.

Citation Information

Patent Citations

  • Shared Processing with Deep Neural Networks

    CN109426262A

  • System and method for real-time large image homography processing

    US20190147621A1

  • Learning robotic tasks using one or more neural networks

    US20190228495A1