System and method for robot system for handling object

JP2024019690A5Pending Publication Date: 2026-05-25MUJIN INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
MUJIN INC
Filing Date
2023-12-25
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

Robots often struggle to perform complex tasks that require sophisticated human-like interaction and motion, particularly in handling objects within containers where objects are irregularly arranged, leading to difficulties in detection, identification, and retrieval.

Method used

A computing system is provided that includes a control system to communicate with a robot equipped with a robotic arm and a camera, enabling the generation of arm and end effector device trajectories for object handling, allowing the robot to identify, approach, and grasp target objects within a container, even when objects are randomly located.

Benefits of technology

The system enhances the speed, precision, and accuracy of object detection, identification, and retrieval from containers by improving the robot's ability to handle irregularly arranged objects, facilitating more efficient and precise robotic interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a computing system that can equally improve detection, identification and extraction of objects which are arranged in a regular manner or in a semiregular manner.SOLUTION: A computing system includes an end effector device or includes a processing circuit that performs communication with a robot having a robot arm mounted on the device. The processing circuit identifies a target object out of a plurality of objects in an object supply source, determines an approach-track for the robot arm and the end effector device to approach the plurality of objects, determines gripping motion for the end effector device to grip the target object, and controls the robot arm and the end effector device, so that the target object is extracted through the determined track. The processing circuit determines an approach-track to a destination, and controls the robot arm and the end effector device gripping the target object, so that the arm and the device approach the destination and release the target object in the destination.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 317,558, filed March 8, 2022, entitled "ROBOTIC SYSTEM WITH OBJECT HANDLING," the entire contents of which are incorporated herein by reference.

[0002] The present technology is directed generally to robotic systems, and more particularly to systems, processes, and techniques for detecting and handling objects. More particularly, the present technology can be used to detect and handle objects in a container. [Background technology]

[0003] With ever-increasing capabilities and decreasing costs, many robots (e.g., machines configured to automatically / autonomously perform physical actions) are now widely used in a variety of different fields. Robots can be used to perform a variety of tasks (e.g., manipulating or transporting objects through space), for example, in manufacturing and / or assembly, packing and / or packaging, transporting and / or shipping, etc. In performing a task, a robot can replicate the actions of a human, thereby replacing or reducing human involvement required to perform an otherwise dangerous or repetitive task.

[0004] However, despite advances in technology, robots often lack the sophistication necessary to replicate the human interactions required to perform larger and / or more complex tasks. Thus, a need remains for improved techniques and systems for managing the actions and / or interactions between robots. Summary of the Invention

[0005] In an embodiment, a computing system is provided that includes a control system configured to communicate with a robot having a robotic arm including or attached to an end effector device, and to communicate with a camera. The at least one processing circuit can be configured, when the robot is in an object handling environment including a source of objects for transfer to a destination within the object handling environment, to identify a target object from among a plurality of objects in the source of objects to transfer the target object from the source of objects to the destination; generate an arm approach trajectory for the robot arm to approach the plurality of objects; generate an end effector device approach trajectory for the end effector device to approach the target object; generate a grasping operation to grasp the target object with the end effector device; control the robot arm according to the arm approach trajectory to approach the plurality of objects; output an arm approach command to control the robot arm according to the end effector device approach trajectory to approach the target object; and output an end effector device control command to control the end effector device in a grasping operation to grasp the target object.

[0006] In another embodiment, a method for picking a target object from a source of objects is provided, the method including the steps of identifying the target object from among a plurality of objects in the source of objects, determining an arm approach trajectory for a robot arm having an end effector device to approach the plurality of objects, generating an end effector device approach trajectory for the end effector device to approach the target object, generating a grasping motion for grasping the target object with the end effector device, controlling the robot arm according to the arm approach trajectory to approach the plurality of objects, outputting an arm approach command for controlling the robot arm according to the end effector device approach trajectory to approach the target object, and outputting an end effector device control command for controlling the end effector device to grasp the object.

[0007] In another embodiment, a non-transitory computer-readable medium is provided that is operable by at least one processing circuit via a communication interface configured to communicate with a robotic system and has executable instructions for implementing a method for picking a target object from a source of objects, the method including: identifying the target object from among a plurality of objects in the source of objects, generating an arm approach trajectory for a robot arm having an end effector device to approach the plurality of objects, generating an end effector device approach trajectory for the end effector device to approach the target object, generating a grasping action to grasp the target object with the end effector device, outputting an arm approach command for controlling the robot arm according to the arm approach trajectory to approach the plurality of objects, outputting an end effector device approach command for controlling the robot arm according to the end effector device approach trajectory to approach the target object, and outputting an end effector device control command for controlling the end effector device to grasp the object. [Brief description of the drawings]

[0008] [Figure 1A] 1 illustrates a system for performing or facilitating object detection, identification, and retrieval, in accordance with embodiments herein. [Figure 1B] 1 illustrates an embodiment of a system for performing or facilitating object detection, identification, and retrieval in accordance with embodiments herein. [Figure 1C] 1 illustrates another embodiment of a system for performing or facilitating object detection, identification, and retrieval in accordance with embodiments herein. [Figure 1D] 1 illustrates yet another embodiment of a system for performing or facilitating object detection, identification, and retrieval in accordance with embodiments herein. [Figure 2A] FIG. 1 is a block diagram illustrating a computing system configured to perform or facilitate object detection, identification and retrieval consistent with embodiments herein. [Figure 2B] FIG. 1 is a block diagram illustrating one embodiment of a computing system configured to perform or facilitate object detection, identification, and retrieval consistent with embodiments herein. [Figure 2C] FIG. 1 is a block diagram illustrating another embodiment of a computing system configured to perform or facilitate object detection, identification, and retrieval consistent with embodiments herein. [Figure 2D] FIG. 1 is a block diagram illustrating yet another embodiment of a computing system configured to perform or facilitate object detection, identification, and retrieval consistent with embodiments herein. [Figure 2E] 1 is an example of image information processed by the system and consistent with embodiments herein. [Figure 2F] 1 is an example of image information processed by the system and consistent with embodiments herein. [Figure 3A] 1 illustrates an exemplary environment for operating a robotic system, according to embodiments herein. [Figure 3B]1 illustrates an exemplary environment for object detection, identification, and retrieval by a robotic system consistent with embodiments herein. [Figure 3C] 1 illustrates a robotic system having an arm, a base, and an end effector device. [Figure 3D] 1 illustrates another exemplary embodiment of a robotic system having an arm, a base, and an end effector device. [Figure 4] A flow diagram is provided illustrating the overall flow of methods and operations for detecting, planning, picking, transporting, and placing target objects according to embodiments herein. [Figure 5A] Illustrate the location of a container or source that contains multiple objects. [Figure 5B] 1 illustrates a visual depiction of the detection results described herein for multiple detected objects from multiple objects within a container or source location. [Figure 5C] 1 illustrates an example of object recognition from detection results consistent with embodiments herein. [Figure 6A] 1 illustrates various grasp models utilized by a robotic system to grasp an object. [Figure 6B] 1 illustrates various grasp models utilized by a robotic system to grasp an object. [Figure 6C] 1 illustrates various grasping models utilized by a robotic system to grasp an object. [Figure 7A] 1 illustrates a motion plan for a transfer cycle of an object by a robotic arm from a source to a destination. [Figure 7B] 1 illustrates an embodiment of a system and method for object handling via a robotic system as described herein. [Figure 7C] 1 illustrates a visual depiction of the detection results described herein for multiple detected objects from multiple objects within a container or source location, where primary and secondary objects are selected via operations further described herein. [Figure 7D]1 illustrates an example of the use of bounding boxes via a robotic system during a grasping operation, as described herein. [Figure 8A] 1 illustrates an end effector device grasp approach trajectory. [Figure 8B] 1 illustrates an object chucking operation. [Figure 8C] 1 illustrates an object grasp departure trajectory. [Figure 9A] 1 illustrates a second object grasping approach trajectory. [Figure 9B] 4 illustrates a second object chucking operation. [Figure 9C] 1 illustrates a second object grasping departure trajectory. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Systems and methods related to object detection, identification, and retrieval are described herein. In particular, the disclosed systems and methods can facilitate object detection, identification, and retrieval when the object is located in a container. As discussed herein, the object may be metal or other material and may be located in a source, including containers such as boxes, bins, crates, etc. Objects may be unorganized or randomly placed in a container, such as a box full of screws. While object detection and identification in such situations may be difficult due to the random placement of objects, the systems and methods discussed herein may equally improve object detection, identification, and object retrieval for regularly or semi-regularly placed objects. Thus, the systems and methods described herein are designed to identify individual objects among multiple objects, where the individual objects may be located in different locations, at different angles, etc. The systems and methods discussed herein may include robotic systems. A robotic system configured according to embodiments herein may autonomously perform integrated tasks by coordinating the operation of multiple robots. A robotic system may include any suitable combination of robotic devices, actuators, sensors, cameras, and computing systems configured to control, issue commands, receive information from the robotic devices and sensors, access, analyze, and process data generated by the robotic devices, sensors, and cameras, generate data or information usable for control of the robotic system, and plan actions for the robotic devices, sensors, and cameras, as described herein. As used herein, a robotic system does not necessarily have immediate access to or control of the robotic actuators, sensors, or other devices. A robotic system may be a computing system configured to improve the performance of such robotic actuators, sensors, and other devices through receiving, analyzing, and processing information, as described herein.

[0010] The technology described herein provides technical improvements to robotic systems configured for use in identifying, detecting, and retrieving objects. The technical improvements described herein increase the speed, precision, and accuracy of these tasks, making it easier to detect, identify, and retrieve objects from containers. The robotic and computing systems described herein address the technical problem of identifying, detecting, and retrieving objects from containers, where the objects may be irregularly positioned. Addressing this technical problem improves the technology of identifying, detecting, and retrieving objects.

[0011] This application refers to systems and robotic systems. A robotic system may include robotic actuator components (e.g., a robotic arm, a robotic gripper, etc.), various sensors (e.g., a camera, etc.), and various computing or control systems, as discussed herein. As discussed herein, a computing system or control system may be referred to as "controlling" the various robotic components, such as the robotic arm, the robotic gripper, the camera, etc. Such "control" may refer to the direct control and interaction of the various actuators, sensors, and other functional aspects of the robotic components. For example, a computing system may control a robotic arm by issuing or providing all of the necessary signals to the various motors, actuators, and sensors to cause robotic movements. Such "control" may also refer to the issuance of abstract or indirect commands to a further robotic control system that converts such commands into the necessary signals to cause robotic movements. For example, a computing system may control a robotic arm by issuing commands describing a trajectory or destination location along which the robotic arm should move, and a further robotic control system associated with the robotic arm may receive and interpret such commands and then provide the necessary direct signals to the various actuators and sensors of the robotic arm to cause the necessary movements.

[0012] Specifically, the technology described herein assists a robotic system in interacting with a target object among multiple objects in a container. Detection, identification, and removal of an object from a container requires several steps, including generating a suitable object recognition template, extracting features that can be used for identification, and generating, refining, and validating a detection hypothesis. For example, due to the possibility of irregular placement of objects, it may be necessary to recognize and identify objects in multiple different poses (e.g., angles and locations) and when potentially obscured by parts of other objects.

[0013] In the following, specific details are described to provide an understanding of the technology of the present disclosure. In an embodiment, the techniques introduced herein may be implemented without including each specific detail disclosed herein. In other instances, well-known features, such as specific functions or routines, are not described in detail to avoid unnecessarily obscuring the present disclosure. References in this description to "embodiment," "one embodiment," or the like, mean that the particular feature, structure, material, or characteristic described is included in at least one embodiment of the present disclosure. Thus, appearances of such phrases in this specification do not necessarily all refer to the same embodiment. On the other hand, such references are not necessarily mutually exclusive. Furthermore, a particular feature, structure, material, or characteristic described with respect to any one embodiment can be combined in any suitable manner with that of any other embodiment, unless such items are mutually exclusive. It should be understood that the various embodiments shown in the figures are merely illustrative representations and are not necessarily drawn to scale.

[0014] For purposes of clarity, certain details describing structures or processes that are well known and often associated with robotic systems and subsystems, but that may unnecessarily obscure certain important aspects of the techniques of the present disclosure, are not included in the following description. Furthermore, although the following disclosure describes several embodiments of different aspects of the present technology, several other embodiments may have different configurations or different components than those described in this section. Thus, the disclosed technology may have other embodiments that have additional elements or that do not have some of the elements described below.

[0015] Many embodiments or aspects of the disclosure described below may take the form of computer or controller executable instructions, including routines executed by a programmable computer or controller. Those skilled in the relevant art will appreciate that the disclosed techniques may be practiced on or with computer or controller systems other than those shown and described below. The techniques described herein may be embodied in a special purpose computer or data processor that is specially programmed, configured, or constructed to execute one or more of the computer executable instructions described below. Thus, the terms "computer" and "controller" as used generally herein refer to any data processor and may include Internet appliances and handheld devices, including palmtop computers, wearable computers, cellular or mobile phones, multiprocessor systems, processor-based or programmable consumer electronics, network computers, minicomputers, and the like. Information handled by these computers and controllers may be presented in any suitable display medium, including liquid crystal displays (LCDs). Instructions for performing computer or controller executable tasks may be stored in or on any suitable computer readable medium, including hardware, firmware, or a combination of hardware and firmware. The instructions may be contained in any suitable memory device, including, for example, a flash drive, a USB device, and / or other suitable medium.

[0016] The terms "coupled" and "connected," along with their derivatives, may be used herein to describe a structural relationship between components. It should be understood that these terms are not intended as synonyms for each other. Rather, in certain embodiments, "connected" may be used to indicate that two or more elements are in direct contact with each other. Unless otherwise clear in the context, the term "coupled" may be used to indicate that two or more elements are in direct or indirect contact with each other (with other intervening elements therebetween), or that two or more elements cooperate or interact with each other (e.g., in a causal relationship, such as for signal transmission / reception or for function calls), or both.

[0017] Any reference herein to image analysis by a computing system may be performed in accordance with or using spatial structure information, which may include depth information describing the respective depth values ​​of various locations relative to a selected point. The depth information may be used to identify an object or estimate how an object is spatially arranged. In some instances, the spatial structure information may include, or be used to generate, a point cloud describing the location of one or more surfaces of an object. Spatial structure information is only one form of possible image analysis, and other forms known to those skilled in the art may be used in accordance with the methods described herein.

[0018] FIG. 1A illustrates a system 1000 for performing object detection, or more specifically, object recognition. More specifically, the system 1000 may include a computing system 1100 and a camera 1200. In this example, the camera 1200 may be configured to generate image information that describes or otherwise represents an environment in which the camera 1200 is located, or more specifically, represents an environment in a field of view of the camera 1200 (also referred to as a camera field of view). The environment may be, for example, a warehouse, a manufacturing plant, a retail space, or other facility. In such an instance, the image information may represent objects located in such a facility, such as boxes, bins, cases, crates, pallets, or other containers. The system 1000 may be configured to generate, receive, and / or process the image information, such as using the image information to distinguish between individual objects in the camera field of view, to perform object recognition or object registration based on the image information, and / or to perform a robot interaction plan based on the image information, as discussed in more detail below (the terms "and / or" and "or" are used interchangeably in this disclosure). The robot interaction plan may be used to control a robot at a facility, for example, to facilitate robot interaction between the robot and a container or other object. The computing system 1100 and the camera 1200 may be located at the same facility or may be located remotely from one another. For example, the computing system 1100 may be part of a cloud computing platform hosted in a data center remote from a warehouse or retail space and may communicate with the camera 1200 via a network connection.

[0019] In an embodiment, the camera 1200 (which may also be referred to as an image sensing device) may be a 2D camera and / or a 3D camera. For example, FIG. 1B illustrates a computing system 1100 and a system 1500A (which may be an embodiment of the system 1000) including a camera 1200A and a camera 1200B, both of which may be an embodiment of the camera 1200. In this example, the camera 1200A may be a 2D camera configured to generate 2D image information that includes or forms a 2D image describing the visual appearance of an environment in the field of view of the camera. The camera 1200B may be a 3D camera (also referred to as a spatial structure sensing camera or a spatial structure sensing device) configured to generate 3D image information that includes or forms spatial structure information about the environment in the field of view of the camera. The spatial structure information may include depth information (e.g., a depth map) that describes the respective depth values ​​of various locations relative to the camera 1200B, such as locations on the surface of various objects in the field of view of the camera 1200B. These locations in the field of view of the camera or on the surface of the object may also be referred to as physical locations. The depth information of this example can be used to estimate how the object is spatially arranged in three-dimensional (3D) space. In some instances, the spatial structure information can include, or can be used to generate, a point cloud that describes locations on one or more surfaces of the object within the field of view of the camera 1200B. More specifically, the spatial structure information can describe various locations on the structure of the object (also referred to as the object structure).

[0020] In an embodiment, the system 1000 may be a robotic motion system for facilitating robotic interaction between a robot and various objects in the environment of the camera 1200. For example, FIG. 1C illustrates a robotic motion system 1500B, which may be an embodiment of the system 1000 / 1500A of FIGS. 1A and 1B. The robotic motion system 1500B may include a computing system 1100, a camera 1200, and a robot 1300. As described above, the robot 1300 may be used to interact with one or more objects, such as boxes, crates, bins, pallets, or other containers, in the environment of the camera 1200. For example, the robot 1300 may be configured to pick containers from one location and move them to another location. In some cases, the robot 1300 may be used to perform an unpalletizing operation, where a group of containers or other objects are unloaded and moved, for example, to a conveyor belt. In some implementations, the camera 1200 may be attached to the robot 1300 or the robot 3300, discussed below. This is also known as a camera handheld or handheld solution. The camera 1200 can be attached to the robot arm 3320 of the robot 1300. The robot arm 3320 can then move to various pick areas and generate image information for those areas. In some implementations, the camera 1200 can be separate from the robot 1300. For example, the camera 1200 can be mounted on the ceiling of a warehouse or other structure and can remain stationary relative to the structure. In some implementations, multiple cameras 1200 can be used, including multiple cameras 1200 separate from the robot 1300 and / or cameras 1200 separate from the robot 1300 used in combination with handheld cameras 1200. In some implementations, a camera 1200 or multiple cameras 1200 can be mounted or fixed to a dedicated robotic system separate from the robot 1300 used for object manipulation, such as a robotic arm, gantry, or other automated system configured for camera movement.Throughout this specification, we may talk about "controlling" or "controlling" the camera 1200. For a handheld camera solution, controlling the camera 1200 also includes controlling the robot 1300 to which the camera 1200 is attached or mounted.

[0021] In an embodiment, the computing system 1100 of FIGS. 1A-1C may form or be incorporated into the robot 1300, which may be referred to as a robot controller. A robot control system may be included in the system 1500B and configured to generate commands for the robot 1300, such as, for example, robot interaction movement commands for controlling a robot interaction between the robot 1300 and a container or other object. In such an embodiment, the computing system 1100 may be configured to generate such commands based on, for example, image information generated by the camera 1200. For example, the computing system 1100 may be configured to determine a motion plan based on the image information, where the motion plan may be, for example, intended to grip or otherwise grasp an object. The computing system 1100 may generate one or more robot interaction movement commands to execute the motion plan.

[0022] In an embodiment, the computing system 1100 may form or be part of a vision system. The vision system may be, for example, a system that generates visual information describing an environment in which the robot 1300 is located, or alternatively or additionally describing an environment in which the camera 1200 is located. The visual information may include 3D image information and / or 2D image information discussed above, or some other image information. In some scenarios, if the computing system 1100 forms a vision system, the vision system may be part of the robot control system discussed above, or may be separate from the robot control system. If the vision system is separate from the robot control system, the vision system may be configured to output information describing the environment in which the robot 1300 is located. The information may be output to the robot control system, which may receive such information from the vision system and perform motion planning based on the information and / or generate robot interaction motion commands. More information regarding the vision system is described in detail below.

[0023] In an embodiment, the computing system 1100 may communicate with the camera 1200 and / or the robot 1300 via a dedicated wired communication interface, such as an RS-232 interface, a universal serial bus (USB) interface, and / or via a direct connection, such as a connection provided via a local computer bus, such as a peripheral component interconnect (PCI) bus. In an embodiment, the computing system 1100 may communicate with the camera 1200 and / or with the robot 1300 via a network. The network may be of any type and / or form, such as a personal area network (PAN), a local area network (LAN), e.g., an intranet, a metropolitan area network (MAN), a wide area network (WAN), or the Internet. The network may utilize different techniques and layers or stacks of protocols, including, for example, Ethernet protocols, Internet Protocol Suite (TCP / IP), ATM (Asynchronous Transfer Mode) techniques, SONET (Synchronous Optical Network) protocols, or SDH (Synchronous Digital Hierarchy) protocols.

[0024] In an embodiment, the computing system 1100 may communicate information directly with the camera 1200 and / or with the robot 1300, or may communicate via an intermediate storage device, or more generally, an intermediate non-transitory computer-readable medium. For example, FIG. 1D illustrates a system 1500C, which may be an embodiment of the systems 1000 / 1500A / 1500B, that includes a non-transitory computer-readable medium 1400 that may be external to the computing system 1100 and may act as an external buffer or repository for storing image information generated by the camera 1200, for example. In such an embodiment, the computing system 1100 may retrieve or otherwise receive image information from the non-transitory computer-readable medium 1400. Examples of the non-transitory computer-readable medium 1400 include electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. The non-transitory computer readable medium may form, for example, a computer diskette, a hard disk drive (HDD), a solid state drive (SDD), a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read only memory (CD-ROM), a digital versatile disc (DVD), and / or a memory stick.

[0025] As described above, the camera 1200 may be a 3D camera and / or a 2D camera. The 2D camera may be configured to generate a 2D image, such as a color image or a grayscale image. The 3D camera may be a depth-sensing camera, such as a time-of-flight (TOF) camera or a structured light camera, or any other type of 3D camera. In some cases, the 2D camera and / or the 3D camera may include an image sensor, such as a charge-coupled device (CCD) sensor and / or a complementary metal-oxide semiconductor (CMOS) sensor. In an embodiment, the 3D camera may include a laser, a LIDAR device, an infrared device, a light / dark sensor, a motion sensor, a microwave detector, an ultrasonic detector, a radar detector, or any other device configured to capture depth information or other spatial structure information.

[0026] As described above, image information may be processed by computing system 1100. In an embodiment, computing system 1100 may include or be configured as a server (e.g., having one or more server blades, processors, etc.), a personal computer (e.g., a desktop computer, a laptop computer, etc.), a smartphone, a tablet computing device, and / or any other computing system. In an embodiment, any or all of the functionality of computing system 1100 may be implemented as part of a cloud computing platform. Computing system 1100 may be a single computing device (e.g., a desktop computer) or may include multiple computing devices.

[0027] 2A provides a block diagram illustrating an embodiment of a computing system 1100. The computing system 1100 in this embodiment includes at least one processing circuit 1110 and a non-transitory computer-readable medium (or media) 1120. In some instances, the processing circuit 1110 may include a processor (e.g., a central processing unit (CPU), a special purpose computer, and / or an on-board server) configured to execute instructions (e.g., software instructions) stored on the non-transitory computer-readable medium 1120 (e.g., computer memory). In some embodiments, the processor may be included in a separate / standalone controller operably coupled to other electronic / electrical devices. The processor may implement program instructions to control / interface with other devices, thereby causing the computing system 1100 to perform actions, tasks, and / or operations. In an embodiment, the processing circuit 1110 includes one or more processors, one or more processing cores, a programmable logic controller ("PLC"), an application specific integrated circuit ("ASIC"), a programmable gate array ("PGA"), a field programmable gate array ("FPGA"), any combination thereof, or any other processing circuit.

[0028] In an embodiment, the non-transitory computer readable medium 1120 that is part of the computing system 1100 may be a replacement for or in addition to the intermediate non-transitory computer readable medium 1400 discussed above. The non-transitory computer readable medium 1120 may be a storage device such as an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, such as, for example, a computer diskette, a hard disk drive (HDD), a solid state drive (SSD), a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, any combination thereof, or any other storage device. In some instances, the non-transitory computer readable medium 1120 may include multiple storage devices. In certain implementations, the non-transitory computer readable medium 1120 is configured to store image information generated by the camera 1200 and received by the computing system 1100. In some instances, the non-transitory computer readable medium 1120 may store one or more object recognition templates used to perform the methods and operations discussed herein. The non-transitory computer readable medium 1120 may alternatively or additionally store computer readable program instructions that, when executed by the processing circuit 1110, cause the processing circuit 1110 to perform one or more methodologies described herein.

[0029] FIG. 2B depicts one embodiment of the computing system 1100, the computing system 1100A, including a communication interface 1131. The communication interface 1131 may be configured to receive image information generated by, for example, the camera 1200 of FIGS. 1A-1D. The image information may be received via an intermediate non-transitory computer-readable medium 1400 or a network discussed above, or via a more direct connection between the camera 1200 and the computing system 1100 / 1100A. In an embodiment, the communication interface 1131 may be configured to communicate with the robot 1300 of FIG. 1C. If the computing system 1100 is external to the robot control system, the communication interface 1131 of the computing system 1100 may be configured to communicate with the robot control system. The communication interface 1131 may also be referred to as a communication component or communication circuitry, and may include, for example, communication circuitry configured to implement communication over a wired or wireless protocol. As examples, the communications circuitry may include an RS-232 port controller, a USB controller, an Ethernet controller, a Bluetooth controller, a PCI bus controller, any other communications circuitry, or combinations thereof.

[0030] 2C, the non-transitory computer-readable medium 1120 may include a storage space 1125 configured to store one or more data objects discussed herein. For example, the storage space may store object recognition templates, detection hypotheses, image information, object image information, robotic arm movement commands, and any additional data objects that the computing system discussed herein may need to access.

[0031] In an embodiment, the processing circuit 1110 may be programmed by one or more computer readable program instructions stored on a non-transitory computer readable medium 1120. For example, FIG. 2D illustrates a computing system 1100C, an embodiment of computing systems 1100 / 1100A / 1100B, in which the processing circuit 1110 is programmed by one or more modules, including an object recognition module 1121, a motion planning module 1129, and an object manipulation planning module 1126. The processing circuit 1110 may be further programmed with a hypothesis generation module 1128, an object registration module 1130, a template generation module 1132, a feature extraction module 1134, a hypothesis refinement module 1136, and a hypothesis verification module 1138. Each of the above modules may represent computer readable program instructions configured to perform a particular task when instantiated in one or more of the processors, processing circuits, computing systems, etc. described herein. Each of the above modules may operate in conjunction with one another to accomplish the functions described herein. Various aspects of the functionality described herein may be performed by one or more of the software modules described above, and the software modules and their descriptions are not to be understood as limiting the computational structure of the systems disclosed herein. For example, a particular task or function may be described with respect to a particular module, but that task or function may be performed by different modules as appropriate. Furthermore, the system functionality described herein may be performed by different sets of software modules configured to provide different breakdowns or allocations of functionality.

[0032] In embodiments, the object recognition module 1121 may be configured to obtain and analyze image information as discussed throughout this disclosure. Methods, systems, and techniques discussed herein with respect to image information may use the object recognition module 1121. The object recognition module may be further configured for object recognition tasks related to object identification as discussed herein.

[0033] The motion planning module 1129 can be configured to plan and execute robot movements. For example, the motion planning module 1129 can interact with other modules described herein to plan the motion of the robot 3300 for object pick and camera placement operations. The methods, systems, and techniques discussed herein with respect to robot arm movements and trajectories can be implemented by the motion planning module 1129.

[0034] The object manipulation planning module 1126 may be configured to plan and execute object manipulation activities of the robotic arm, such as, for example, grasping and releasing an object, and executing robotic arm commands to assist and facilitate such grasping and releasing. The object manipulation planning module 1126 may be configured to perform processing related to trajectory determination, pick-and-grip procedure determination, and end effector interaction with an object. The operation of the object manipulation planning module 1126 is described in further detail with respect to FIG.

[0035] The hypothesis generation module 1128 may be configured to perform template matching and recognition tasks to generate detection hypotheses. The hypothesis generation module 1128 may be configured to interact or communicate with any other necessary modules.

[0036] The object registration module 1130 may be configured to acquire, store, generate, and otherwise process object registration information that may be required for various tasks discussed herein. The object registration module 1130 may be configured to interact or communicate with any other necessary modules.

[0037] The template generation module 1132 may be configured to complete the object recognition template generation task. The template generation module 1132 may be configured to interact with the object registration module 1130, the feature extraction module 1134, and any other necessary modules.

[0038] The feature extraction module 1134 may be configured to complete feature extraction and generation tasks. The feature extraction module 1134 may be configured to interact with the object registration module 1130, the template generation module 1132, the hypothesis generation module 1128, and any other necessary modules.

[0039] The hypothesis refinement module 1136 may be configured to complete the hypothesis refinement task. The hypothesis refinement module 1136 may be configured to interact with the object recognition module 1121 and the hypothesis generation module 1128, as well as any other necessary modules.

[0040] The hypothesis verification module 1138 may be configured to complete the hypothesis verification task. The hypothesis verification module 1138 may be configured to interact with the object registration module 1130, the feature extraction module 1134, the hypothesis generation module 1128, the hypothesis refinement module 1136, and any other necessary modules.

[0041] With reference to Figures 2E, 2F, 3A and 3B, methods related to the object recognition module 1121 that may be implemented for image analysis are described. Figures 2E and 2F illustrate example image information associated with the image analysis method, while Figures 3A and 3B illustrate an example robot environment associated with the image analysis method. References herein related to image analysis by a computing system may be implemented according to or using spatial structure information, which may include depth information describing respective depth values ​​of various locations relative to a selected point. The depth information may be used to identify an object or estimate how an object is spatially arranged. In some instances, the spatial structure information may include, or may be used to generate, a point cloud describing the location of one or more surfaces of an object. Spatial structure information is just one form of possible image analysis, and other forms known to those skilled in the art may be used according to the methods described herein.

[0042] In an embodiment, the computing system 1100 may acquire image information representing an object in a camera field of view (e.g., 3200) of the camera 1200. The steps and techniques described below for acquiring image information may hereinafter be referred to as image information capture operation 3001. In some instances, the object may be one object 5012 from a plurality of objects 5012 in a scene 5013 of the field of view 3200 of the camera 1200. The image information 2600, 2700 may be generated by the camera (e.g., 1200) when the object 5012 is (or was) in the camera field of view 3200 and may describe one or more of the individual objects 5012 or the scene 5013. The object appearance describes the appearance of the object 5012 from the viewpoint of the camera 1200. If there are multiple objects 5012 in the camera field of view, the camera may generate image information representing multiple objects or a single object (such image information regarding a single object may be referred to as object image information) as appropriate. The image information may be generated by a camera (eg, 1200) when a group of objects is (or was) in the camera's field of view, and may include, for example, 2D image information and / or 3D image information.

[0043] As an example, Figure 2E depicts a first set of image information, more specifically, 2D image information 2600, which is generated by camera 1200, as described above, and which represents object 3410A / 3410B / 3410C / 3410D / 3401 of Figure 3A. More specifically, 2D image information 2600 may be a grayscale or color image and may describe the appearance of object 3410A / 3410B / 3410C / 3410D / 3401 from the perspective of camera 1200. In an embodiment, 2D image information 2600 may correspond to a single color channel (e.g., a red, green, or blue color channel) of a color image. When the camera 1200 is disposed above the object 3410A / 3410B / 3410C / 3410D / 3401, the 2D image information 2600 may represent an appearance of a top surface of each of the objects 3410A / 3410B / 3410C / 3410D / 3401. In the embodiment of FIG. 2E, the 2D image information 2600 may include a respective portion 2000A / 2000B / 2000C / 2000D / 2550, also referred to as an image portion or object image information, representing a respective surface of the object 3410A / 3410B / 341C / 3410D / 3401. In FIG. 2E, each image portion 2000A / 2000B / 2000C / 2000D / 2550 of the 2D image information 2600 may be an image region, or more specifically, a pixel region (where an image is formed by pixels). Each pixel in the pixel region of the 2D image information 2600 may be characterized as having a location described by a set of coordinates [U, V], which may have values ​​relative to the camera coordinate system, as shown in FIG. 2E and FIG. 2F, or some other coordinate system. Each of the pixels may also have an intensity value, such as a value from 0 to 255 or 0 to 1023. In further embodiments, each of the pixels may include any additional information associated with the pixel in various formats (e.g., hue, saturation, intensity, CMYK, RGB, etc.).

[0044] As mentioned above, the image information may be all or a portion of an image, such as the 2D image information 2600, in some embodiments. In an example, the computing system 1100 may be configured to extract the image portion 2000A from the 2D image information 2600 to obtain only image information associated with the corresponding object 3410A. When an image portion (such as the image portion 2000A) is directed to a single object, it may be referred to as object image information. The object image information need not encompass only information about the object of interest. For example, the object of interest may be near, under, on, or otherwise in the vicinity of one or more other objects. In such a case, the object image information may include information about the object of interest as well as one or more adjacent objects. The computing system 1100 may extract the image portion 2000A by performing an image segmentation or other analysis or processing operation based on the 2D image information 2600 and / or the 3D image information 2700 illustrated in FIG. 2F. In some implementations, image segmentation or other operations may include detecting image locations where physical edges of objects in the 2D image information 2600 appear (e.g., edges of the object) and using such image locations to identify object image information that is limited to represent individual objects in the camera field of view (e.g., 3200) and to substantially exclude other objects. By "substantially exclude," it is meant that the image segmentation or other processing techniques are designed and configured to exclude non-target objects from the object image information, while it is understood that errors may occur, noise may be present, and various other factors may result in the inclusion of portions of other objects.

[0045] 2F depicts an example in which the image information is 3D image information 2700. More specifically, the 3D image information 2700 may include, for example, a depth map or point cloud indicating respective depth values ​​of various locations on one or more surfaces (e.g., top surfaces, or other exterior surfaces) of the object 3410A / 3410B / 3410C / 3410D / 3401. In some implementations, an image segmentation operation for extracting image information may include detecting image locations where physical edges of an object (e.g., edges of a box) appear in the 3D image information 2700, and using such image locations to identify image portions (e.g., 2730) that are limited to representing individual objects within the camera field of view (e.g., 3410A).

[0046] Each depth value may be relative to the camera 1200 generating the 3D image information 2700, or may be relative to some other reference point. In some embodiments, the 3D image information 2700 may include a point cloud including respective coordinates for various locations on the structure of the object that are within the camera field of view (e.g., 3200). In the example of FIG. 2F, the point cloud may include respective sets of coordinates describing respective surface locations of the objects 3410A / 3410B / 3410C / 3410D / 3401. The coordinates may be 3D coordinates, such as [XYZ] coordinates, and may have values ​​relative to the camera coordinate system, or some other coordinate system. For example, the 3D image information 2700 may include locations 27101-2710, also referred to as physical locations on the surface of the object 3410D. n 27. The 3D image information 2700 may further include a first image portion 2710, also referred to as image portion, which indicates respective depth values ​​for a set of 3D images. Furthermore, the 3D image information 2700 may further include a second portion, a third portion, a fourth portion, and a fifth portion 2720, 2730, 2740, and 2750. These portions are then respectively denoted as 27201 to 2720. n , 27301~2730 n , 27401~2740 n , and 27501 to 2750 n3A , the image portion 2710 may be narrowed to refer to only the image portion 2710. Similar to the discussion of the 2D image information 2600, the identified image portion 2710 may relate to an individual object and may be referred to as object image information. Thus, object image information, as used herein, may include 2D and / or 3D image information.

[0047] In an embodiment, an image normalization operation may be performed by the computing system 1100 as part of acquiring image information. The image normalization operation may involve transforming an image or image portion generated by the camera 1200 to generate a transformed image or a transformed image portion. For example, where the acquired image information, which may include 2D image information 2600, 3D image information 2700, or a combination of the two, may be subjected to an image normalization operation to attempt to change the image information in terms of viewpoint, object pose, and lighting conditions associated with the visual description information. Such normalization may be performed to facilitate a more accurate comparison between the image information and the model (e.g., template) information. The viewpoint may refer to the pose of the object relative to the camera 1200 and / or the angle at which the camera 1200 is looking at the object when the camera 1200 generates an image representing the object.

[0048] For example, image information may be generated during an object recognition operation when a target object is within the camera field of view 3200. The camera 1200 may generate image information representing a target object when the target object has a particular pose relative to the camera. For example, the target object may have a pose that makes its top surface perpendicular to the optical axis of the camera 1200. In such an example, the image information generated by the camera 1200 may represent a particular viewpoint, such as a top view of the target object. In some instances, when the camera 1200 is generating image information during an object recognition operation, the image information may be generated at a particular lighting condition, such as a lighting intensity. In such an example, the image information may represent a particular lighting intensity, lighting color, or other lighting condition.

[0049] In an embodiment, the image normalization operation may involve adjusting an image or image portion of a scene generated by a camera to better match the image or image portion to viewpoint and / or lighting conditions associated with the object recognition template information. The adjustment may involve transforming the image or image portion to generate a transformed image that matches at least one of the object pose or lighting conditions associated with the visual description information of the object recognition template.

[0050] Viewpoint adjustment may involve processing, warping, and / or shifting an image of a scene so that the image represents the same perspective as the visual description information that may be included in the object recognition template. Processing may include, for example, changing the color, contrast, or lighting of the image, warping the scene may include changing the size, dimensions, or ratio of the image, and shifting the image may include changing the position, orientation, or rotation of the image. In an exemplary embodiment, processing, warping, and / or shifting may be used to change objects in the image of a scene to have an orientation and / or size that matches or better corresponds to the visual description information of the object recognition template. If the object recognition template describes a front view (e.g., a top view) of some objects, the image of the scene may be warped to also represent the front view of the objects in the scene.

[0051] Further aspects of the object recognition methods embodied in this specification are described in more detail in U.S. patent application Ser. No. 16 / 991,510, filed Aug. 12, 2020, and U.S. patent application Ser. No. 16 / 991,466, filed Aug. 12, 2020, each of which is incorporated by reference herein.

[0052] In various embodiments, the terms "computer readable instructions" and "computer readable program instructions" are used to describe software instructions or computer code configured to perform various tasks and operations. In various embodiments, the term "module" refers broadly to a collection of software instructions or code configured to cause the processing circuitry 1110 to perform one or more functional tasks. The modules and computer readable instructions may be described as performing various operations or tasks when a processing circuitry or other hardware component executes the module or computer readable instructions.

[0053] 3A and 3B illustrate an exemplary environment in which computer readable program instructions stored on a non-transitory computer readable medium 1120 are utilized via a computing system 1100 to increase the efficiency of object identification, detection, and retrieval operations and methods. Image information acquired by the computing system 1100 and illustrated in FIG. 3A influences the system's decision-making procedures and command output to a robot 3300 present within the object environment.

[0054] 3A and 3B illustrate an example environment in which the processes and methods described herein may be implemented. FIG. 3A depicts an environment having a system 3000 (which may be an embodiment of systems 1000 / 1500A / 1500B / 1500C of FIGS. 1A-1D) including at least a computing system 1100, a robot 3300, and a camera 1200. Camera 1200 may be an embodiment of camera 1200 and may be configured to generate image information representing a scene 5013 within a camera field of view 3200 of camera 1200, or more specifically, representing objects (e.g., boxes) within camera field of view 3200, such as objects 3000A, 3000B, 3000C, and 3000D. In one example, each of objects 3000A-3000D may be a container, such as, for example, a box or crate, while object 3550 may be, for example, a pallet on which the container is disposed. Additionally, each of objects 3000A-3000D may be a container that further contains individual objects 5012. Each object 5012 may be, for example, a rod, a bar, a gear, a bolt, a nut, a screw, a nail, a rivet, a spring, a linkage, a gear tooth, or any other type of physical object, as well as an assembly of multiple objects. FIG. 3A illustrates an embodiment that includes multiple containers of objects 5012, while FIG. 3B illustrates an embodiment that includes a single container of objects 5012.

[0055] In an embodiment, the system 3000 of FIG. 3A may include one or more light sources. The light sources may be, for example, light emitting diodes (LEDs), halogen lamps, or any other light source, and may be configured to emit visible light, infrared light, or any other form of light toward the surface of the objects 3000A-3000D. In some implementations, the computing system 1100 may be configured to communicate with the light sources to control when the light sources are activated. In other implementations, the light sources may operate independently of the computing system 1100.

[0056] In an embodiment, the system 3000 may include a camera 1200 or multiple cameras 1200, including a 2D camera configured to generate 2D image information 2600 and a 3D camera configured to generate 3D image information 2700. The camera 1200 or multiple cameras 1200 may be mounted on or fixed to the robot 3300, may be stationary in the environment, and / or may be fixed to a dedicated robotic system separate from the robot 3300 used for object manipulation, such as a robotic arm, gantry, or other automated system configured for camera movement. FIG. 3A shows an example with a stationary camera 1200 and a handheld camera 1200, while FIG. 3B shows an example with only a stationary camera 1200. The 2D image information 2600 (e.g., a color image or a grayscale image) may describe the appearance of one or more objects, such as objects 3000A / 3000B / 3000C / 3000D or object 5012, in the camera field of view 3200. For example, the 2D image information 2600 may capture or otherwise represent visual details disposed on an exterior surface (e.g., a top surface) of each of the objects 3000A / 3000B / 3000C / 3000D and 5012, and / or a contour of their exterior surfaces. In an embodiment, the 3D image information 2700 may describe a structure of one or more of the objects 3000A / 3000B / 3000C / 3000D / 3550 and 5012, which structure for an object may also be referred to as the structure of the object or the physical structure of the object. For example, the 3D image information 2700 may include a depth map, or more generally, depth information that may describe respective depth values ​​of various locations within the camera field of view 3200 relative to the camera 1200 or relative to some other reference point. The locations corresponding to each depth value may be locations on various surfaces (also called physical locations) within the camera field of view 3200, such as locations on the top surfaces of each of objects 3000A / 3000B / 3000C / 3000D / 3550 and 5012.In some instances, the 3D image information 2700 may include a point cloud, which may include multiple 3D coordinates describing various locations on one or more exterior surfaces of the objects 3000A / 3000B / 3000C / 3000D / 3550 and 5012, or some other object within the camera field of view 3200. The point cloud is shown in FIG.

[0057] In the example of Figures 3A and 3B, a robot 3300 (which may be an embodiment of the robot 1300) may include a robot arm 3320 attached at one end to a robot base 3310 and attached at the other end to or formed by an end effector device 3330, such as a robot gripper. The robot base 3310 may be used to mount the robot arm 3320, while the robot arm 3320, and more specifically, the end effector device 3330, may be used to interact with one or more objects in the environment of the robot 3300. The interaction (also referred to as robot interaction) may include, for example, gripping or otherwise grasping at least one of the objects 3000A-3000D and 5012. For example, the robot interaction may be part of an object picking operation to identify, detect, and remove the object 5012 from the container. The end effector device 3330 may have a suction cup or other component to grip or grasp the object 5012. The end effector device 3330 may be configured to grasp or grip an object through contact with a single face or surface of the object, for example, via an upper surface, using a suction cup or other gripping component.

[0058] The robot 3300 may further include additional sensors configured to obtain information used to implement a task, such as to manipulate a structural member and / or to transport the robotic unit. The sensors may include devices configured to detect or measure one or more physical characteristics of the robot 3300 (e.g., its state, condition, and / or location of one or more structural members / joints) and / or one or more physical characteristics of the surrounding environment. Some examples of sensors may include accelerometers, gyroscopes, force sensors, strain gauges, tactile sensors, torque sensors, position encoders, etc.

[0059] The computing system 1100 includes a robot arm 3320 including or attached to an end effector device 3330, and a control system configured to communicate with a robot 3300 having a camera 1200 attached to the robot arm 3320. Figures 3C and 3D illustrate an embodiment of a robot 3300 with which the computing system 1100 can communicate and command / control to accomplish the methods described herein. In an embodiment, the camera 1200 is disposed elsewhere in the object handling environment 3400 while communicating with the control system of the computing system 1100 via either a wireless or hard-wired connection. The robot 3300 may include physical or structural members 3321a, 3321b connected at joints 3320a, 3320b to form the robot arm 3320 and end effector device 3330, allowing a larger range of motion (e.g., rotational and / or translational displacement). The physical or structural member 3321a may further connect to the robot base 3310 via joint 3320a. The robot 3300 may include actuation devices such as motors, actuators, wires, artificial muscles, electroactive polymers, etc. (not shown) configured to drive or manipulate (e.g., displace and / or reorient) the structural members 3321a, 3321b around or at the corresponding joints 3320a, 3320b. For example, the robot arm 3300 may be able to rotate 360° in all directions about the joint 3320a relative to the robot base 3310, or the structural members 3321a, 3321b may rotate 360° in all directions at any point that connects to the joint 3320a, 3320b connection. The robot arm 3300 can further translate anywhere within a hemispherical three-dimensional space, where the fully extended length of the robot arm 3300 (i.e., straightened, or 180°) acts as the radius of the hemispherical three-dimensional space, measured from the central axis of the robot base 3310 (i.e., where the robot arm 3320 connects to the robot base 3310) to the tip or end of the end effector instrument 3330.

[0060] The connected structural members 3321a, 3321b and joints 3320a, 3320b may form a power chain configured to manipulate an end effector device 3330 configured to perform one or more tasks (e.g., gripping, spinning, welding, etc.) depending on a desired use of the robot 3300. The robot 3300 may include actuation devices such as motors, actuators, wires, artificial muscles, electroactive polymers, etc. (not shown) configured to drive or manipulate (e.g., displace and / or reorient) the end effector device 3330. In general, the end effector device 3330 may provide the ability to grip objects 3410A / 3410B / 3410C / 3410D / 3401 of various sizes and shapes. The object 3410A / 3410B / 3410C / 3410D / 3401 may be any object, including, for example, a rod, a bar, a gear, a bolt, a nut, a screw, a nail, a rivet, a spring, a linkage, a gear tooth, a disk, a washer, or any other type of physical object, as well as an assembly of multiple objects. The end effector device 3330 may include at least one gripper 3332 having gripping fingers 3332a, 3332b, as illustrated in FIG. 3C. The gripping fingers 3332a, 3332b may translate relative to one another to clamp, grasp, or otherwise secure the object 3410A / 3410B / 3410C / 3410D / 3401. In an embodiment, the end effector device 3330 includes at least two grippers 3332, 3334 having gripping fingers 3332a, 3332b, 3334a, 3334b, respectively, as illustrated in FIG. 3D. The gripping fingers 3332a, 3332b can translate relative to one another, and the gripping fingers 3334a, 3334b can translate relative to one another to clamp, grip, or otherwise secure an object 3410A / 3410B / 3410C / 3410D / 3401. In an embodiment, the end effector device 3330 may include three or more grippers (not shown) and / or grippers with three or more gripping fingers (not shown), each having a translation capability designed to clamp, grip, or otherwise secure an object.

[0061] The robot 3300 can be configured for a location within the object handling environment 3400, including a container 3420 having an object 3410A / 3410B / 3410C / 3410D / 3401 disposed thereon or therein for delivery or transfer to a destination 3440 within the object processing environment 3400. The container 3420 may be any container suitable for holding the object 3410A / 3410B / 3410C / 3410D / 3401, such as, for example, a bin, box, bucket, or pallet. The object 3410A / 3410B / 3410C / 3410D / 3401 may be any object, including, for example, a rod, bar, gear, bolt, nut, screw, nail, rivet, spring, linkage, gear tooth, disk, washer, or any other type of physical object, as well as an assembly of multiple objects. In an embodiment, object 3410A / 3410B / 3410C / 3410D / 3401 may refer to an object accessible from container 3420, e.g., having a mass ranging from a few grams to a few kilograms, and a size ranging from, e.g., 5 mm to 500 mm. For example and illustrative purposes, the description of method 4000 herein refers to a ring-shaped object as target object 3510a (FIGS. 5B and 6A-6C) within a plurality of objects 3500 (shown in FIG. 5B) with which computer system 1100 and robot 3300 may interact using the methods described herein. The plurality of objects 3500 may be substantially identical in terms of size, shape, weight, and material composition. In an embodiment, the plurality of objects 3500 may differ from one another in size, shape, weight, and material composition, as previously described. The particular shapes of objects discussed herein, e.g., are used for example purposes only, and the methods and processes described herein may be used or employed with objects of different shapes, as desired.

[0062] Thus, in regard to the above, the computing system 1100 may be configured to operate as follows to transfer a target object from a source or container 3420 to a destination 3440.

[0063] 4 provides a flow diagram illustrating the overall flow of methods and operations for detecting, planning, picking, transferring, and placing a target object according to embodiments herein. The method 4000 of detecting, planning, picking, transferring, and placing may include any combination of features of the sub-methods and operations described herein. The method 4000 may include any or all of an object detection operation 4002, an object graspability determination operation 4003, a target selection operation 4004, a trajectory determination operation 4005, a pick / grip procedure determination operation 4006, a robot arm / end effector device trajectory execution operation 4008, an end effector interaction operation 4010, and a destination trajectory execution operation 4012 for controlling the robot arm 3320. The object detection operation 4002 may be performed in real time or in a pre-processing or offline environment outside the context of the robot operation. Thus, in some embodiments, these operations and methods may be performed in advance to facilitate subsequent actions by the robot. The detect object operation 4002 and determine object graspability operation 4003 may be the first steps in the planning portion of the method 4000. The select target operation 4004, determine trajectory operation 4005, and determine pick / grip procedure operation 4006 may provide the remaining steps in the planning portion and may be performed multiple times during the method 4000. The execute robot arm / end effector device trajectory operation 4008, the execute end effector interaction operation 4010, and the execute destination trajectory operation 4012 for controlling the robot arm 3320 may each be performed in the context of a robot operation to detect, identify, and remove a target object from a container.

[0064] In operation 4002, the method 4000 includes detecting, via the camera 1200, a plurality of objects 3500 in a container or source of objects 3420. The objects 3500 may represent a plurality of physical real-world objects (FIG. 5A). The operation 4002 may generate a detection result 3520 for one or more of the objects 3500 in the container 3420. The detection result 3520 may include a digital representation of the plurality of objects 3500 in the container 3420 (FIG. 5B), which may be referred to as individually detected objects 3510. Further operations of the method 4000 may determine from the detected objects 3510 which are target objects 3510a or target objects 3511a / 3511b, and / or non-graspable objects 3510b (e.g., as discussed with respect to FIG. 7B).

[0065] Operation 4002 may include analyzing information (e.g., image information) received from camera 1200 to generate detection result 3520 (FIG. 5C) according to methods described herein. Information received from camera 1200 may include images of environment 3400 of object container 3420 of a plurality of objects 3500. As discussed above, the plurality of objects 3500 may include detected object 3510.

[0066] Generating the detection result 3520 may include identifying the plurality of objects 3500 in the object container 3420 and then identifying the detected objects 3510 from which a target object 3510a or target object 3511a / 3511b is later determined to be picked via the robot 3300 and transported to the destination 3440. FIG. 5B provides a visual depiction of the detection result 3520 for the plurality of detected objects 3510 among the plurality of objects 3500 in the container 3420 (their physical representations are provided as FIG. 5A). FIG. 5C illustrates the physical objects 3500 present in the physical world, while the detected objects 3510 refer to a representation of the physical objects 3500 described by the detection result 3520. The detection results 3520 may include multiple object representations 4013 for each of the detected objects 3510, including information about each of the detected objects 3510, such as the location of the detected object 3510 within the container 3420, the location of the detected object 3510 relative to the other detected objects 3510 (e.g., whether the detected object 3510 is on top of a pile of multiple objects 3500 or under other adjacent detected objects 3510), the orientation and pose of the detected object 3510, the confidence of the object detection, available gripping models 3350a / 3350b / 3350c (as described in more detail below), or combinations thereof.

[0067] Thus, operation 4002 of method 4000 may include obtaining a detection result 3520 including a plurality of object representations 4013 based on the object detection. The computer system 1100 may use the plurality of object representations 4013 of all detected objects 3510 from the detection result 3520 in determining the valid grasp model 3350a / 3350b / 3350c. Each of the detected objects 3510 may have a corresponding detection result 3520 representing digital information (i.e., object representations 4013) for each of the detected objects 3510. In an embodiment, the corresponding detection result 3520 may incorporate the plurality of detected objects 3510 into a plurality of objects 3500 that are physically present in the real world. The detected objects 3510 may represent digital information (i.e., object representations 4013) for each of the detected objects 3500.

[0068] In one embodiment, identifying the plurality of objects 3500 to obtain the detection result 3520 may be accomplished by any suitable means. In an embodiment, identifying the plurality of objects 3500 may include processes including object registration, template generation, feature extraction, hypothesis generation, hypothesis refinement, and hypothesis verification, as performed, for example, by the hypothesis generation module 1128, the object registration module 1130, the template generation module 1132, the feature extraction module 1134, the hypothesis refinement module 1136, and the hypothesis verification module 1138. These processes are described in detail in U.S. Patent Application No. 17 / 884,081, filed August 9, 2022, the entire contents of which are incorporated herein.

[0069] Object registration is a process that involves obtaining and using object registration data, e.g., known previously stored information related to the object 3500, to generate an object recognition template for use in identifying and recognizing similar objects in a physical scene. Template generation is a process that involves generating a set of object recognition templates for a computing system to use in identifying the object 3500 for further operations related to object picking. Feature extraction (also called feature generation) is a process that involves the extraction or generation of features from object image information for use in object recognition template generation. Hypothesis generation is a process that involves generating one or more object detection hypotheses, e.g., based on a comparison of the object image information to one or more object recognition templates. Hypothesis refinement is a process for refining the match between the object recognition template and the object image information even in scenarios where the object recognition template does not exactly match the object image information. Hypothesis validation is a process in which a single hypothesis from multiple hypotheses is selected as the best fit or best choice for the object 3500.

[0070] In operation 4003, the method 4000 includes identifying graspable objects from among the plurality of objects 3500. As a step in the planning portion of the method 4000, operation 4003 includes determining graspable and non-graspable objects from the detected objects 3510. Operation 4003 may be performed based on the detected objects 3510 to assign a grasp model to each detected object 3510 or to determine that the detected object 3510 is a non-graspable object 3510b.

[0071] The grasping models 3350a / 3350b / 3350c explain how the detected object 3510 may be grasped by the end effector device 3330. For illustrative purposes, Figures 6A-6C illustrate three different grasping models 3350a / 3350b / 3350c for gripping the target object 3510a, although it should be understood that other grasping models are possible.

[0072] Figure 6A, illustrated as gripping model 3350a, demonstrates an inner chuck in which the gripper fingers 3332a / 3332b / 3334a / 3334b perform a reverse clamping motion against the inner wall of the ring of the target object 3510a (i.e., the gripper fingers 3332a / 3332b / 3334a / 3334b translate outward or away from each other once both enter the ring of the target object 3510a).

[0073] FIG. 6B illustrates a gripping model 3350b demonstrating an inner and outer chuck in which gripper fingers 3332a / 3332b / 3334a / 3334b clamp the inner and outer walls of a ring of target object 3510a.

[0074] FIG. 6C, illustrated as gripping model 3350c, demonstrates a side chuck in which gripper fingers 3332a / 3332b / 3334a / 3334b grip the outer disk portion of the ring of target object 3510a.

[0075] Each of the gripping models 3350a / 3350b / 3350c may be ranked according to factors such as predicted grip stability 4016, which may have an associated transfer speed modifier that may determine the speed, acceleration, and / or deceleration at which an object may be moved by the robotic arm 3320. For example, the associated transfer speed modifier is a value that determines the speed at which the robotic arm 3320 and / or the end effector device 3330 moves. The value may be set between zero and one, with zero representing a complete stop (e.g., no movement, completely still) and one representing the maximum operating speed of the robotic arm 3320 and / or the end effector device 3330. The transfer speed modifier may be determined offline (e.g., through real-world testing) or in real-time (e.g., through computer model simulations to account for friction, gravity, and momentum).

[0076] The predicted grip stability 4016 may further be an indication of how safe the target object 3510a will be once gripped by the end effector device 3330. For example, grip model 3350a may have a higher predicted grip stability 4016 than grip model 3350b, which may have a higher predicted grip stability 4016 than grip model 3350c. In other embodiments, different grip models 3350 may be ranked differently according to their predicted grip stability 4016.

[0077] Processing of the detection results 3520 may provide data indicating whether each of the detected objects 3510 may be grasped by one or more of the grasping models based on a plurality of object representations 4013 for each of the detected objects 3510, including the location of the detected object 3510 within the container 3420, the location of the detected object 3510 relative to the other detected objects 3510 (e.g., whether the detected object 3510 is on top of a pile of multiple objects 3500 or under other adjacent detected objects 3510), the orientation and pose of the detected object 3510, the confidence of the object detection, available grasping models 3350a / 3350b / 3350c (described in more detail below), or a combination thereof. For example, one of the detected objects 3510 may be grasped according to grasping models 3350a and 3350b, but not according to grasping model 3350c.

[0078] A detected object 3510 may be determined as an ungraspable object 3510b if no graspable model for the object could be found. For example, the detected object 3510 may not be accessible for grasping by any of the graspable models 3350a / 3350b / 3350c (because they are covered at odd angles, partially buried, partially invisible, etc.) and therefore not graspable by the end effector device 3330. The ungraspable objects 3510b may be removed from the detection results 3520 such that no further processing is performed on them, for example, by removing them from the detection results 3520 or by flagging them as ungraspable.

[0079] Removing the non-graspable objects 3510b from the plurality of objects 3500 and / or the detected objects 3510 may be further performed according to the following: In an embodiment, based on at least one of the plurality of object representations 4013 of the detection result 3520, further determining and removing the non-graspable objects 3510b from the remaining detected objects 3510 to evaluate the target object 3510a. As described above, the object representation 4013 of each of the detected objects 3510 includes, among other things, the position of the detected object' 3510 in the container 3420, the position of the detected object 3510 relative to other detected objects 3510, the orientation and pose of the detected object 3510, the confidence of the object detection, the available gripping models 3350a / 3350b / 3350c, or a combination thereof. For example, the ungraspable object 3510b may be located within the container 3420 in a manner that does not allow actual access by the end effector device 3330 (e.g., the ungraspable object 3510b is leaning against a wall or corner of the container). The ungraspable object 3510b may be determined to be unavailable for picking / grasping by the end effector device 3330 due to the orientation of the ungraspable object 3510b (e.g., the orientation / pose of the ungraspable object 3510b is such that the end effector device 3330 cannot actually grasp or pick the ungraspable object 3510b using any of the available grasp models 3350a / 3350b / 3350c). The ungraspable object 3510b may be surrounded or covered by other detected objects 3510 in a manner that does not allow practical access by the end effector device 3330 (e.g., the ungraspable object 3510b is located at the bottom of a container covered by other detected objects 3510, the ungraspable object 3510b is wedged between multiple other detected objects 3510). As discussed above in operation 4002, when detecting multiple objects, the computer system 1100 may output a low confidence in detecting the ungraspable object 3510b (e.g., the computer system 1100 is not entirely certain / believes that the ungraspable object 3510b has been properly identified compared to the other detected objects 3510).

[0080] As a further example, the ungraspable object 3510b may be a detected object 3510 that does not have an available gripping model 3350a / 3350b / 3350c based on the detection results 3520. For example, the computer system 1100 may determine that the end effector device 3330 is unable to pick / grasp the ungraspable object 3510b with any of the gripping models 3350a / 3350b / 3350c due to any combination of the object representations 4013 described above, including, among others, the location of the ungraspable object 3510b within the container, its location relative to other detected objects 3510, its orientation, its confidence level, or its object type, as described further herein. The ungraspable object 3510b may be determined by the computer system 1100 due to the ungraspable object 3510b having no available gripping model 3350a / 3350b / 3350c. For example, as further described herein with respect to pick / grip procedure operation 4006, an ungraspable object 3510b may be determined by the computer system 1100 to be ungraspable because the ungraspable object 3510b has a lower predicted grip stability 4016 or other measured variable than other detected objects 3510.

[0081] The remaining graspable objects may be ranked or ordered according to one or more criteria. The graspable objects may be ranked according to any combination of detection confidence (e.g., confidence in the detection result associated with the object), object location (e.g., ease of access, objects that are not clearly visible, obstructed, or embedded may have a higher ranking), and a ranking of the grasp models identified for the graspable objects.

[0082] At operation 4004, the method 4000 includes target selection. At operation 4004, the target object 3510a or the target object 3511a / 3511b may be selected from the graspable objects.

[0083] 7C and 7D, the graspable object identified by operation 4003 may be a candidate object 3512a / 3512b. The candidate objects 3512a / 3512b may be further weeded out by excluding or removing any objects that do not have an inverse kinematic solution. The candidate objects 3512a / 3512b lack an inverse kinematic solution (e.g., a solution for the robot arm 3320 to move itself to a position that allows grasping of the candidate object 3512a / 3512b and then move away from the grasping operation). For example, an inverse kinematic solution may not be found if the calculated configuration of the robot 3300 to reach the candidate object 3512a / 3512b violates the constraints of the robot 3300, the robot arm 3320, and / or the end effector device 3330. In determining whether an inverse kinematics solution exists for the candidate object 3512a / 3512b, the computing system 1100 may determine a trajectory for the candidate object 3512a / 3512b, for example, according to the methods discussed below with respect to operation 4005. In an embodiment, the graspable detected object 3510 may be located in an area of ​​the object source 3420 that does not allow the robot arm 3320 to be correctly positioned or configured to properly grasp that particular candidate object 3512a / 3512b or to leave after grasping the candidate object 3512a / 3512b.

[0084] For each candidate object 3512a / 3512b from the graspable objects the following can be performed: Candidate objects can be selected for processing, for example in an order according to the ranking of the graspable objects described above.

[0085] 7C, candidate object 3512a is referred to as a primary candidate object 3512a, e.g., an object that may be a first object in a double pick-and-pick operation. Candidate object 3512b may be a secondary candidate object 3512b, e.g., an object that may be a second object in a double pick-and-pick operation.

[0086] For each primary candidate object 3512a, the remaining secondary candidate objects 3512b may be filtered or removed according to the following: First, secondary objects 3512b within the obstruction range 3530 of the primary candidate object 3512a may be removed. The obstruction range 3530 represents the minimum distance from the first object at which other nearby objects are unlikely to shift in position or attitude when the first object is removed from the pile of objects. The obstruction range 3530 may depend on the size of the object and / or its shape (larger objects may require a larger range and some object shapes may cause more obstruction when moved). Thus, secondary candidate objects 3512b that are likely to be obstructed or moved during grasping of the primary candidate object 3512a may be removed.

[0087] The remaining secondary candidate objects 3512b may be further filtered or removed according to the similarity of the gripping models 3350a / 3350b / 3350c identified for the primary candidate object 3512a and the secondary candidate object 3512b. In an embodiment, the secondary candidate object 3512b may be removed if it has an assigned gripping model that is different from that of the primary candidate object 3512a. In an embodiment, the secondary candidate object 3512b may be removed if the grip stability of the gripping model 3350a / 3350b / 3350c assigned to the secondary candidate object 3512b differs from the grip stability of the gripping model 3350a / 3350b / 3350c assigned to the primary candidate object 3512a by more than a threshold. Object transfer may be optimized by providing robotic motion at maximum speed. As described above with respect to different gripping models 3350a / 3350b / 3350c, some gripping models 3350a / 3350b / 3350c have greater grip stability, thereby allowing greater speed of robotic movement. Selecting primary candidate objects 3512a and secondary candidate objects 3512b with gripping models 3350a / 3350b / 3350c that have the same or similar grip stability allows an increase in the speed of robotic movement. If the grip stabilities are different, the speed of robotic movement is limited to the speed allowed by the lower grip stability. Thus, in a scenario where multiple objects with high grip stability and multiple objects with low grip stability are available, it is advantageous to pair the object with the high grip stability with the object with the low grip stability.

[0088] The remaining secondary candidate objects 3512b may be further filtered or removed according to an analysis of potential trajectories between the primary candidate objects 3512a and the secondary candidate objects 3512b. If an inverse kinematics solution cannot be generated between the primary candidate objects 3512a and the secondary candidate objects 3512b, the secondary candidate object 3512b may be removed. As discussed above, an inverse kinematics solution may be identified through a trajectory determination similar to that described with respect to operation 4005.

[0089] It may then be determined that gripping the primary candidate object 3512a precludes gripping of the secondary candidate object 3512b. Referring now to Figure 7D, a bounding box 3600 may be generated by the computer system 1100 around at least one of the grippers 3332 / 3334 designated for interaction with each of the primary object 3512a and secondary objects 3512b, as illustrated in Figure 7D. When a second one of the grippers 3332 / 3334 attempts to approach, move with, interact with, grasp or move away from the secondary candidate object 3512b, the bounding box 3600 can be used by the computer system 1100 to determine whether the pose of the gripper 3332 / 3334 while gripping the primary candidate object 3512a (which has a bounding box 3600 generated around it) will result in a collision of the bounding box 3600 with other objects in the object handling environment 3400 / object source or container 3420 and / or the plurality of objects 3500. In doing so, the computer system 1100 can determine whether the primary object 3512a and secondary object 3512b grasped by the gripper 3332 / 3334 covered by the bounding box 3600 collide with other objects 3500 and / or the object handling environment 3400 in a manner that may cause the primary candidate object 3512a to be knocked out of the gripper 3332 / 3334's grasp while grasping the secondary candidate object 3512b.

[0090] Other means of filtering or removing secondary candidate objects 3512b may also be employed. For example, in an embodiment, secondary objects 3512b having a different orientation than the primary object 3512a may be removed. In an embodiment, secondary objects 3512b having a different object type or model than the primary object 3512a may be removed.

[0091] After removing the secondary candidate objects 3512b, object pairs between the primary candidate objects 3512a and the non-removed secondary candidate objects 3512b can be generated for trajectory determination. In an embodiment, each primary candidate object 3512a can be assigned a single secondary candidate object 3512b to form an object pair. In the case of multiple non-removed secondary candidate objects 3512b, a single secondary candidate object 3512b can be selected, for example, according to the simplest or fastest trajectory between the primary candidate object 3512a and the secondary candidate object 3512b and / or based on the ranking of graspable objects as described above with respect to operation 4003. In a further embodiment, each primary candidate object 3512a can be assigned multiple secondary candidate objects 3512b to form multiple object pairs, and a trajectory can be computed for each. In such an embodiment, the fastest or easiest trajectory can be selected to finalize the pairing between the primary candidate object 3512a and the secondary candidate object 3512b.

[0092] Once the primary object 3512a is paired with its respective secondary object 3512b from the graspable object, the computer system 1100 can designate each primary object 3512a paired with its respective secondary object 3512b as a target object 3511a / 3511b for grasp determination, robot arm trajectory execution, end effector interaction, and destination trajectory execution, as detailed in operations 4006 / 4008 / 4010 / 4012, respectively, herein.

[0093] In an embodiment, a first target object 3511a of the plurality of target objects 3511a / 3511b is associated with a first gripping model 3350a / 3350b / 3350c, and a second target object 3511b of the plurality of target objects 3511a / 3511b is associated with a second gripping model 3350a / 3350b / 3350c. The gripping model 3350a / 3350b / 3350c selected for the first target object 3511a may be similar or identical to the gripping model 3350a / 3350b / 3350c selected for the second target object 3511b based on at least one of the plurality of object representations 4013 of the detection result 3520, as described above. For example, a first target object 3511a may be gripped by gripper 3332 using gripping model 3350a in which gripper fingers 3332a, 3332b perform an inner chucking, or reverse pinching, motion against the inner wall of a ring of first target object 3511a, as shown in Figures 8A-8C. A second target object 3511b may also be gripped by gripper 3334 using gripping model 3350a in which gripper fingers 3334a, 3334b perform an inner chucking, or reverse pinching, motion against the inner wall of a ring of target object 3511b, as shown in Figures 9A-9C.

[0094] In operation 4005, the method 4000 may include determining a robot trajectory. Operation 4005 may include at least determining an arm approach trajectory 3360, determining an end effector device approach trajectory 3362, and determining a destination approach trajectory 3364.

[0095] The operation 4005 may include determining an arm approach trajectory 3360, determining an end effector device approach trajectory 3362, and determining a destination approach trajectory 3364 for the robot arm 3320 to approach the multiple objects 3500. FIG. 7A illustrates a motion plan for a transfer cycle of a target object 3510a by the robot arm 3320 and end effector device 3330 from a source (i.e., container 3420) to a destination 3440. A transfer cycle refers to a full cycle of movement by the robot arm 3320 to effect movement of an object from an object source or container to a destination 3440. In an embodiment, the operation 4005 includes determining multiple arm approach trajectories 3360a / 3360b for the robot arm 3320 to approach the multiple objects 3500. FIG. 7B illustrates a motion plan for a transfer cycle of multiple target objects 3511a / 3511b by the robotic arm 3320 and end effector device 3330 from a source (i.e., container 3420) to a destination 3440.

[0096] In operation 4005, the computer system 1100 determines an arm approach trajectory 3360, which includes a path along which the robot arm 3320 is controlled to move or translate in a direction toward the vicinity of the source or container 3420. In determining such an arm approach trajectory 3360, the fastest path (e.g., a path that allows the robot arm 3320 to take the least amount of time to translate from its current position to the vicinity of the source or container 3420) is desired based on factors such as the shortest travel distance from the current location of the robot arm 3320 to the container 3420 and / or the maximum available speed of travel of the robot arm 3320. In determining the maximum available speed of travel, the state of the end effector device 3330 is determined, i.e., whether the end effector device 3330 currently has the target object 3510a or target object 3511a / 3511b in its grip. In an embodiment, the end effector device 3330 is not gripping any target object 3510a or target object 3511a / 3511b, and therefore the maximum speed available to the robot arm 3320 can be used for the arm approach trajectory 3360, since the end effector device 3330 is not gripping any target object 3510a or target object 3511a / 3511b, and therefore the instance of the target object 3510a or target object 3511a / 3511b slipping / falling off the end effector device 3330 is nullified. In an embodiment, the end effector device 3330 can have at least one target object 3510a or target object 3511a / 3511b gripped by its grippers 3332 / 3334, and therefore the travel speed of the robot arm 3320 is calculated taking into account the grip stability of the grippers 3332 / 3334 on the gripped target object 3510a or target object 3511a / 3511b, as will be explained in more detail below.

[0097] In operation 4005, the method 4000 may include determining an end effector device approach trajectory 3362 for the end effector device 3330 to approach the target object 3510a or the target object 3511a / 3511b. The end effector device approach trajectory 3362 may represent an expected path of travel of the end effector device 3330 attached to the robot arm 3320. The computer system 1100 may determine the end effector device approach trajectory 3362, in which the robot arm 3320, the end effector device 3330, or a combination of the robot arm 3320 and the end effector device 3330 are controlled to move or translate in a direction toward the target object 3510a or the target object 3511a / 3511b in the container 3420. In an embodiment, once the robot arm trajectory 3362 is determined such that the robot arm 3320 ends its trajectory at or within the vicinity of the source or container 3420, the end effector device approach trajectory 3362 is determined. The end effector device approach trajectory 3362 may be determined in such a way that the gripper fingers 3332a / 3332b / 3334a / 3334b of the gripper 3332 / 3334 are placed adjacent to the target object 3510a or the target object 3511a / 3511b such that the gripper fingers 3332a / 3332b / 3334a / 3334b of the gripper 3332 / 3334 can properly grip the target object 3510a or the target object 3511a / 3511b in a manner consistent with the determined gripping model 3350a / 3350b / 3350c as described above.

[0098] 7B illustrates another example of a motion plan for a transfer cycle of multiple target objects 3511a / 3511b by the robot arm 3320 and end effector device 3330 from a source or container 3420 to a destination 3440. In an embodiment, the computer system 1100 determines an arm approach trajectory 3360, and the robot arm 3320 is controlled to move or translate in a direction toward the vicinity of the source or container 3420. In determining such an arm approach trajectory 3360, the shortest / fastest path is desired based on factors such as the shortest travel distance from the current location of the robot arm 3320 to the container 3420 and / or the maximum available travel speed of the robot arm 3320. In determining the maximum available travel speed, the state of the end effector device 3330 is determined, i.e., whether the end effector device 3330 currently has the target object 3510a or the target object 3511a / 3511b in its grip. In an embodiment of the trajectory, the end effector device 3330 is not gripping any target object 3510a or target object 3511a / 3511b, and therefore the maximum speed available to the robot arm 3320 can be utilized for the arm approach trajectory 3360 since the end effector device 3330 is not gripping any target object 3510a or target object 3511a / 3511b and therefore the instance of the target object 3510a or target object 3511a / 3511b slipping / falling off the end effector device 3330 is nullified. In another embodiment, the end effector device 3330 may have at least one target object 3510a or target object 3511a / 3511b gripped by its grippers 3332 / 3334 and therefore the traveling speed of the robot arm 3320 is calculated by considering the grip stability of the grippers 3332 / 3334 on the gripped target object 3510a or target object 3511a / 3511b as described in more detail below.

[0099] 7B further illustrates multiple end effector device approach trajectories 3362a / 3362b used to pick or grasp the target object 3511a / 3511b. In an embodiment, the computer system 1100 may determine the end effector device approach trajectories 3362 / 3362a / 3362b in which the robot arm 3320, the end effector device 3330, or a combination of the robot arm 3320 and the end effector device 3330 are controlled to move or translate in a direction toward the target object 3510a or the target object 3511a / 3511b in the source or container 3420. In an embodiment, the end effector device approach trajectory 3362 / 3362a / 3362b is determined once the robot arm trajectory 3362 is determined such that the robot arm 3320 ends its trajectory at or within the vicinity of the source or container 3420. The end effector device approach trajectory 3362 / 3362a / 3362b can be determined in such a manner that the gripper fingers 3332a / 3332b / 3334a / 3334b of the gripper 3332 / 3334 are placed adjacent to the target object 3510a or the target object 3511a / 3511b, so that the gripper fingers 3332a / 3332b / 3334a / 3334b of the gripper 3332 / 3334 can properly grasp the target object 3510a or the target object 3511a / 3511b in a manner consistent with the determined grasping model 3350a / 3350b / 3350c, as described above. The end effector approach trajectory 3362 / 3362a / 3362b may further be determined by the state of the grippers 3332 / 3334, i.e., whether the target object 3510a or the target object 3511a / 3511b is currently gripped by at least one gripper 3332 / 3334. In such a scenario, determining the end effector device approach trajectory 3362 / 3362a / 3362b is based on an optimized end effector device approach time of the end effector device 3330 in a gripping operation to grip the target object 3510a or the target object 3511a / 3511b, the optimized end effector device approach time being the most efficient end effector device approach time determined based on the calculations described below.The optimized end effector device approach time is calculated based on the grip stability of the gripper 3332 / 3334 on the grasped target object 3510a or target object 3511a / 3511b.

[0100] In an embodiment, the optimized end effector device approach time is determined according to the available gripping model 3350a / 3350b / 3350c for the target object 3510a or target object 3511a / 3511b. For example, the amount of time required for the end effector device 3330 to properly position the gripper 3332 / 3334 adjacent to the target object 3510a or target object 3511a / 3511b in a manner that allows the gripper fingers 3332a / 3332b / 3334a / 3334b to properly grip the target object 3510a or target object 3511a / 3511b according to the selected gripping model 3350a / 3350b / 3350c is factored into the optimized end effector device approach time. The time required to properly execute grip model 3350a may be less than or greater than the time required to properly execute grip model 3350b or grip model 3350c. Thus, the grip model 3350a / 3350b / 3350c having the determined minimum time required to properly execute the grip may be selected for target object 3510a or target object 3511a / 3511b to be picked or grasped by gripper 3332 / 3334 of end effector device 3330. The selected gripping model can be selected based on a balancing of factors, for example, by balancing a determined minimum time required to properly perform the grip 3350a / 3350b / 3350c against the predicted grip stability 4016, whereby a faster gripping model 3350a / 3350b / 3350c can be discounted against a second faster gripping model 3350a / 3350b / 3350c in order to sacrifice speed over insufficient predicted grip stability 4016 and reduce the likelihood of grip failure (i.e., dropping, displacing, throwing, or otherwise mishandling the target object 3510a or target object 3511a / 3511b after it has been picked or grasped by the gripper 3332 / 3334 of the end effector device 3330).

[0101] In operation 4005, the method 4000 may further include determining one or more destination approach trajectories 3364 (illustrated in FIG. 7B as destination approach trajectories 3364a and 3364b). In an embodiment, determining the destination trajectory 3364a / 3364b of the robotic arm 3320 may be based on an optimized destination trajectory time for the robotic arm 3320 to travel from the vessel 3420 to one or more destinations 3440. The optimized destination trajectory time may be a determined most efficient destination trajectory time for the robotic arm 3320 to travel from the vessel 3420 to the destination 3440. For example, the optimized trajectory time may be determined by the shortest path between the current location of the robotic arm 3320 (e.g., at or near the vessel 3420) and the destination 3364. The optimized trajectory time may be determined by a path that the robotic arm 3320 may travel fastest without obstacles towards the destination 3364. In an embodiment, determining the destination trajectory 3364 of the robot arm 3320 is based on a predicted grip stability 4016 between the end effector device 3330 and the target object 3510a or the target object 3511a / 3511b. For example, a predicted grip stability 4016 having a higher value may indicate a stronger grip or hold that the gripper fingers 3332a / 3332b / 3334a / 3334b of the gripper 3332 / 3334 may have on the target object 3510a or the target object 3511a / 3511b, which may allow for faster movement of the robot arm 3320 and / or the end effector device 3330 while traversing the destination trajectory 3364 towards the destination 3440.Conversely, a predicted grip stability 4016 having a lower value may indicate a weaker grip or hold that the gripper fingers 3332a / 3332b / 3334a / 3334b of the gripper 3332 / 3334 may have on the target object 3510a or target object 3511a / 3511b, which may therefore require slower movement of the robot arm 3320 and / or end effector device 3330 while traversing the destination trajectory 3364 towards the destination 3440 to prevent a failure scenario, i.e., the target object 3510a / 3511a / 3511b being dropped, thrown, or otherwise displaced.

[0102] In an embodiment, a single destination approach trajectory 3364a may be provided to place both target objects 3511a / 3511b at the same destination 3440. The single destination approach trajectory 3364a may include one or more dechucking or degripping operations to release the target objects 3511a / 3511b. In an embodiment, multiple destination approach trajectories 3364a / 3364b may be determined to place the target objects 3511a / 3511b at either different locations of the same destination 3440 or at two different destinations 3440. A second destination approach trajectory 3364b may be determined to transfer the end effector device 3332 / 3334 between locations within the destination 3440 or between two destinations 3440.

[0103] In operation 4006, the method 4000 includes determining a picking or gripping procedure for grasping or gripping the target object 3510a or 3511a / 3511b with the end effector device 3330 when the end effector device 3330 reaches the target object 3510a or 3511a / 3511b at the end of the end effector device approach trajectory 3362 / 3362a / 3362b. The picking or gripping procedure may represent how the end effector device 3330 approaches, interacts with, contacts, feels, or otherwise grasps the target object 3510a or 3511a / 3511b with the grippers 3332 / 3334. The grasping models 3350a / 3350b / 3350c explain how the target object 3510a or the target object 3511a / 3511b may be grasped by the end effector device 3330. For illustrative purposes, Figures 6A-6C illustrate three different grasping models 3350a / 3350b / 3350c for gripping the target object 3510a or the target object 3511a / 3511b, as detailed above, although it should be understood that other grasping models are possible.

[0104] Determining the gripping behavior may include selecting at least one gripping model 3350a, 3350b, or 3350c from the multiple available gripping models 3350a / 3350b / 3350c to be used by the end effector device 3330 in the gripping behavior determination of operation 4006. In an embodiment, the computer system 1100 determines the gripping behavior based on the gripping model 3350a / 3350b / 3350c having the highest rank. The computer system 1100 may be configured to determine a rank for each of the multiple available gripping models 3350a / 3350b / 3350c according to the predicted grip stability 4016 of each of the multiple gripping models 3350a / 3350b / 3350c. Each of the grasp models 3350a / 3350b / 3350c can be ranked according to factors such as a predicted grip stability 4016, which may have an associated transport speed modifier that can determine the speed, acceleration, and / or deceleration at which the target object 3510a or target object 3511a / 3511b may be moved by the robot arm 3320 during execution of the arm approach trajectory 3360 and / or the end effector device approach trajectory 3362. The predicted grip stability 4016 may also be an indication of how the target object 3510a or target object 3511a / 3511b will be secured once picked or grasped by the end effector device 3330. Generally, the stronger the predicted grip stability 4016, or the ability of the end effector device 3330 to hold the target object 3510a / 3511a / 3511b, the greater the probability that the robot 3300 will be able to move the robot arm 3320 and / or end effector device 3330 through the determined arm approach trajectory 3360 and / or end effector approach trajectory 3362 / 3362a / 3362b while holding / grasping the target object 3510a / 3511a / 3511b without resulting in a failure scenario, i.e., the target object 3510a / 3511a / 3511b being dropped, thrown, or otherwise displaced from the gripper.

[0105] In an embodiment determining the rank of each of the grip models 3350a / 3350b / 3350c, the computer system 1100 may determine that grip model 3350a is likely to have a higher predicted grip stability 4016 than grip model 3350b, which is likely to have a higher predicted grip stability 4016 than grip model 3350c. As another example, the detected object 3510 may be inaccessible for grasping by at least one of the grasping models 3350a / 3350b / 3350c based on the multiple object representations 4013 corresponding to the detected object 3510 (i.e., at least one of the detected objects 3510 is in a location or orientation or shape that does not permit effective use of the particular grasping model 3350a / 3350b / 3350c) and therefore cannot be picked up by the end effector device 3330 via the grasping model 3350a / 3350b / 3350c at that time. In such a scenario, the remaining grasping models 3350a / 3350b / 3350c are measured for predicted grip stability 4016. For example, the grip model 3350a may not have the target object 3510a or the target object 3511a / 3511b available as a choice to pick or grasp, e.g., based on the previously determined plurality of object representations 4013. Thus, the grip model 3350a may receive the lowest possible rank value, an empty rank value, or no rank at all (i.e., completely ignored). Thus, the predicted grip stability 4016 of the grip model 3350a may be excluded when calculating the rank to apply during the grip action determination of operation 4006. For example, if the predicted grip stability 4016 of the grip model 3350b is determined to have a higher value than the predicted grip stability of the grip model 3350c, then the grip model 3350b will receive a higher value rank, while the grip model 3350c will receive a lower value rank (but still higher than the grip model 3350a).In other embodiments, inaccessible ones of the gripping models 3350a / 3350b / 3350c may be included in the ranking procedure but may be assigned the lowest rank.

[0106] In an embodiment, determining at least one grasp model 3350a / 3350b / 3350c for use by the end effector device 3330 is based on the rank of the grasp model having the highest determined value of predicted grip stability 4016. The rank of grasp model 3350a may have a predicted grip stability 4016 having a higher value than the rank of grasp model 3350b and / or 3350c, and thus grasp model 3350a may be ranked higher than grasp model 3350b and / or 3350c. To maximize or optimize the speed of transfer of target object 3510a / 3511a / 3511b within each transfer cycle, computer system 1100 may select target objects 3510a / 3511a / 3511b having similar predicted grip stability 4016. In an embodiment, the computer system 1100 can select multiple target objects 3511a / 3511b that have the same grip model 3350a / 3350b / 3350c. The computer system 1100 can calculate a motion plan for a transfer cycle while gripping the target objects 3511a / 3511b based on the detection result 3520. The objective is to reduce the calculation time for picking multiple target objects 3511a / 3511b in the source container 3420 while optimizing the transfer speed between the source container 3420 and the destination 3440. In this way, the robot 3300 can transfer both target objects 3511a / 3511b at maximum speed since both target objects 3511a / 3511b have the same predicted grip stability 4016.Conversely, the computer system 1100 may use the gripping model 3350a / 3350b / 3350c having a higher rank (i.e., a higher judgement value of the predicted grip stability 4016) to determine the target object 3510a / 3511a / 3511b to be grasped by the gripper 3332 / 3334 of the end effector device 3330, and the gripping model 3350a / 3350b / 3350c having a lower rank (i.e., a lower judgement value of the predicted grip stability 4016) to determine the target object 3510a / 3511a / 3511b to be grasped by the gripper 3332 / 3334 of the end effector device 3330. If 3350a / 3350b / 3350c is used to select a second target object 3510a / 3511a / 3511b to be grasped by the gripper 3332 / 3334 of the end effector device 3330, the speed of transfer is limited or upper bounded by the lower predicted grip stability 4016 of the target object 3510a / 3511a / 3511b having the lower ranked grasping model 3350a / 3350b / 3350c. That is, for successive transport cycles, selecting two target objects 3511a / 3511b with gripping models 3350a / 3350b / 3350c with higher predicted grip stability 4016 and higher transport speed, and then selecting two target objects 3511a / 3511b with gripping models 3350a / 3350b / 3350c with lower predicted grip stability 4016 and lower transport speed, is more optimal than successive transport cycles including one target object 3511a with gripping models 3350a / 3350b / 3350c with higher predicted grip stability 4016 and one target object 3511b with gripping models 3350a / 3350b / 3350c with lower predicted grip stability 4016, since both transport cycles are limited to the slower transport speed in the later scenario.

[0107] The various trajectory determinations of operation 4005 and grasp operation determination 4006 are described sequentially with respect to operations of method 4000. Where suitable and appropriate, it will be understood that the various operations of method 4000 may occur simultaneously with one another or in different orders presented below. For example, trajectory determinations (such as destination approach trajectory 3364) may be made during the execution of other trajectories. Thus, target approach trajectory 3364 may be determined during the execution of arm approach trajectory 3362.

[0108] In operation 4008, the method 4000 may include outputting a first command (e.g., an arm approach command) to control the robotic arm 3300 in the arm approach trajectory 3360 to approach the plurality of objects 3500. As illustrated in FIG. 7B, the computer system 1100 may output the first command to control the robotic arm 3320 from an area outside the vicinity of the source or container 3420 to a location at or within the vicinity of the source or container 3420. The first command may control the robotic arm 3320 to move from an area at or near the destination 3440 to a location at or within the vicinity of the source or container 3420. In operation 4008, the method 4000 may include outputting a second command (e.g., an end effector device approach command) to control the robot arm 3320 in the end effector device approach trajectory 3362 to approach the target object 3510a / 3511a / 3511b (e.g., to cause the end effector device 3330 to approach the target object 3510a / 3511a / 3511b). As illustrated in FIG. 7B, the end effector device approach trajectory 3362a / 3362b can be used to approach multiple target objects 3511a / 3511b.

[0109] In operation 4010, the method 4000 includes outputting a third command (e.g., an end effector device control command) to control the end effector device 3330 in a grasping operation to grasp the target object 3510a or the target object 3511a / 3511b. The end effector device 3330 can grasp the target object 3510a / 3511a / 3511b using the gripping fingers 3332a / 3332b / 3334a / 3334b of the gripper 3332 / 3334 using the gripping model 3350a / 3350b / 3350c previously determined to have the highest ranked and / or predicted grip stability 4016. The gripping fingers 3332a / 3332b / 3334a / 3334b can be controlled to move or translate in a manner consistent with a predetermined gripping model 3350a / 3350b / 3350c when the end effector device 3330 contacts the target object 3510a / 3511a / 3511b.

[0110] In operation 4012, the method 4000 may further include executing the destination trajectory 3364 to control the robot arm 3320 to approach the destination. Operation 4012 may include outputting a fourth command (e.g., a robot arm control command) to control the robot arm 3320 in the destination trajectory 3364. In an embodiment, the destination trajectory 3364 may be determined during the trajectory determination operation 4005 described above. In an embodiment, the destination trajectory 3364 may be determined after the trajectory execution operation 4008 and the end effector interaction operation 4010. In an embodiment, the destination trajectory 3364 may be determined by the computer system 1100 at any time prior to execution of the destination trajectory 3364, including during performance of other operations. In an embodiment, operation 4012 may further include outputting a fifth command (e.g., an end effector device release command) to control the end effector device 3330 to release, unclasp, or de-chuck the target object 3510a or target objects 3511a / 3511b within or at the destination 3440 when the robot arm 3320 and end effector device 3330 reach the destination 3440 at the end of the destination trajectory 3364.

[0111] At a high level, the motion plan for a transfer cycle of target object 3510a or target object 3511a / 3511b by the robot arm 3320 from a source container 3420 to a destination 3440 involves the operations illustrated in FIG. 7A , picking target object 3510a or target object 3511a / 3511b from a source container 3420 3420 location, transporting target object 3510a or target object 3511a / 3511b to a destination 3440 location, placing target object 3510a or target object 3511a / 3511b at the destination 3440 location, and returning to the source container 3420 location. The overall transfer cycle time is upper bounded by the transfer of the target object 3510a or target object 3511a / 3511b from the source container 3420 3420 to the destination 3440 due to the predicted gripping stability 4016 of the target object 3510a or target object 3511a / 3511b by the end effector device 3330 on the robot arm 3320.

[0112] In general, the method 4000 described herein may be used to manipulate (e.g., move and / or reorient) a target object (e.g., one of a package, box, case, cage, pallet, etc. corresponding to a task to be performed) from a start / source location to a task / destination location. For example, a loading unit (e.g., a debunking robot) may be configured to transfer a target object from a location in a carrier (e.g., a truck) to a location on a conveyor. Also, a transfer unit may be configured to transfer a target object from one location (e.g., a conveyor, a pallet, or a bin) to another location (e.g., a pallet, a bin, etc.). As another example, a transfer unit (e.g., a palletizing robot) may be configured to transfer a target object from a source location (e.g., a pallet, a picking area, and / or a conveyor) to a destination pallet. Upon completion of the operation, a transport unit (e.g., a conveyor, an automated guided vehicle (AGV), a shelf-transporting robot, etc.) can transfer the target object from an area associated with the transfer unit to an area associated with the loading unit, and the loading unit can transfer the target object from the transfer unit (e.g., by moving a pallet carrying the target object) to a storage location (e.g., a location on a shelf). More details regarding the tasks and associated actions are provided above.

[0113] For illustrative purposes, the computer system 1100 system is described in the context of a packaging and / or shipping center, however, it is understood that the computer system 1100 can be configured to perform tasks in other environments / purposes, such as for manufacturing, assembly, storage / inventory, medical, and / or other types of automation. It is also understood that the computer system 1100 can include other units (not shown), such as manipulators, service robots, modular robots, etc. For example, in some embodiments, the computer system 1100 can include a depalletizing unit for transferring objects from a cage cart or pallet onto a conveyor or other pallet, a bin switching unit for transferring objects from one bin to another, a packaging unit for wrapping / storing objects, a sorting unit for grouping objects according to one or more characteristics thereof, a piece pick unit for differently manipulating (e.g., sorting, grouping, and / or transporting) objects according to one or more characteristics thereof, or combinations thereof, etc.

[0114] It will be apparent to those skilled in the relevant art that other suitable modifications and adaptations to the methods and applications described herein can be made without departing from the scope of any of the embodiments. The embodiments described above are illustrative examples, and the disclosure should not be construed as being limited to these particular embodiments. It should be understood that the various embodiments disclosed herein may be combined in different combinations than those specifically presented in the description and accompanying figures. It should also be understood that, by way of example, certain acts or events of any of the processes or methods described herein may be performed in a different sequence, or may be added, integrated, or omitted entirely (e.g., not all acts or events described may be required to perform a method or process). In addition, although certain features of the embodiments herein are described for clarity as being implemented by a single component, module, or unit, it should be understood that the features and functions described herein may be implemented by any combination of components, units, or modules. Thus, various changes and modifications may be effected by those skilled in the art without departing from the spirit or scope of the invention, as defined in the appended claims.

[0115] Further embodiments include the following: Embodiment 1 is a computing system including a control system configured to communicate with a robot having a robot arm including an end effector device or attached to the end effector device and to communicate with a camera, and at least one processing circuit, the control system configured to: when the robot is in an object handling environment including a source of objects for transfer to a destination in the object handling environment, identify a target object from among a plurality of objects in the source of objects for transferring the target object from the source of objects to a destination in the object handling environment; generate an arm approach trajectory for the robot arm to approach the plurality of objects; and generate an arm approach trajectory for the end effector device to approach the target object. and at least one processing circuit configured to perform the following: generating an end effector device approach trajectory; generating a grasping operation for grasping a target object with the end effector device; outputting an arm approach command to control a robot arm according to the arm approach trajectory to approach a plurality of objects; outputting an end effector device approach command to control the robot arm within the end effector device approach trajectory to approach the target object; and outputting an end effector device control command to control the end effector device in a grasping operation to grasp the target object. Embodiment 2 is a computer system of embodiment 1, further including generating a destination trajectory for the robot arm to approach the destination, outputting a robot arm control command to control the robot arm according to the destination trajectory, and outputting an end effector device release command to control the end effector device to release the target object at the destination. Example 3 is the computer system of Example 2, wherein determining the destination trajectory of the robot arm is based on an optimized destination trajectory time for the robot arm to travel from the source to the destination. Example 4 is the computer system of example 2, wherein determining the destination trajectory of the robot arm is based on a predicted grip stability between the end effector device and the target object. Example 5 is the computer system of example 1, wherein determining the end effector device approach trajectory is based on an optimized end effector device approach time for the end effector device to grasp the target object in a grasping operation. Example 6 is the computer system of example 5, wherein the optimized end effector device approach time is determined based on an available grasp model for the target object. Example 7 is the computer system of example 1, wherein determining the grasp operation includes determining at least one grasp model from a plurality of available grasp models for use by the end effector device in the grasp operation. Example 8 is the computer system of example 7, wherein the at least one processing circuit is further configured to determine a rank for each of the plurality of available grasp models according to a predicted grip stability of each of the plurality of grasp models. Example 9 is the computer system of example 8, wherein determining at least one grasp model for use by the end effector device is based on a rank having the highest judgement value of predicted grip stability. Embodiment 10 is the computer system of embodiment 1, wherein at least one processing circuit is further configured for generating one or more detection results, each of which represents a detected object of one or more objects in the source of objects and includes a corresponding object representation that defines at least one of an object orientation of the detected object, a location of the detected object within the source of objects, a location of the detected object relative to other objects, and a confidence determination. Example 11 is the computer system of example 1, wherein the objects are substantially identical in terms of size, shape, weight, and material composition. Example 12 is the computer system of example 1, wherein the objects vary from one another in size, shape, weight, and material composition. Example 13 is the computer system of example 10, wherein identifying the target object from the one or more detection results includes determining whether there is a gripping model available for the detected object, and removing from the detected objects any detected objects without a gripping model available. Example 14 is the computer system of example 13, further comprising removing the detected object based on at least one of the object orientation, the location of the detected object within the source of the object, and / or the object distance. Example 15 is the computer system of example 1, wherein the at least one processing circuit is further configured to identify a plurality of target objects including the target object from the detection result. Example 16 is a computer system of example 15, wherein the target object is a first target object of a plurality of target objects associated with a first grasping model, and a second target object of the plurality of target objects is associated with a second grasping model. Example 17 is a computer system of example 15, wherein identifying the multiple target objects includes selecting a first target object for grasping by the end effector device and a second target object for grasping by the end effector device. Example 18 is the computer system of Example 17, wherein at least one processing circuit is further configured to output a second end effector device approach command to control the robot arm to approach the second target object, output a second end effector device control command to control the end effector device to grasp the second target object, generate a destination trajectory for the robot arm to approach the destination, output a robot arm control command to control the robot arm according to the destination trajectory, and output an end effector device release command to control the end effector device to release the first target object and the second target object at the destination. Embodiment 19 is a method for selecting a target object from a source of objects, the method including: identifying the target object from among a plurality of objects in the source of objects; generating an arm approach trajectory for a robot arm having an end effector device to approach the plurality of objects; generating an end effector device approach trajectory for the end effector device to approach the target object; generating a grasping operation for grasping the target object with the end effector device; outputting an arm approach command for controlling the robot arm according to the arm approach trajectory to approach the plurality of objects; outputting an end effector device approach command for controlling the robot arm according to the end effector device approach trajectory to approach the target object; and outputting an end effector device control command for controlling the end effector device in a grasping operation to grasp the object. Embodiment 20 is a non-transitory computer-readable medium operable by at least one processing circuit via a communications interface configured to communicate with a robotic system, the non-transitory computer-readable medium being configured with executable instructions for implementing a method for picking a target object from a source of objects, the method including: identifying the target object from among a plurality of objects in the source of objects; generating an arm approach trajectory for a robot arm having an end effector device to approach the plurality of objects; generating an end effector device approach trajectory for the end effector device to approach the target object; generating a grasping operation for grasping the target object with the end effector device; outputting an arm approach command for controlling the end effector device in the arm approach trajectory to approach the plurality of objects; outputting an end effector device approach command for controlling the robot arm in the end effector device approach trajectory to approach the target object; and outputting an end effector device control command for controlling the end effector device in a grasping operation to grasp the object.

Claims

1. A computing system, A control system configured to communicate with a robot having an end effector device or a robot arm attached to the end effector device, and to communicate with a camera, It comprises at least one processing circuit, The at least one processing circuit, when the robot is in an object handling environment which includes a source of objects to be transported to a destination in the object handling environment, transports a target object from the source of objects to the destination. Identifying a graspable object, including a primary candidate object and a secondary candidate object, from among multiple objects in the source of the object. Identifying the target object from the graspable object, including removing the secondary candidate object if it is within the interference range of the primary candidate object. In order for the robot arm to approach the plurality of objects, an arm approach trajectory is generated. In order for the end effector device to approach the target object, the end effector device generates an approach trajectory. The end effector device generates a gripping motion for gripping the target object. To control the robot arm according to the arm approach trajectory and approach the plurality of objects, an arm approach command is output. To control the robot arm within the approach trajectory of the end effector device and approach the target object, an end effector device approach command is output, and In the gripping operation, the end effector device is controlled to grip the target object, and an end effector device control command is output. A computing system configured to perform [a specific action].

2. The robot arm generates a destination trajectory for approaching the destination, In order to control the robot arm according to the aforementioned destination trajectory, a robot arm control command is output, The computing system according to claim 1, further comprising: controlling the end effector device to release the target object at the destination, and outputting an end effector device release command.

3. The computing system according to claim 2, wherein determining the destination trajectory of the robot arm is based on a destination trajectory time optimized for the robot arm to move from the supply source to the destination.

4. The computing system according to claim 2, wherein the destination trajectory of the robot arm is determined based on the predicted grip stability between the end effector device and the target object.

5. The computing system according to claim 1, wherein determining the approach trajectory of the end effector device is based on an end effector device approach time optimized for the end effector device to grasp the target object in the gripping operation.

6. The computing system according to claim 5, wherein the optimized end effector device approach time is determined based on an available gripping model for the target object.

7. The computing system according to claim 1, wherein determining the gripping operation includes determining at least one gripping model from a plurality of available gripping models for use by the end effector device in the gripping operation.

8. The computing system according to claim 7, wherein the at least one processing circuit is further configured to determine a rank for each of the plurality of available gripping models according to the predicted grip stability of each of the plurality of gripping models.

9. The computing system according to claim 8, wherein determining the at least one gripping model for use by the end effector device is based on the rank having the highest determination value of the predicted grip stability.

10. The at least one processing circuit, Each is further configured to generate one or more detection results representing one or more detected objects among the one or more objects in the source of the object, The computing system according to claim 1, wherein each of the one or more detection results includes a corresponding object representation that defines at least one of the object orientation of the detected object, the location of the detected object within the source of the object, the location of the detected object relative to other objects, and a confidence determination.

11. The computing system according to claim 1, wherein the plurality of objects are substantially identical in terms of size, shape, weight and material composition.

12. The computing system according to claim 1, wherein the plurality of objects differ from each other in size, shape, weight, and material composition.

13. Identifying the target object from one or more of the above detection results, To determine whether there is an available gripping model for the detected object, The computing system according to claim 10, comprising removing the detected object from the detected object without an available gripping model.

14. The computing system according to claim 13, further comprising removing the detected object based on at least one of the object orientation, the location of the detected object within the source of the object, and / or the distance between objects.

15. The computing system according to claim 1, wherein the at least one processing circuit is further configured to identify a plurality of target objects, including the target object, from the detection results.

16. The target object is the first target object of the plurality of target objects associated with the first gripping model, The computing system according to claim 15, wherein a second target object of the plurality of target objects is associated with a second gripping model.

17. The computing system according to claim 15, wherein identifying the plurality of target objects includes selecting the first target object for gripping by the end effector device and the second target object for gripping by the end effector device.

18. The at least one processing circuit, To control the robot arm and approach the second target object, a second end effector device approach command is output. To control the end effector device and grasp the second target object, a second end effector device control command is output, and a destination trajectory is generated for the robot arm to approach the destination. In order to control the robot arm according to the aforementioned destination trajectory, a robot arm control command is output, The computing system according to claim 17, further configured to output an end effector release command in order to control the end effector device to release the first target object and the second target object at the destination.

19. A method for selecting a target object from a source of objects, Identifying a graspable object, including a primary candidate object and a secondary candidate object, among multiple objects in the source of the object, Identifying the target object from the graspable objects, including removing the secondary candidate object if it is within the interference range of the primary candidate object; In order for a robot arm having an end effector device to approach the plurality of objects, an arm approach trajectory is generated, In order for the end effector device to approach the target object, the end effector device generates an approach trajectory, The end effector device generates a gripping motion for gripping the target object, To control the robot arm according to the arm approach trajectory and approach the plurality of objects, an arm approach command is output. To control the robot arm within the approach trajectory of the end effector device and approach the target object, an end effector device approach command is output. A method comprising outputting an end effector device control command in order to control the end effector device to grasp the target object during the gripping operation.

20. A non-temporary computer-readable medium having executable instructions, which is operable by at least one processing circuit via a communication interface configured to communicate with a robot system, The aforementioned instruction is for implementing a method for selecting a target object from a source of objects, The aforementioned method, Identifying a graspable object, including a primary candidate object and a secondary candidate object, from among multiple objects in the source of the object, Identifying the target object from the graspable objects, including removing the secondary candidate object if it is within the interference range of the primary candidate object; In order for a robot arm having an end effector device to approach the plurality of objects, an arm approach trajectory is generated, In order for the end effector device to approach the target object, the end effector device generates an approach trajectory, The end effector device generates a gripping motion for gripping the target object, To control the robot arm according to the arm approach trajectory that approaches the aforementioned multiple objects, an arm approach command is output, To control the robot arm along the approach trajectory of the end effector device as it approaches the target object, an end effector device approach command is output. A non-temporary computer-readable medium that includes outputting an end-effector device control command in order to control the end-effector device to grasp the target object in the gripping operation.