Systems and methods for object detection

A computing system enhances robotic object retrieval by identifying and resolving occlusions in irregularly stacked objects through cost map segmentation, improving efficiency and accuracy in robotic tasks.

WO2026053176A1PCT designated stage Publication Date: 2026-03-12MUJIN INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2026-03-12

Smart Images

  • Figure IB2025059023_12032026_PF_FP_ABST
    Figure IB2025059023_12032026_PF_FP_ABST
Patent Text Reader

Abstract

A computing system configured for identifying occlusions among objects in a scene is provided. The computing system includes at least one processing circuit configured to generate a cost map indicating surface variations among objects in the scene. The cost map may be segmented to identify individual objects within the scene. Objects within the scene may be compared to neighboring objects to identify occlusions.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR OBJECT DETECTIONCross-Reference to Related Application(s)

[0001] The present application claims the benefit of U.S. Provisional Appl. No. 63 / 692,231, filed September 9, 2024, and U.S. Provisional Appl. No. 63 / 692,226, filed September 9, 2024, the entire contents of which are incorporated by reference herein.Field of the Invention

[0002] The present technology is directed generally to robotic systems and, more specifically, to systems, processes, and techniques for identifying and detecting objects. More particularly, the present technology may be used for identifying occluded objects in a scene containing multiple objects.Background

[0003] With their ever- increasing performance and lowering cost, many robots (e.g., machines configured to automatically / autonomously execute physical actions) are now extensively used in various different fields. Robots, for example, can be used to execute various tasks (e.g., manipulate or transfer an object through space) in manufacturing and / or assembly, packing and / or packaging, transport and / or shipping, etc. In executing the tasks, the robots can replicate human actions, thereby replacing or reducing human involvements that are otherwise required to perform dangerous or repetitive tasks.

[0004] However, despite the technological advancements, robots often lack the sophistication necessary to duplicate human interactions required for executing larger and / or more complex tasks. Accordingly, there remains a need for improved techniques and systems for managing operations and / or interactions between robots.Brief Summary

[0005] In some aspects, the techniques described herein relate to a computing system configured to identify object occlusions within a scene including: a control system configured to communicate with a robot and to communicate with a camera; at least one processing circuit configured for: obtaining image information captured by the camera of a plurality of objects in the scene; identifying and removing a background portion of the image information; generating a cost map based on the image information, the cost map being indicative of surface variations among the objects in the scene; segmenting the cost map according to a comparison betweencost map values and a reference cost threshold; and identifying object occlusions between segmented pol gons of the cost map.

[0006] In some aspects, the techniques described herein relate to a method of identifying object occlusions within a scene performed by a control system having at least one processing circuit and being configured to communicate with a robot and to communicate with a camera, the method including: obtaining image information captured by the camera of a plurality of objects in the scene; identifying and removing a background portion of the image information; generating a cost map based on the image information, the cost map being indicative of surface variations among the objects in the scene; segmenting the cost map according to a comparison between cost map values and a reference cost threshold; and identifying object occlusions between segmented polygons of the cost map.Brief Description of the Figures

[0007] FIG. 1A illustrates a system for performing or facilitating the detection, identification, and retrieval of objects according to embodiments hereof.

[0008] FIG. IB illustrates an embodiment of the system for performing or facilitating t the detection, identification, and retrieval of objects according to embodiments hereof.

[0009] FIG. 10 illustrates another embodiment of the system for performing or facilitating the detection, identification, and retrieval of objects according to embodiments hereof.

[0010] FIG. ID illustrates yet another embodiment of the system for performing or facilitating the detection, identification, and retrieval of objects according to embodiments hereof.

[0011] FIG. 2A is a block diagram that illustrates a computing system configured to perform or facilitate the detection, identification, and retrieval of objects, consistent with embodiments hereof.

[0012] FIG. 2B is a block diagram that illustrates an embodiment of a computing system configured to perform or facilitate the detection, identification, and retrieval of objects, consistent with embodiments hereof.

[0013] FIG. 2C is a block diagram that illustrates another embodiment of a computing system configured to perform or facilitate the detection, identification, and retrieval of objects, consistent with embodiments hereof.

[0014] FIG. 2D is a block diagram that illustrates yet another embodiment of a computing system configured to perform or facilitate the detection, identification, and retrieval of objects, consistent with embodiments hereof.

[0015] FIG. 2E is an example of image information processed by systems and consistent with embodiments hereof.

[0016] FIG. 2F is another example of image information processed by systems and consistent with embodiments hereof.

[0017] FIG. 3A illustrates an exemplary environment for operating a robotic system, according to embodiments hereof.

[0018] FIG. 3B illustrates an exemplary environment for the detection, identification, and retrieval of objects by a robotic system, consistent with embodiments hereof.

[0019] FIG. 4 provides a flow diagram illustrating an overall flow of methods and operations for the detection, identification, and retrieval of objects, according to embodiments hereof.

[0020] FIG. 5 illustrates a method of occlusion determination consistent with embodiments hereof.

[0021] FIG. 6 illustrates a subprocess for removing a background from a captured image, according to embodiments hereof.

[0022] FIG. 7A-7C illustrate the operations of the subprocess for removing a background from a captured image, according to embodiments hereof.

[0023] FIG. 7D illustrates a background masking process.

[0024] FIG. 8 illustrates a subprocess for generating a cost map, according to embodiments hereof.

[0025] FIGS. 9 A and 9B are provided to better illustrate the operations of the subprocess for generating a cost map, according to embodiments hereof.

[0026] FIG. 10 illustrates a subprocess method for sharpening gradient angle magnitudes, according to embodiments hereof..

[0027] FIG. 11 illustrates a subprocess for obtaining image data, according to embodiments hereof..

[0028] FIG. 12 illustrates a subprocess for computing gradient angles and magnitudes, according to embodiments hereof..

[0029] FIG. 13 illustrates a subprocess for weighting gradient angle magnitudes, according to embodiments hereof..

[0030] FIG. 14A illustrates an example image showing image portions satisfying the threshold.

[0031] FIG. 14B illustrates an image having sharpened edges as a result of the gradient angle magnitude weighting described herein.

[0032] FIG. 15 illustrates a subprocess for sharpening gradients, according to embodiments hereof.

[0033] FIGS. 16A-16F illustrate the subprocess for sharpening gradients, according to embodiments hereof.

[0034] FIG. 17 illustrates a subprocess for cost map segmentation, according to embodiments hereof.

[0035] FIGS. 18A- 18D provide illustrations related to cost map segmentation, according to embodiments hereof.

[0036] FIG. 19 illustrates a subprocess for identifying occluded objects, according to embodiments hereof.

[0037] FIGS. 20A-20I provide illustrative support for the subprocess for identifying occluded objects.Detailed Description

[0038] Systems and methods related to object detection and occlusion determination are described herein. The disclosed systems and methods may facilitate object detection, identification, and retrieval where the objects are located amongst other objects. As discussed herein, a group of commingled objects may be referred to as a “stack.” In particular, the disclosed systems and methods may operate to identify occlusions between objects. Identifying such occlusions may be particularly advantageous when operating a robotic system to retrieve objects that are commingled. Even where a particular occluded object may be graspable and moveable by the robotic system, moving such an object is likely to cause other objects within the stack to move. Such movement of the other objects may require the stack to be imaged and processed again to permit further retrieval. Operating a system to retrieve unoccluded orminimally occluded objects may permit object retrieval without disturbing the stack, thereby allowing the system to continue to use previously computed information for retrieval. Thus, the system may operate more efficiently.

[0039] As discussed herein, objects in a stack may be of any material and may be located in containers such as boxes, bins, crates, etc., on pallets, on conveyors, and / or on or in any other suitable surface or container. The container or surface where the objects are located may be referred to as a retrieval location. The objects may be situated in the retrieval location in an unorganized or irregular fashion, for example, a box full various envelopes, packages, and letters. Object detection, identification, and retrieval in such circumstances may be challenging due to the irregular arrangement of the objects and the tendency for irregularly placed objects to be stacked atop one another and thereby occlude one another.

[0040] Robotic systems configured in accordance with embodiments hereof may autonomously execute integrated tasks by coordinating operations of multiple robots. Robotic systems, as described herein, may include any suitable combination of robotic devices, actuators, sensors, cameras, and computing systems configured to control, issue commands, receive information from robotic devices and sensors, access, analyze, and process data generated by robotic devices, sensors, and camera, generate data or information usable in the control of robotic systems, and plan actions for robotic devices, sensors, and cameras. As used herein, robotic systems are not required to have immediate access or control of robotic actuators, sensors, or other devices. Robotic systems, as described here, may be computational systems configured to improve the performance of such robotic actuators, sensors, and other devices through reception, analysis, and processing of information.

[0041] The technology described herein provides technical improvements to a robotic system configured for use in object identification, detection, and retrieval, and specifically related to occlusion detection. Technical improvements described herein increase the speed, precision, and accuracy of retrieval tasks from a retrieval location by identifying occlusions between objects. The robotic systems and computational systems described herein address the technical problem of identifying overlap and other types of occlusions between objects that may be irregularly arranged. By addressing this technical problem, the technology of object identification, detection, and retrieval is improved.

[0042] The present application refers to systems and robotic systems. Robotic systems, as discussed herein, may include robotic actuator components (e.g., robotic arms, roboticgrippers, etc.), various sensors (e.g., cameras, etc.), and various computing or control systems. As discussed herein, computing systems or control systems may be referred to as “controlling” various robotic components, such as robotic arms, robotic grippers, cameras, etc. Such “control” may refer to direct control of and interaction with the various actuators, sensors, and other functional aspects of the robotic components. For example, a computing system may control a robotic arm by issuing or providing all of the required signals to cause the various motors, actuators, and sensors to cause robotic movement. Such “control” may also refer to tire issuance of abstract or indirect commands to a further robotic control system that then translates such commands into the necessary signals for causing robotic movement. For example, a computing system may control a robotic arm by issuing a command describing a trajectory or destination location to which the robotic arm should move to and a further robotic control system associated with the robotic arm may receive and interpret such a command and then provide the necessary direct signals to the various actuators and sensors of the robotic arm to cause the required movement.

[0043] In particular, the present technology described herein assists a robotic system to interact with a target object among a plurality of objects in a container, conveyor belt, or other retrieval location. Specifically, by recognizing occlusions and overlap between objects, the systems and methods described herein are able to facilitate the efficient retrieval of objects from an irregular stack.

[0044] In the following, specific details are set forth to provide an understanding of the presently disclosed technology. In embodiments, the techniques introduced here may be practiced without including each specific detail disclosed herein. In other instances, well- known features, such as specific functions or routines, are not described in detail to avoid unnecessarily obscuring the present disclosure. References in this description to “an embodiment,” “one embodiment,” or the like mean that a particular feature, structure, material, or characteristic being described is included in at least one embodiment of the present disclosure. Thus, the appearances of such phrases in this specification do not necessarily all refer to the same embodiment. On the other hand, such references are not necessarily mutually exclusive either. Furthermore, the particular features, structures, materials, or characteristics described with respect to any one embodiments can be combined in any suitable manner with those of any other embodiment, unless such items are mutually exclusive. It is to be understood that the various embodiments shown in the figures are merely illustrative representations and are not necessarily drawn to scale.

[0045] Several details describing structures or processes that are well-known and often associated with robotic systems and subsystems, but that can unnecessarily obscure some significant aspects of the disclosed techniques, are not set forth in the following description for purposes of clarity. Moreover, although the following disclosure sets forth several embodiments of different aspects of the present technology, several other embodiments may have different configurations or different components than those described in this section. Accordingly, tire disclosed techniques may have other embodiments with additional elements or without several of the elements described below.

[0046] Many embodiments or aspects of the present disclosure described below may take the form of computer- or controller-executable instructions, including routines executed by a programmable computer or controller. Those skilled in the relevant art will appreciate that the disclosed techniques can be practiced on or with computer or controller systems other than those shown and described below. The techniques described herein can be embodied in a special-purpose computer or data processor that is specifically programmed, configured, or constructed to execute one or more of the computer-executable instructions described below. Accordingly, the terms "computer" and "controller" as generally used herein refer to any data processor and can include Internet appliances and handheld devices (including palm-top computers, wearable computers, cellular or mobile phones, multi-processor systems, processor-based or programmable consumer electronics, network computers, minicomputers, and the like). Information handled by these computers and controllers can be presented at any suitable display medium, including a liquid crystal display (LCD). Instructions for executing computer- or controller-executable tasks can be stored in or on any suitable computer-readable medium, including hardware, firmware, or a combination of hardware and firmware. Instructions can be contained in any suitable memory device, including, for example, a flash drive, USB device, and / or other suitable medium.

[0047] The terms “coupled” and “connected,” along with their derivatives, can be used herein to describe structural relationships between components. It should be understood that these terms are not intended as synonyms for each other. Rather, in particular embodiments, “connected” can be used to indicate that two or more elements are in direct contact with each other. Unless otherwise made apparent in the context, the term “coupled” can be used to indicate that two or more elements are in either direct or indirect (with other intervening elements between them) contact with each other, or that the two or more elements co-operateor interact with each other (e.g., as in a cause-and-effect relationship, such as for signal transmission / reception or for function calls), or both.

[0048] Any reference herein to image analysis by a computing system may be performed according to or using spatial structure information that may include depth information which describes respective depth value of various locations relative a chosen point. The depth information may be used to identify objects or estimate how objects are spatially arranged. In some instances, the spatial structure information may include or may be used to generate a point cloud that describes locations of one or more surfaces of an object. Spatial structure information is merely one form of possible image analysis and other forms known by one skilled in the art may be used in accordance with the methods described herein.

[0049] FIG. 1A illustrates a system 1000 for performing object detection, or, more specifically, object recognition. More particularly, the system 1000 may include a computing system 1100 and a camera 1200. In this example, the camera 1200 may be configured to generate image information which describes or otherwise represents an environment in which tire camera 1200 is located, or, more specifically, represents an environment in the camera’s 1200 field of view (also referred to as a camera field of view). The environment may be, e.g., a warehouse, a manufacturing plant, a retail space, or other premises. In such instances, the image information may represent objects located at such premises, such as boxes, bins, cases, crates, pallets, or other containers. The system 1000 may be configured to generate, receive, and / or process the image information, such as by using the image information to distinguish between individual objects in the camera field of view, to perform object recognition or object registration based on the image information, and / or perform robot interaction planning based on the image information, as discussed below in more detail (the terms “and / or” and “or” are used interchangeably in this disclosure). The robot interaction planning may be used to, e.g., control a robot at the premises to facilitate robot interaction between the robot and the containers or other objects. The computing system 1100 and the camera 1200 may be located at the same premises or may be located remotely from each other. For instance, the computing system 1100 may be part of a cloud computing platform hosted in a data center which is remote from the warehouse or retail space and may be communicating with the camera 1200 via a network connection.

[0050] In an embodiment, the camera 1200 (which may also be referred to as an image sensing device) may be a 2D camera and / or a 3D camera. For example, FIG. IB illustrates a system 1500A (which may be an embodiment of the system 1000) that includes the computingsystem 1100 as well as a camera 1200A and a camera 1200B, both of which may be an embodiment of the camera 1200. In this example, the camera 1200A may be a 2D camera that is configured to generate 2D image information which includes or forms a 2D image that describes a visual appearance of the environment in the camera’s field of view. The camera 1200B may be a 3D camera (also referred to as a spatial structure sensing camera or spatial structure sensing device) that is configured to generate 3D image information which includes or forms spatial structure information regarding an environment in the camera’s field of view. The spatial structure information may include depth information (e.g., a depth map) which describes respective depth values of various locations relative to the camera 1200B, such as locations on surfaces of various objects in the camera 1200B’s field of view. These locations in the camera’s field of view or on an object’s surface may also be referred to as physical locations. The depth information in this example may be used to estimate how the objects are spatially arranged in three-dimensional (3D) space. In some instances, the spatial structure information may include or may be used to generate a point cloud that describes locations on one or more surfaces of an object in the camera 1200B’s field of view. More specifically, the spatial structure information may describe various locations on a structure of the object (also referred to as an object structure).

[0051] In an embodiment, the system 1000 may be a robot operation system for facilitating robot interaction between a robot and various objects in the environment of tire camera 1200. For example, FIG. 1C illustrates a robot operation system 1500B, which may be an embodiment of the system 1000 / 1500A of FIGS. 1A and IB. The robot operation system 1500B may include the computing system 1100, the camera 1200, and a robot 1300. As stated above, the robot 1300 may be used to interact with one or more objects in the environment of the camera 1200, such as with boxes, crates, bins, pallets, or other containers. For example, the robot 1300 may be configured to pick up the containers from one location and move them to another location. In some cases, the robot 1300 may be used to perform a de-palletization operation in which a group of containers or other objects are unloaded and moved to, e.g., a conveyor belt. In some implementations, the camera 1200 may be attached to the robot 1300 or the robot 3300, discussed below. This is also known as a camera in -hand or a camera on- hand solution., The camera 1200 may be attached to a robot arm 3320 of the robot 1300. The robot arm 3320 may then move to various picking regions to generate image information regarding those regions. In some implementations, the camera 1200 may be separate from tire robot 1300. For instance, the camera 1200 may be mounted to a ceiling of a warehouse orother structure and may remain stationary relative to the structure. In some implementations, multiple cameras 1200 may be used, including multiple cameras 1200 separate from the robot 1300 and / or cameras 1200 separate from the robot 1300 being used in conjunction with inhand cameras 1200. In some implementations, a camera 1200 or cameras 1200 may be mounted or affixed to a dedicate robotic system separate from the robot 1300 used for object manipulation, such as a robotic arm, gantry, or other automated system configured for camera movement. Throughout the specification, “control” or “controlling” the camera 1200 may be discussed. For camera in-hand solutions, control of the camera 1200 also includes control of the robot 1300 to which the camera 1200 is mounted or attached.

[0052] In an embodiment, the computing system 1100 of FIGS. 1A-1C may form or be integrated into the robot 1300, which may also be referred to as a robot controller. A robot control system may be included in the system 1500B, and is configured to e.g., generate commands for the robot 1300, such as a robot interaction movement command for controlling robot interaction between the robot 1300 and a container or other object. In such an embodiment, the computing system 1100 may be configured to generate such commands based on, e.g., image information generated by the camera 1200. For instance, the computing system 1100 may be configured to determine a motion plan based on the image information, wherein the motion plan may be intended for, e.g., gripping or otherwise picking up an object. The computing system 1100 may generate one or more robot interaction movement commands to execute the motion plan.

[0053] In an embodiment, the computing system 1100 may form or be part of a vision system. The vision system may be a system which generates, e.g., vision information which describes an environment in which the robot 1300 is located, or, alternatively or in addition to, describes an environment in which the camera 1200 is located. The vision information may include the 3D image information and / or the 2D image information discussed above, or some other image information. In some scenarios, if the computing system 1100 forms a vision system, the vision system may be part of the robot control system discussed above or may be separate from the robot control system. If the vision system is separate from the robot control system, the vision system may be configured to output information describing the environment in which the robot 1300 is located. The information may be outputted to the robot control system, which may receive such information from the vision system and performs motion planning and / or generates robot interaction movement commands based on the information. Further information regarding the vision system is detailed below.

[0054] In an embodiment, the computing system 1100 may communicate with the camera 1200 and / or with the robot 1300 via a direct connection, such as a connection provided via a dedicated wired communication interface, such as a RS-232 interface, a universal serial bus (USB) interface, and / or via a local computer bus, such as a peripheral component interconnect (PCI) bus. In an embodiment, the computing system 1100 may communicate with the camera 1200 and / or with the robot 1300 via a network. The network may be any type and / or form of network, such as a personal area network (PAN), a local-area network (LAN), e.g., Intranet, a metropolitan area network (MAN), a wide area network (WAN), or the Internet. The network may utilize different techniques and layers or stacks of protocols, including, e.g., the Ethernet protocol, the internet protocol suite (TCP / IP), the ATM (Asynchronous Transfer Mode) technique, the SONET (Synchronous Optical Networking) protocol, or the SDH (Synchronous Digital Hierarchy) protocol.

[0055] In an embodiment, the computing system 1100 may communicate information directly with the camera 1200 and / or with the robot 1300, or may communicate via an intermediate storage device, or more generally an intermediate non-transitory computer- readable medium. For example, FIG. ID illustrates a system 1500C, which may be an embodiment of the system 1000 / 1500A / 1500B, that includes a non-transitory computer- readable medium 1400, which may be external to the computing system 1100, and may act as an external buffer or repository for storing, e.g., image information generated by the camera 1200. In such an example, the computing system 1100 may retrieve or otherwise receive the image information from the non-transitory computer-readable medium 1400. Examples of the non-transitory computer readable medium 1400 include an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. The non-transitory computer-readable medium may form, e.g., a computer diskette, a hard disk drive (HDD), a solid-state drive (SDD), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), and / or a memory stick.

[0056] As stated above, the camera 1200 may be a 3D camera and / or a 2D camera. The 2D camera may be configured to generate a 2D image, such as a color image or a grayscale image. The 3D camera may be, e.g., a depth-sensing camera, such as a time-of-flight (TOF) camera or a structured light camera, or any other type of 3D camera. In some cases, the 2Dcamera and / or 3D camera may include an image sensor, such as a charge coupled devices (CCDs) sensor and / or complementary metal oxide semiconductors (CMOS) sensor. In an embodiment, the 3D camera may include lasers, a LIDAR device, an infrared device, a light / dark sensor, a motion sensor, a microwave detector, an ultrasonic detector, a RADAR detector, or any other device configured to capture depth information or other spatial structure information.

[0057] As stated above, the image information may be processed by the computing system 1100. In an embodiment, the computing system 1100 may include or be configured as a server (e.g., having one or more server blades, processors, etc.), a personal computer (e.g., a desktop computer, a laptop computer, etc.), a smartphone, a tablet computing device, and / or other any other computing system. In an embodiment, any or all of the functionality of the computing system 1100 may be performed as part of a cloud computing platform. The computing system 1100 may be a single computing device (e.g., a desktop computer), or may include multiple computing devices.

[0058] FIG. 2A provides a block diagram that illustrates an embodiment of the computing system 1100. The computing system 1100 in this embodiment includes at least one processing circuit 1110 and a non-transitory computer-readable medium (or media) 1120. In some instances, the processing circuit 1110 may include processors (e.g., central processing units (CPUs), special-purpose computers, and / or onboard servers) configured to execute instructions (e.g., software instructions) stored on the non-transitory computer-readable medium 1120 (e.g., computer memory). In some embodiments, the processors may be included in a separate / stand-alone controller that is operably coupled to the other electronic / electrical devices. The processors may implement the program instructions to control / interface with other devices, thereby causing the computing system 1100 to execute actions, tasks, and / or operations. In an embodiment, the processing circuit 1110 includes one or more processors, one or more processing cores, a programmable logic controller (“PLC”), an application specific integrated circuit (“ASIC”), a programmable gate array (“PGA”), a field programmable gate array (“FPGA”), any combination thereof, or any other processing circuit.

[0059] In an embodiment, the non-transitory computer-readable medium 1120, which is part of the computing system 1100, may be an alternative or addition to the intermediate non- transitory computer-readable medium 1400 discussed above. The non-transitory computer- readable medium 1120 may be a storage device, such as an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, asemiconductor storage device, or any suitable combination thereof, for example, such as a computer diskette, a hard disk drive (HDD), a solid state drive (SSD), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory7(SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, any combination thereof, or any other storage device. In some instances, the non -transitory computer-readable medium 1120 may include multiple storage devices. In certain implementations, the non-transitory computer-readable medium 1120 is configured to store image information generated by the camera 1200 and received by the computing system 1100. In some instances, the non- transitory computer-readable medium 1120 may store one or more object recognition template used for performing methods and operations discussed herein. The non-transitory computer- readable medium 1120 may alternatively or additionally store computer readable program instructions that, when executed by the processing circuit 1110, causes the processing circuit 1110 to perform one or more methodologies described here.

[0060] FIG. 2B depicts a computing system 1100A that is an embodiment of the computing system 1100 and includes a communication interface 1131. The communication interface 1131 may be configured to, e.g., receive image information generated by the camera 1200 of FIGS. 1A-1D. The image information may be received via the intermediate non- transitory computer-readable medium 1400 or the network discussed above, or via a more direct connection between the camera 1200 and the computing system 1100 / 1100A. In an embodiment, the communication interface 1131 may be configured to communicate with the robot 1300 of FIG. 1C. If the computing system 1100 is external to a robot control system, the communication interface 1131 of the computing system 1100 may be configured to communicate with the robot control system. The communication interface 1131 may also be referred to as a communication component or communication circuit, and may include, e.g., a communication circuit configured to perform communication over a wired or wireless protocol. As an example, the communication circuit may include a RS-232 port controller, a USB controller, an Ethernet controller, a Bluetooth® controller, a PCI bus controller, any other communication circuit, or a combination thereof.

[0061] In an embodiment, as depicted in FIG. 2C, the non-transitory computer-readable medium 1120 may include a storage space 1125 configured to store one or more data objects discussed herein. For example, the storage space may store object recognition templates, detection hypotheses, image information, object image information, robotic arm movecommands, and any additional data objects the computing systems discussed herein may require access to.

[0062] In an embodiment, the processing circuit 1110 may be programmed by one or more computer-readable program instructions stored on the non-transitory computer-readable medium 1120. For example, FIG. 2D illustrates a computing system 1100C, which is an embodiment of the computing system 1100 / 1100A / 1100B, in which the processing circuit 1110 is programmed by one or more modules, including an object recognition module 1121, a motion planning module 1129, an image preprocessing module 1126, a hypothesis generation module 1128, and a hypothesis validation module 1138. Each of the above modules may represent computer-readable program instructions configured to cany7out certain tasks when instantiated on one or more of the processors, processing circuits, computing systems, etc., described herein. Each of the above module may operate in concert with one another to achieve the functionality described herein. Various aspects of the functionality described herein may be carried out by one or more of the software modules described above and the software modules and their descriptions are not to be understood as limiting the computational structure of systems disclosed herein. For example, although a specific task or functionality may be described with respect to a specific module, that task or functionality may also be performed by a different module as required. Further, the system functionality described herein may be performed by a different set of software modules configured with a different breakdown or allotment of functionality.

[0063] In an embodiment, the object recognition module 1121 may be configured to obtain and analyze image information as discussed throughout the disclosure. Methods, systems, and techniques discussed herein with respect to image information may use the object recognition module 1121. In embodiments, these may include image capture. The object recognition module may further be configured for object recognition tasks related to object and edge identification, as discussed herein. In embodiments, as discussed below, the object recognition module 1121 may employ machine learning techniques. In embodiments, the object recognition module 1121 may be configured for identifying depth values in the image information.

[0064] The motion planning module 1129 may be configured plan and execute the movement of a robot. For example, the motion planning module 1129 may interact with other modules described herein to plan motion of a robot 3300 for object retrieval operations and for camera placement operations. Methods, systems, and techniques discussed herein with respectto robotic arm movements and trajectories may be performed by the motion planning module 1129. The motion planning module 1129 may be configured to receive occlusion determination information, as discussed herein, and incorporate such into motion planning operations.

[0065] The image preprocessing module 1126 may be configured to perform any tasks related to preprocessing of image information, as may be required by the various methods and techniques described herein.

[0066] The hypothesis generation module 1128 may be configured to generate a detection hypothesis. Detection hypotheses may be understood as identification of a pickable region in an object scene for which robotic retrieval may be planned. The hypothesis generation module 1128 may be configured to interact or communicate with any other necessary module.

[0067] The hypothesis validation module 1138 may be configured to complete hypothesis validation tasks as discussed herein The hypothesis validation module 1138 may be configured to interact with the object recognition module 1121, the image preprocessing module 1126, the hypothesis generation module 1128, and any other necessary modules. In particular, the hypothesis validation module may operate to perform occlusion determination methods. Occlusion determination may, in embodiments, be considered a form of hypothesis validation in which a detection hypothesis is validated by determining the occlusion status of an object associated with a pickable region of the detection hypothesis.

[0068] With reference to FIGS. 2E, 2F, 3A, and 3B, methods related to the object recognition module 1121 that may be performed for image analysis are explained. FIGS. 2E and 2F illustrate example image information associated with image analysis methods while FIGS. 3A and 3B illustrate example robotic environments associated with image analysis methods. References herein related to image analysis by a computing system may be performed according to or using spatial structure information that may include depth information which describes respective depth value of various locations relative a chosen point. The depth information may be used to identify objects or estimate how objects are spatially arranged. In some instances, the spatial structure information may include or may be used to generate a point cloud that describes locations of one or more surfaces of an object. Spatial structure information is merely one form of possible image analysis and other forms known by one skilled in the art may be used in accordance with the methods described herein.

[0069] In embodiments, the computing system 1 100 may obtain image information representing an object in a camera field of view (e.g., 3200) of a camera 1200. The steps and techniques described below for obtaining image information may be referred to below as an image information capture operation 3001. In some instances, the object may one object 5012 from a plurality of objects 5012 in a scene 5013 in the field of view 3200 of a camera 1200. The image information 2600, 2700 may be generated by the camera (e.g., 1200) when the objects 5012 are (or have been) in the camera field of view 3200 and may describe one or more of the individual objects 5012 or the scene 5013. The object appearance describes the appearance of an object 5012 from the viewpoint of the camera 1200. If there are multiple objects 5012 in the camera field of view, the camera may generate image information that represents the multiple objects or a single object (such image information related to a single object may be referred to as object image information), as necessary. The image information may be generated by the camera (e.g., 1200) when the group of objects is (or has been) in the camera field of view, and may include, e.g., 2D image information and / or 3D image information.

[0070] As an example, FIG. 2E depicts a first set of image information, or more specifically, 2D image information 2600, which, as stated above, is generated by the camera 1200 and represents the objects 5012 of FIG. 3 A. The 2D image information is depicted here in simplified form showing four square objectes. More specifically, the 2D image information 2600 may be a grayscale or color image and may describe an appearance of the objects 5012 from a viewpoint of the camera 1200. In an embodiment, the 2D image information 2600 may correspond to a single-color channel (e.g., red, green, or blue color channel) of a color image. If the camera 1200 is disposed above the objects 5012, then the 2D image information 2600 may represent an appearance of respective top surfaces of the objects 5012. In the example of FIG. 2E, the 2D image information 2600 may include respective portions 2000A / 2000B / 2000C / 2000D / 2550, also referred to as image portions or object image information, that represent respective surfaces of the objects 5012. In FIG. 2E, each image portion 2000A / 2000B / 2000C / 2000D / 2550 of the 2D image information 2600 may be an image region, or more specifically a pixel region (if the image is formed by pixels). Each pixel in the pixel region of the 2D image information 2600 may be characterized as having a position that is described by a set of coordinates [U, V] and may have values that are relative to a camera coordinate system, or some other coordinate system, as shown in FIGS. 2E and 2F. Each of the pixels may also have an intensity value, such as a value between 0 and 255 or 0 and 1023.In further embodiments, each of the pixels may include any additional information associated with pixels in various formats (e.g., hue, saturation, intensity, CMYK, RGB, etc.)

[0071] As stated above, the image information may in some embodiments be all or a portion of an image, such as the 2D image information 2600. In examples, the computing system 1100 may be configured to extract an image portion 2000A from the 2D image information 2600 to obtain only the image information associated with a corresponding object 3410A. Where an image portion (such as image portion 2000A) is directed towards a single object it may be referred to as object image information. Object image information is not required to contain information only about an object to which it is directed. For example, the object to which it is directed may be close to, under, over, or otherwise situated in the vicinity of one or more other objects. In such cases, the object image information may include information about the object to which it is directed as well as to one or more neighboring objects. The computing system 1100 may extract the image portion 2000A by performing an image segmentation or other analysis or processing operation based on the 2D image information 2600 and / or 3D image information 2700 illustrated in FIG. 2F. In some implementations, an image segmentation or other processing operation may include detecting image locations at which physical edges of objects appear (e.g., edges of the object) in the 2D image information 2600 and using such image locations to identify object image information that is limited to representing an individual object in a camera field of view (e.g., 3200) and substantially excluding other objects. By “substantially excluding,” it is meant that the image segmentation or other processing techniques are designed and configured to exclude non-target objects from the object image information but that it is understood that errors may be made, noise may be present, and various other factors may result in the inclusion of portions of other objects.

[0072] FIG. 2F depicts an example in which the image information is 3D image information 2700. More particularly, the 3D image information 2700 may include, e.g., a depth map or a point cloud that indicates respective depth values of various locations on one or more surfaces (e.g. , top surface or other outer surface) of the objects 5012. In some implementations, an image segmentation operation for extracting image information may involve detecting image locations at which physical edges of objects appear (e.g., edges of a box or envelope) in the 3D image information 2700 and using such image locations to identify an image portion (e.g., 2730) that is limited to representing an individual object in a camera field of view.

[0073] The respective depth values may be relative to the camera 1200 which generates the 3D image information 2700 or may be relative to some other reference point. In some embodiments, the 3D image information 2700 may include a point cloud which includes respective coordinates for various locations on structures of objects in the camera field of view (e.g., 3200). In the example of FIG. 2F, the point cloud may include respective sets of coordinates that describe the location of the respective surfaces of the objects 5012. The coordinates may be 3D coordinates, such as [X Y Z] coordinates, and may have values that are relative to a camera coordinate system, or some other coordinate system. For instance, the 3D image information 2700 may include a first image portion 2710, also referred to as an image portion, that indicates respective depth values for a set of locations 27101 -2710n, which are also referred to as physical locations on a surface of the object 3410D. Further, the 3D image information 2700 may further include a second, a third, a fourth, and a fifth portion 2720, 2730, 2740, and 2750. These portions may then further indicate respective depth values for a set of locations, which may be represented by 2720i-2720n, 2730i-2730n, 2740i-2740n, and 2750]- 2750nrespectively. These figures are merely examples, and any number of objects with corresponding image portions may be used. Similarly to as stated above, the 3D image information 2700 obtained may in some instances be a portion of a first set of 3D image information 2700 generated by the camera. In the example of FIG. 2E, if the 3D image information 2700 obtained represents an object 5012 of FIG. 3A, then the 3D image information 2700 may be narrowed as to refer to only the image portion 2710. Similar to the discussion of 2D image information 2600, an identified image portion 2710 may pertain to an individual object and may be referred to as object image information. Thus, object image information, as used herein, may include 2D and / or 3D image information.

[0074] In an embodiment, an image normalization operation may be performed by the computing system 1100 as part of obtaining the image information. The image normalization operation may involve transforming an image or an image portion generated by the camera 1200, so as to generate a transformed image or transformed image portion. For example, if the image information, which may include the 2D image information 2600, the 3D image information 2700, or a combination of the two, obtained may undergo an image normalization operation to attempt to cause the image information to be altered in viewpoint, object pose, lighting condition associated with the visual description information. Such normalizations may be performed to facilitate a more accurate comparison between the image information and model (e.g., template) information. The viewpoint may refer to a pose of an object relative tothe camera 1200, and / or an angle at which the camera 1200 is viewing the object when the camera 1200 generates an image representing the object.

[0075] For example, the image information may be generated during an object recognition operation in which a target object is in the camera field of view 3200. The camera 1200 may generate image information that represents the target object when the target object has a specific pose relative to the camera. For instance, the target object may have a pose which causes its top surface to be perpendicular to an optical axis of the camera 1200. In such an example, the image information generated by the camera 1200 may represent a specific viewpoint, such as a top view of the target object. In some instances, when the camera 1200 is generating the image information during the object recognition operation, the image information may be generated with a particular lighting condition, such as a lighting intensity. In such instances, the image information may represent a particular lighting intensity, lighting color, or other lighting condition.

[0076] In an embodiment, the image normalization operation may involve adjusting an image or an image portion of a scene generated by the camera, so as to cause the image or image portion to better match a viewpoint and / or lighting condition associated with information of an object recognition template. The adjustment may involve transforming the image or image portion to generate a transformed image which matches at least one of an object pose or a lighting condition associated with the visual description information of the object recognition template.

[0077] The viewpoint adjustment may involve processing, warping, and / or shifting of the image of the scene so that the image represents the same viewpoint as visual description information that may be included within an object recognition template. Processing, for example, may include altering the color, contrast, or lighting of the image, warping of the scene may include changing the size, dimensions, or proportions of the image, and shifting of the image may include changing the position, orientation, or rotation of the image. In an example embodiment, processing, warping, and or / shifting may be used to alter an object in the image of the scene to have an orientation and / or a size which matches or better corresponds to the visual description information of the object recognition template. If the object recognition template describes a head-on view (e.g., top view) of some object, the image of the scene may be warped so as to also represent a head-on view of an object in the scene.

[0078] Further aspects of the object recognition methods performed herein are described in greater detail in U.S. Application No. 16 / 991,510, filed August 12, 2020, and U.S. Application No. 16 / 991,466, filed August 12, 2020, each of which is incorporated herein by reference.

[0079] In various embodiments, the terms “computer-readable instructions” and “computer-readable program instructions” are used to describe software instructions or computer code configured to carry out various tasks and operations. In various embodiments, the term “module” refers broadly to a collection of software instructions or code configured to cause the processing circuit 1110 to perform one or more functional tasks. The modules and computer-readable instructions may be described as performing various operations or tasks when a processing circuit or other hardware component is executing the modules or computer- readable instructions.

[0080] FIGS. 3A-3B illustrate exemplary environments in which the computer-readable program instructions stored on the non-transitory computer-readable medium 1120 are utilized via the computing system 1100 to increase efficiency of object identification, detection, and retrieval operations and methods. The image information obtained by the computing system 1100 and exemplified in FIG. 3A influences the system’s decision-making procedures and command outputs to a robot 3300 present within an object environment.

[0081] FIGS. 3A - 3B illustrate an example environment in which the process and methods described herein may be performed. FIG. 3A depicts an environment having a system 3000 (which may be an embodiment of the system 1000 / 1500A / 1500B / 1500C of FIGS. 1A- 1D) that includes at least the computing system 1100, a robot 3300, and a camera 1200. The camera 1200 may be an embodiment of the camera 1200 and may be configured to generate image information which represents a scene 5013 in a camera field of view 3200 of the camera 1200, or more specifically represents objects in the camera field of view 3200, such as objects 5012. In an example, the object 5012 may include envelopes and packages or other objects having a generally flat structure. The invention, however, is not limited to such object shapes. FIG. 3A illustrates an embodiment including a conveyor belt as a retrieval location for objects 5012 while FIG. 3B illustrates an embodiment including a container as a retrieval location of objects 5012.

[0082] In an embodiment, the system 3000 of FIG. 3 A may include one or more light sources. The light source may be, e.g., a light emitting diode (LED), a halogen lamp, or anyother light source, and may be configured to emit visible light, infrared radiation, or any other form of light toward surfaces of the objects 5012. In some implementations, the computing system 1100 may be configured to communicate with the light source to control when the light source is activated. In other implementations, the light source may operate independently of the computing system 1100.

[0083] In an embodiment, the system 3000 may include a camera 1200 or multiple cameras 1200, including a 2D camera that is configured to generate 2D image information 2600 and a 3D camera that is configured to generate 3D image information 2700. The camera 1200 or cameras 1200 may be mounted or affixed to the robot 3300, may be stationary within the environment, and / or may be affixed to a dedicated robotic system separate from the robot 3300 used for object manipulation, such as a robotic arm, gantry, or other automated system configured for camera movement. FIG. 3 A shows an example having a stationary camera 1200 and an on-hand camera 1200, while FIG. 3B shows an example having only a stationary camera 1200. The 2D image information 2600 (e.g., a color image or a grayscale image) may describe an appearance of one or more objects, such as the objects 5012 or the object 5012 in the camera field of view 3200. For instance, the 2D image information 2600 may capture or otherwise represent visual detail disposed on respective outer surfaces (e.g., top surfaces) of the objects 5012, and / or contours of those outer surfaces. In an embodiment, the 3D image information 2700 may describe a structure of one or more of the objects 5012, wherein the structure for an object may also be referred to as an object structure or physical structure for the object. For example, the 3D image information 2700 may include a depth map, or more generally include depth information, which may describe respective depth values of various locations in the camera field of view 3200 relative to the camera 1200 or relative to some other reference point. The locations corresponding to the respective depth values may be locations (also referred to as physical locations) on various surfaces in the camera field of view 3200, such as locations on respective top surfaces of the objects 5012. In some instances, the 3D image information 2700 may include a point cloud, which may include a plurality of 3D coordinates that describe various locations on one or more outer surfaces of the objects 5012, or of some other objects in the camera field of view 3200. The point cloud is shown in FIG. 2F

[0084] In the example of FIGS. 3A and 3B, the robot 3300 (which may be an embodiment of the robot 1300) may include a robot arm 3320 having one end attached to a robot base 3310 and having another end that is attached to or is formed by an end effector apparatus 3330, such as a robot gripper. The robot base 3310 may be used for mounting therobot arm 3320, while the robot arm 3320, or more specifically the end effector apparatus 3330, may be used to interact with one or more objects in an environment of the robot 3300. The interaction (also referred to as robot interaction) may include, e.g., gripping or otherwise picking up at least one of the objects 5012. For example, the robot interaction may be part of an object picking operation to identify, detect, and retrieve the objects 5012 from containers. The end effector apparatus 3330 may have suction cups or other components for grasping or grabbing the object 5012. The end effector apparatus 3330 may be configured, using a suction cup or other grasping component, to grasp or grab an object through contact with a single face or surface of the object, for example, via a top face.

[0085] The robot 3300 may further include additional sensors configured to obtain information used to implement the tasks, such as for manipulating the structural members and / or for transporting the robotic units. The sensors can include devices configured to detect or measure one or more physical properties of the robot 3300 (e.g., a state, a condition, and / or a location of one or more structural members / joints thereof) and / or of a surrounding environment. Some examples of the sensors can include accelerometers, gyroscopes, force sensors, strain gauges, tactile sensors, torque sensors, position encoders, etc.

[0086] FIG. 4 provides a flow diagram illustrating an overall flow of methods and operations for the detection, identification, and retrieval of objects, according to embodiments hereof. The object detection, identification, and retrieval method 4000 may include any combination of features of the sub-methods and operations described herein. The method 4000 may include any or all of an image information capture operation 3001, an image preprocessing operation 4001, a feature extraction operation 5000, a hypothesis generation operation 10000, a hypothesis validation operation 13000, and a robotic control operation 15000, motion planning, and motion execution. In embodiments, the hypothesis validation method 13000 may further include occlusion determination techniques, as discussed with respect to FIG. 5. These operations and methods may be performed to facilitate action by a robot. The image information capture operation 3001, the hypothesis generation operation 10000, the hypothesis validation method 13000, and the robotic control operation 15000 may each be performed in the context of robotic operation for detecting, identifying, and retrieving objects from a retrieval location.

[0087] A series of methods, processes, sub-processes, and operations are now described for occlusion determination as part of the robotic image processing pipeline. FIG. 5 illustrates a method of occlusion determination consistent with embodiments hereof. The method of FIG.5 represents the overall structure of the occlusion determination method, while the following figures are used to illustrate various operations and subprocesses of the method 500. Although discussed as part of a hypothesis validation method, various aspects of the method 500 may occur at different parts of the robotic image processing pipeline. For example, the method 500 may use image information captured during a different portion of the pipeline. In another example, the method 500 may operate at least partially in parallel with other aspects of the pipeline. For example, the method 500 may operate to identify occlusions while the system is generating hypotheses. In further examples, the method 500 may be employed in the context of other robotic operations outside of the specific robotic image processing pipeline discussed herein.

[0088] In an embodiment, the occlusion determination method 500 (and all operations and subprocesses thereof) may be performed by, e.g., the computing system 1100 (or 1 100A / 1100B / 1 100C) of FIGS. 2A-2D or the computing system 1100 of FIGS. 3A-3B, or more specifically by the at least one processing circuit 1110 of the computing system 1100. In some scenarios, the computing system 1100 may perform the occlusion determination method 500 by executing instructions stored on a non-transitory computer-readable medium (e.g., 1120). For instance, the instructions may cause the computing system 1100 to execute one or more of the modules illustrated in FIG. 2D, which may perform occlusion determination method 500 (and all operations and subprocesses thereof). For example, in embodiments, steps of the occlusion determination method 500 (and all operations and subprocesses thereof) may be performed by the image preprocessing module 1126, the object recognition module 1121, and the hypothesis validation module 1138 operating in concert to perform an occlusion check on a pickable region identified according to a detection hypothesis. The steps of the occlusion determination method 500 may be employed to identify whether an identified object occludes another object or is occluded by another object. This object occlusion determination may be understood as hypothesis validation to be used for motion planning.

[0089] The at least one processing circuit 1110 may perform specific steps of the occlusion determination method 500 (and all operations and subprocesses thereof) for generating an object occlusion determination. Although the operations and subprocesses are detailed in a specific order, the method 500 is not limited to such. The described steps may be performed in a different order and / or with some steps excluded or repeated as may be appropriate. The description below provides one example of carry ing out steps of the method 500.

[0090] The occlusion determination method 500 may include a plurality of subprocesses, including at least a subprocess 502 for removing an image information background , a subprocess 504 for generating a cost map, a subprocess 506 for segmenting the cost map, and a subprocess 508 for determining object occlusions.

[0091] The occlusion determination method 500 may begin with or otherwise include a subprocess 502, which may be a background removal subprocess, as illustrated by FIGS. 6-7D. FIG. 6 illustrates the subprocess 502 for removing a background from a captured image. FIG. 7A-7C are provided to better illustrate the operations of subprocess 502. The background removal subprocess 502 is configured to identify, isolate, and remove background information from a captured imaged (either 2D, 3D, or both). The background removal subprocess 502 removes the background and retains the objects in the image information.

[0092] In an operation 602, 2D image information, for example, 2D image information 2600 may be obtained. The 2D image information 2600 may be obtained from any suitable source, e.g., from any of the cameras or optical sensing devices discussed herein and / or from an interim location such as a memory device. The 2D image information may include an image or image information about one or more objects in a scene. In the example, provided in FIG. 7A, the 2D image information 2600 is a 2D color image 701 captured by a camera positioned above a conveyor belt with objects 5012, here shown as envelopes, stacked thereon. In the 2D color image 701, the background 5035 represents the conveyor belt.

[0093] In an operation 604, a candidate background color may be identified. Identifying a candidate background color may include identifying color characteristics of the pixels (or other image units) in the 2D image information 2600. Color characteristics of the pixels may include HSV (hue, saturation, value) model values. Other suitable color characteristics may be based on other models, such as RGB, CMYK, Lab, XYZ, YCbCr, HSL, grayscale, or any other suitable model for describing pixel color characteristics.

[0094] From the determined color characteristics, the system 1100 may estimate the color of the background 5035. In this example the color of the background 5035 is represented by the color of the conveyor belt. FIG. 7B illustrates a probability distribution 702 of the estimated color of the background as determined by the system 1100, shown by a black circle. In FIG. 7B, the x-axis represents hue and the y-axis represents coloredness (e.g., a combination of saturation and value). In other color models, the x and y axes may be adjusted accordingly. The magnitude at each point in tire probability distribution 702 is represented in grayscale,where white equals a low probability and black equals a high probability. In some embodiments, these colors may be reversed. The center of the black circle represents a known or expected background color, e.g., a known color of a container, conveyor belt, etc. Pixels having color characteristics that exactly match the expected background color (e.g., in hue and coloredness) will fall in the center of the high probability portion 5037. As the hue and / or coloredness varies away from an exact match, the background color likeliness falls. Pixels away from the high probability portion 5037 are unlikely to be a background color at all.

[0095] The system 1100 may select one or more points within this distribution as candidate background pixels according to their likelihood of matching the background color (e.g., whether the pixels fall near the high probability portion 5037). The system 1100 may, in embodiments, be preconfigured to reject certain colors as candidate background portions. The preconfigured rejected values may be based on the type of object that is likely to enter the field of view of the imaging device. In embodiments, candidate background colors may be rejected if they match the expected color of objects 5012. In embodiments, candidate background colors may be rejected if they fail to match expected color of a background 5035.

[0096] In an operation 606, 3D image information, for example, 3D image information 2700 is obtained or accessed, by methods described herein. The 3D image information 2700 may be obtained from a suitable imaging device as described herein, including, for example, the same imaging device that obtained the 2D imaging information 2600. The precision of the imaging device may depend on the type of object that is expected to enter the field of view.

[0097] In an operation 608, candidate background image data portions may be identified from the 3D image information 2700. The system 1100 may determine color characteristics of pixels and / or pixel regions and then determine how closely the color characteristics (e.g., HSV values) of the 3D image information 2700 correspond to the predetermined background color as determined in operation 604. The system 1100 may generate a probabilistic map to represent portions having higher probability other portions having a lower probability of being the background based on how closely the color characteristics match those of the predetermined background color.

[0098] In embodiments, the system 1100 may further employ chromatic weighting to determine how much trust to place in whether an identified hue within the image corresponds to the actual hue. The chromatic weight may refer to a likelihood that an identified attribute of tire color characteristics corresponds to an actual attribute of the physical obj ect represented bythe color characteristics. In the HSV model, the chromatic weight may be based on both the saturation and the value of the pixels and / or pixel regions as identified from the HSV determination. If the value in HSV is very low (e.g. below a hue threshold), all hues may essentially look the same (black). If the saturation is too low (e.g. below a saturation threshold), the hue may get washed out. To have sufficient chromatic weight, a function of the saturation and the value must be satisfied. In embodiments, the desired or sufficient chromatic weight (e.g., chromatic strictness) may be described as an exponential line on a plane with value along the y-axis and saturation along the x-axis, as illustrated in FIG. 7C. The chromatic strictness can correspond to the percentage of this plane above (greater saturation and value) the exponential line 5036. Saturation / Value points with sufficient chromatic weight are in the upper right of the graph of FIG. 7C, satisfying the chromatic strictness threshold represented by the exponential line 5036. Saturation / V alue points in the lower left, that do not satisfy the chromatic strictness threshold may be more difficult to confirm that the identified hue is accurate.

[0099] The system 1100 may use the probabilistic map and the chromatic weighting as determined from the 3D image information 2700 to further refine the probability that the regions of the 2D image information 2600 corresponding to the candidate background color are actually background. For instance, the system 1100 may use a function of the probabilistic determination and the chromatic weight to obtain a chromatic weighted probability that any given pixel or pixel region matches the candidate background color. The function to obtain a chromatic weighted probability may include multiplying the probabilistic determination and tire chromatic weight or multiplying weighted or scaled versions of these.

[0100] The system 1100 may use a predetermined area mask to identify portions of the 2D image information 2600 that will have portions that correspond to the background or may indicate a region of interest. FIG. 7D illustrates this within the dotted line. The predetermined area mask may be an area mask that is predetermined according to known characteristics of the observing cameras and the observational space. For example, when the objects 5012 are in a container, the predetermined area mask may be used to mask any area outside of the container. When the objects 5012 are on a conveyor belt, the predetermined area mask be used to mask any area off of the conveyor belt.

[0101] In further embodiments, the system 1100 may use depth data of the 3D image information 2700 to facilitate the background determination. For example, within the predetermined portion of the image (e.g., based on the predetermined area mask), tire system1 100 may determine the depth of various portions of the scene. The system 1100 may generally determine that information representing a distance further from the imaging device corresponds to the background (e.g., a conveyor belt or container bottom). The system 1100 may also take into consideration the relative number and / or consistency of these additional measurements. Accordingly, the system 1100 may determine a depth probability for individual pixels and / or pixel regions representing a likelihood, based on depth, that the pixel or pixels represent an object scene background. In embodiments, depth data outside the region of interest as determined by the predetermined area mask may be ignored or not computed.

[0102] The system 1100 may combine the information from the chromatic weighted probability that the hue corresponds to the background and the probability that the depth data corresponds to the background. This combination may be used to estimate which portions of the image data (e.g., 2D image information 2600 and / or 3D image information 2700) correspond to the background. Accordingly, the system 1100 may identify tire probability of regions representing background regions of the image data (e.g., 2D image information 2600 and / or 3D image information 2700) according to a combination of chromatic weighted probability and depth probability. Where the probability of a region representing the background is sufficiently high, e.g., surpassing a predetermined threshold, these regions may be determined or identified as background regions.

[0103] In an operation 610, identified background image data portions may be removed, e.g., masked off, from the 2D image information 2600 and / or the 3D image information 2700. As discussed above, tire system 1100 may use any combination of the determined probability of the 2D image information 2600 and / or 3D image information corresponding to the background. For instance, the system 1100 may use a combination of the probabilistic determinations (depth and chromatic) to determine whether there is a sufficiently high, for example over a certain threshold, probability' that the portion of the image data corresponds to the background. This may be used to logically remove (e.g., by applying a mask) the portions of the image information determined to correspond to the background. The background masked image 7040 illustrated in FIG. 7D illustrates an object holding container with a masked background shown in cross-hatching.

[0104] Returning now to FIG. 5, a cost map generation subprocess 504 may be performed. FIG. 8 illustrates the subprocess 504 for generating a cost map. FIGS. 9A-9B are provided to better illustrate the operations of subprocess 504. In the cost map generationsubprocess 504, a cost map indicative of surface variation among objects in a scene may be generated based on the obtained image information (e.g., image information 2600 / 2700).

[0105] A cost map may be used to represent information of an area. The cost map may include cost values associated with different x-y locations of the area and / or with different segments or portions of the area. In embodiments, the cost map may describe the degree of similarity or differences between adjacent image data units in the image data. Image data unitsmay include discretized units of the image data such as an individual pixel or groupings of multiple adjacent pixels (e.g., 2x2, 3x3, etc. groupings), 3D points in a 3D point cloud, or a combination thereof. Each of the image data units may have a cost value describing a degree of difference or similarity of an image characteristic with neighboring image data units. In embodiments, the cost value can be based on one or more image characteristics. Each image characteristic on which the cost value is based may be considered a layer of the cost map. Each layer of the cost map may include partial cost values based on a specific image characteristic associated with the cost map layer. The total cost value at any given image data point may be represented by a function of the partial cost values of the cost map layers.

[0106] In embodiments, cost map layers may include a plurality of layers based on, for example, differences in image characteristics, such as depth value between adjacent image points, differences between surface normal angle magnitudes of adjacent image points, differences between surface normal angle directions of adjacent image points, and / or whether the image point has been identified as part of an object edge. In embodiments, the cost value for any one image data unit in a cost map layer may be calculated as an average of the difference in an image characteristic between the image data unit and the 8 neighboring image data units (in a grid system). In embodiments, the cost map can include one, two, three, or four layers, including any combination of the different layers described herein. Embodiments described below discuss the generation of each individual layer. It will be understood that, in embodiments including fewer than four layers, steps associated with generation of unincluded layers may not be performed. Although layers may be described herein as “first,” “second,” etc., this description is purely for example purposes and the numbering of layers may differ depending on the combination of layers selected. In embodiments, the cost map can be used by the system 1100 to determine whether image data units are part of the same surface of an object or part of surfaces of different objects. These embodiments will be described in detail below.

[0107] In embodiments, the cost map values may be determined based on identified surface variations among object in the scene, as discussed below. In embodiments, the cost map may include a plurality of layers, as discussed further below. Each of the plurality of layers may be computed or determined according to different factors. The layers may be combined according to weighting factors to provide a complete cost map. In embodiments, as discussed below, partial cost values for a layer may be between 0 and 1 (although any range of values may be used as well), with 0 representing a low cost indicating that an image data point has no differences in image characteristics with adjacent image data units and 1 representing a high cost indicating that an image data point has maximum differences in image characteristics with adjacent image data units. The partial cost values for each layer may be combined, as discussed below, to generate cost values for the overall cost map.

[0108] In embodiments consistent with the present disclosure, a cost map may include cost values that generally correspond to the amount of effort (or difficulty) to enter into tire associated portion of the area. For instance, the cost value may represent the amount of effort for a robotic system to enter a particular portion of the area. “Effort” does not necessarily correspond to physical exertion, but may alternatively or additionally include other factors such as undesirable areas (such as areas which raise safety concerns or would result in an undesirable outcome), areas that may require significant additional computation to properly plan entry to, etc. In some situations, the system 1100 may attempt to avoid portions of the area that have higher costs, preferring to navigate toward lower cost portions.

[0109] FIG. 9A and FIG. 9B are provided to better illustrate the cost map generation subprocess 504. FIGS. 9A and 9B illustrate a plurality of objects 5012 and 5012’ in a scene. The object edges 5045 and object surface 5046 are shown. FIG. 9A represents a cost map layer associated with gradient angle magnitudes. FIG. 9B represents a cost map layer associated with gradient angles. In embodiments, when the cost map is completed, portions of the image data corresponding to object surfaces 5046, e.g., areas that are surrounded or partially surrounded by object edges 5045, may be grouped into segments according to their respective cost values. This may indicate that a physical object edge 5045, having a higher cost map value, may represent a boundary between objects

[0110] In an operation 802, the system 1100 may identify object edges and surfaces. The object edges may be used to generate a first layer of the cost map, i.e.., an object edge layer. The system 1100 may determine which portions of 2D image information 2600 and / or 3D image information 2700 correspond to objects 5012 in the scene. The system 1100 may furtherbe configured to estimate location and orientation of physical object edges 5045 of the objects 5012 in the scene. Additionally, the system 1100 may be configured to remove or ignore image textures, e.g., logos, labels, hand writing, etc. the system 1100 may estimate object edges.

[0111] In embodiments, this may be performed via computational algorithm and / or by a neural network or other artificial intelligence or machine learning mode. As used herein, the term “artificial intelligence” (Al) is used to refer to a wide variety of methods and systems, including, but not limited to machine learning algorithms and systems. Suitable machine learning systems and algorithms may include, for example, supervised learning models, unsupervised learning models, reinforcement learning models, deep learning neural network models such as feedforw ard neural networks (FNNs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs), transformer-based models including at least large language models (LLMs), generative models, hybrid or specialized models, etc. The following description refers to “neural netw orks” for convenience, but the scope of the present disclosure may further include any suitable artificial intelligence or machine learning model.

[0112] In embodiments, the system 1100 may employ the computational algorithm or neural network to identify the objects, determine portions of image information belonging to objects, remove texturedness, and / or estimate the location and orientation of objects within the scene. The system 1 100 may utilize, for example, a mobile neural network, trained on a public database and / or private database. Identified object edges and surfaces may be assigned an edge parameter, for example, an edge probability value indicating a likelihood that the associated image data unit belongs to an edge. In embodiments, image portions close to object edges may have higher and higher edge parameter values as the edge is approached and lower and lower cost values as the edge is retreated from. The distribution of edge parameter values at an edge may be linear, gaussian, or may take another distribution. Cost map values of the object edge layer may be assigned according to the identified object edges and surfaces. Image portions identified as belonging to object edges may be assigned higher cost values (e.g., 1) and image portions identified as belonging to object surfaces may be assigned lower cost values (e.g., 0). In embodiments, the cost value for any one image data unit in the object edge layer may be calculated as a function of the edge parameter value of itself and the 8 neighboring image data units, for example, as an average of the differences in the edge parameter values between the image data unit and its neighbors.

[0113] In an operation 804, gradient angle magnitudes of the object corresponding portions are determined. The gradient angle magnitudes may be used to determine an anglemagnitude layer (e.g., a second layer) of tire cost map. Gradient angle magnitudes represent an amount of curvature at an image data point and thus may be referred to as a measure of curvature. In embodiments, as discussed below, other measures of curvature may be used without departing from the scope of this disclosure. Although the methods described herein are illustrated and described with respect to gradient angle magnitudes, all operations discussed herein may alternatively be performed with other curvature measures as may be appropriate. Thus, the angle magnitude layer of the cost map may be a curvature measure layer.

[0114] The gradient angle magnitudes may represent angles between changing surface normal vectors of objects in the scene. As an initial step, the system 1100 may identify surface normal vectors of image data units according to the 2D image information 2600 and / or the 3D image information 2700. Normal vectors correspond to unit vectors projecting orthogonally from the surface of the objects 5012 (e.g. envelopes, in the examples shown) at each corresponding image data point. The gradient angle magnitudes may be determined based on the difference between angles of two surface normal vectors. If the two surface normal vectors are provided with the same origin point, the total angle between them is the gradient angle magnitude. The gradient angle magnitudes represent a measure of how great the change in direction of the surface normal vectors is, or a measure of curvature of the objects.

[0115] In further embodiments, other curvature measures may be used, for example, a magnitude of the gradient vector between surface normal vectors. Surface normal vectors having a greater gradient angle magnitude between them will also have a greater magnitude of tire gradient vector between them. Other curvature measures may also be used without deviating from the scope of this disclosure. As described herein, all operations performed based on a gradient angle magnitude may also be performed based on an alternative curvature measure.

[0116] In embodiments, for a given image data point, the gradient angle magnitude may be computed based on the surface normal vectors of surrounding image data units. In embodiments, the surface normal vector of the image data point may be compared to the surface normal vectors of the 8 surrounding image data units in a grid. The gradient angle magnitude assigned to the image data point may be a function of the angle magnitudes between the surface normal vector of the image data point and tire surround image data units. In embodiments, the function may be the Euclidean norm or the root-sum-square.

[0117] In embodiments, for a given image data point, the gradient angle magnitude is determined according to differences between the surface normal vector of an image data point on one side and the surface normal vector of an image data point on the other side. In embodiments, this computation may be performed in the x direction and the y direction. In embodiments, this computation may also be performed in the diagonal directions. In embodiments, the gradient angle magnitude may be determined as a function of the x-direction angle magnitude and the y-direction angle magnitude, or as a function of the four angle magnitudes. In embodiments, the gradient angle magnitude may be computed as a Euclidean norm or the root-sum-square of the (up to four) individual directional angle magnitudes.

[0118] The gradient angle magnitude at a given point within the 3D image information 2700 therefore is representative of how much the neighboring surface normal vectors differ from the surface normal vector at that point. The gradient angle magnitudes may be used to generate the angle magnitude layer of the plurality of layers of the cost map. The system 1100 may assign a higher cost to data points indicating a grater gradient angle magnitude.

[0119] For two very similar vectors (e.g., two vectors next to each other on a largely flat surface 5046), the gradient angle magnitude will be very small and possibly zero. For two very different vectors (e.g., vectors near an edge 5046, which may include a first vector on a surface and a second vector on a neighboring perpendicular surface), the gradient angle magnitude would be 90°. A maximum gradient angle magnitude may be 180° and may appear, for example, when an object is folded in half. In some embodiments, the system 1100 may not be capable of identifying two surface normal vectors having a gradient angle magnitude of 180° . Accordingly, the maximum gradient angle magnitude may be set as the largest detectable gradient angle magnitude, which may be, for example, 175 °, 170 °, 165 °, 160,0etc. The maximum gradient angle magnitude may have a cost value of 1 while the minimum gradient angle magnitude may have a cost value of 0.

[0120] In an operation 806, gradient angles are determined. The gradient angles may be used to determine an angle difference layer (e.g., a third layer) of the cost map. Gradient angles represent a direction of greatest curvature at an image data point and thus may be referred to as indicative of a direction of curvature. In embodiments, as discussed below, other measures of curvature direction may be used without departing from the scope of this disclosure. Although the methods described herein are illustrated and described with respect to gradient angles, all operations discussed herein may alternatively be performed with other curvaturedirection measures as may be appropriate. Accordingly, the angle difference layer may be a curvature direction layer.

[0121] In embodiments, the angle difference layer may be an optional layer. The gradient angles represent changing angles of surface normal vectors of objects in the scene. From the gradients determined in the x and y directions, the system 1100 may also determine the gradient angle. While the gradient angle magnitudes represent the magnitude of the angular difference betw een surface normal vectors, the gradient angles represent the directional angular difference in the x-y plane, which may also be referred to as the azimuthal difference. In an example, the system may determine that the gradient angle is the arctan of the ratio of an x- direction gradient vector and a y-direction gradient vector. The gradient angle may represent the angle, in the X-Y plane, between the two compared surface normal vectors and therefore may be representative of the direction of curvature at the reference point. As shown in FIG. 9B, the objects on the furthest right hand side are not lying completely flat on the conveyor. Some objects may also have a slight concave or convex bulge to them, which would be shown as the normal vectors continuously changing angle. Where the normal vectors are continuously changing angle as in object 5012’, gradient angles may consistently be elevated. The right most envelope in the original image (e.g. object 5012’), for example, appears to have some bulge to it.

[0122] The system 1100 may determine a layer of the cost map from the gradient angle determinations. For instance, the system may determine which parts of the gradient angle data have a change, relative to nearby values, greater than a threshold. That is, the system 1100 may determine where the gradient angle is most rapidly changing. Such rapid changes may help determine a location of an object edge. A high cost may be assigned to the image portions that indicate a rapid change in gradient angle, and a lower cost may be assigned to slow7or no changes in the gradient angle.

[0123] In an operation 808, an optional subprocess for sharpening gradient angle magnitudes may be performed. FIG. 10 illustrates a subprocess method 1000 for sharpening gradient angle magnitudes. The subprocess method 1000 for sharpening gradient angle magnitudes may involve the modification of gradient angle magnitudes according to weights determined from the gradient angles. The subprocess method 1000 for sharpening gradient angle magnitudes may include subprocess 1002 for obtaining image data, subprocess 1004 for computing gradient angle and gradient angle magnitude, subprocess 1006 for weighting gradient angle magnitudes, and subprocess 1008 for sharpening gradients. In the optionalsubprocess for sharpening gradient angle magnitudes as described below, several of the operations may overlap or otherwise share functionality with method 500 and / or subprocess 504. For example, obtaining image data, determining gradient angle magnitudes, determining gradient angles, etc. In embodiments, these overlapping operations may be conducted twice or may be conducted only a single time as may be necessary.

[0124] FIG. I l illustrates the subprocess 1002 for obtaining image data. Subprocess 1002 for obtaining image data may include operation 1102 for generating 3D image information 2700 and operation 1104 for calculating normal vectors of the 3D image information 2700.

[0125] In an operation 1102, 3D image information, for example, 3D image information 2700 may be generated, captured, or otherwise obtained, as described throughout this disclosure. The 3D image information 2700 may be obtained from a suitable imaging device as described herein, including, for example, the same imaging device that obtained the 2D imaging information 2600. The precision of the imaging device may depend on the type of object that is expected to enter the field of view. Operation 1102 may be similar to operation 606, as described above.

[0126] In an operation 1104, normal vectors corresponding to pixels (or other image unit) of the 3D image information may be obtained, calculated, and / or generated. In some embodiments, the normal vectors may be calculated by system 1100. In other embodiments, tire normal vectors may be generated by a camera unit capturing the 3D image information.

[0127] FIG. 12 illustrates the subprocess 1004 for computing gradient angles and magnitudes. The subprocess 1004 for computing gradient angles and magnitudes may include an operation 1202 for accessing image normal vector data, an operation 1204 for calculating gradient angle magnitude, and an operation 1206 for calculating gradient angle.

[0128] In an operation 1202, the system 1 100 may access or obtain the image normal vector data, as obtained, e.g., at operation 1104.

[0129] In an operation 1204, which may correspond to operation 804 for calculating gradient angle magnitude, the system 1100 may calculate gradient angle magnitudes from the image normal vector data. At operation 1204, the system 1100 calculates the gradient angle magnitudes of the surface normal vectors obtained at operation 1 104. The system 1100 may employ the same methods as in operation 804 for determining gradient angle magnitudes.

[0130] In an operation 1206, which may correspond to operation 806 for calculating gradient angles, the system 1100 may calculate gradient angles from the image normal vector data. At operation 1206, the system 1100 calculates the gradient angles of the surface normal vectors obtained at operation 1104. The system 1100 may employ the same methods as in operation 806 for determining gradient angles.

[0131] FIG. 13 illustrates the subprocess 1006 for weighting gradient angle magnitudes. FIGS. 14A-14B are provided to better illustrate the operations of subprocess 1004. The subprocess 1006 for computing gradient angles and magnitudes may include an operation 1302 for selecting a reference angle, an operation 1304 for comparing reference and gradient angles, an operation 1306 for calculating a weighting factor, and an operation 1308 for modifying a gradient angle magnitude. The steps of subprocess 1006 may be repeated several times for each of a plurality of reference angles or may be repeated a single time for one reference angle.

[0132] In an operation 1302, a reference angle is selected from a plurality of reference angles. The plurality of reference angles may include any suitable number of reference angles between 0° and 180°. The reference angles may be evenly spaced (e.g., every 1°, every 5°, every 10 °, every 15 °, every 20 °, every 30 °, etc.) In embodiments, the reference angles may be unevenly spaced within the range. In embodiments, specific angles may be chosen to provide the best magnitude sharpening.

[0133] In an operation 1304, the system 1100 may determine differences between the gradient angle, e.g, as determined in operation 1206, and the selected reference angle. As discussed above, gradient angle is a measure of curvature direction. Small differences between the gradient angle and the selected reference angle indicate that curvature at the corresponding image data point is occurring in the same direction as the reference angle.

[0134] In an operation 1306, the differences determined at operation 1304 are compared to a threshold. The system 1100 determines whether the difference associated with each image unit is small enough, e.g., below a specific threshold. If the difference is below a threshold, this indicates that the direction of change in curvature aligns with the reference angle within the thresholded amount.

[0135] In an operation 1308, weighting factors are applied to the gradient angle magnitude data. The system 1100 may first assign weighting factors to image data units based on whether or not the threshold is satisfied. In embodiments, a single weighting factor may be assigned to all of the image data units having difference values that satisfy the threshold and adifferent weighting factor may be assigned to all of the image data units that do not satisfy tire threshold. In some embodiments, rather than using the threshold, the weighting factors may be assigned according to the difference between the gradient angle and the reference angle. As an example, FIG. 14A illustrates an example image showing image portions that satisfy the threshold for a reference angle of zero. The light color image portions of FIG. 14A satisfy a threshold when compared to a reference angle of 0. As can be seen, the lighter color portions correspond to vertical edges.

[0136] Image data units (e.g., pixels) corresponding to the difference data values that satisfy the threshold, have their gradient angle magnitudes increased by a weighting factor, while pixels that do not satisfy the threshold may have their gradient angle magnitudes decreased by a weighting factor or left alone. This operation functions to increase the gradient angle magnitudes of image data units that have a direction of curvature corresponding to the selected reference angle while decreasing gradient angle magnitudes of image data units that have a direction of curvature that does not correspond to the selected reference angle.

[0137] Compared to the original gradient angle magnitude data this magnitude sharpening step may reduce noise and emphasize the edges of the perceived objects. FIG. 14B illustrates an image having emphasized edges as a result of the gradient angle magnitude weighting described herein.

[0138] The subprocess 1006 for weighting gradient angle magnitudes may be performed for each of the plurality of reference angles, resulting in a corresponding plurality of weighted gradient angle magnitude maps, with each map being weighted based on a corresponding reference angle.

[0139] FIG. 15 illustrates the subprocess 1008 for sharpening gradients. FIGS. 16A-16B are provided to better illustrate the operations of subprocess 1008. The subprocess 1008 may include an operation 1502 for accessing weighted gradient angle magnitude data (e.g., he weighted gradient angle magnitude maps), an operation 1504 for accessing sharpening filter information, an operation 1506 for processing the weighted gradient angle magnitude data, an operation 1508 for repeating operations 1502-1506, and an operation 1508 for combining sharpened data.

[0140] In an operation 1502, the weighted gradient angle magnitude data associated with a selected reference angle, as computed in subprocess 1006, may be accessed for further use in subprocess 1008.

[0141] In an operation 1504, the sharpening filter associated with the selected reference angle is accessed by the system 1100. FIG. 16A illustrates, e.g., the sharpening filter 1800 associated with the reference angle 0°. The filter may be used to further process the weighted gradient angle magnitude data as determined at operation 1308. The sharpening filter 1800 may differ from a sobel operator for edge detection in that it is larger to more strongly emphasize the direction or angle of the edge. In the illustrated sharpening filter 1800, the central, darker values are areas the indicate that gradient angle magnitudes will be increased, the perimeter darker values are areas that indicate that gradient angle magnitudes will be decreased, and the area therebetween represent areas where the gradient angle magnitudes will be maintained.

[0142] In an operation 1506, the sharpening filter is applied to the weighted gradient angle magnitudes previously obtained. The sharpening filter operates to sharpen or emphasize objects edges within the scene that have angular orientations corresponding (e.g., similar to and / or within a range of) to the reference angle associated with the selected sharpening filter. For example, for the sharpening filter associated with the reference angle 0 (filter 1800), object edges that correspond (e.g., are substantially parallel) to 0° (e.g., vertical) are emphasized or sharpened. Substantially parallel edges may include those that are within 20%, 15%, 10%, 5%, and / or 1% of parallel, depending on the strictness of filtering parameters. Thus, as shown, e.g., at FIG. 16B, object edges that are generally vertically oriented may be emphasized. The sharpening filter may operate by increasing weighted gradient angle magnitudes that correspond with the center of the filter and decreasing weighted gradient angle magnitudes that correspond with the edges of the filter.

[0143] When the filter is applied across the corresponding map of weighted gradient angle magnitude values it may operate as follows. Image data units that do not represent edges of any sort have small weighted gradient angle magnitudes and therefore experience little change, positive or negative. Image data units that represent edges not corresponding to the reference angle have already been reduced in the weighting step and will also experience little change or have their values further reduced because their orientation does not correspond to the filter. Image data units that have large gradient angle magnitudes and represent edges corresponding to the reference angle will have their gradient angle magnitudes increase. Image data units that neighbor those representing edges corresponding to the reference angle (e.g., to either side of the represented edge) will have their gradient angle magnitudes decreased. This results in the edge being both emphasized and tightened. 1

[0144] In an operation 1508, the system 1100 is configured repeat operations 1502, 1504, and 1506 for different reference angles within the plurality of reference angles. Accordingly, as the system 1100 performs the above described methods for gradient angle magnitude sharpening across reference angles ranging from 0 to 180, many, most, substantially all, or all of the object edges may be sharpened / emphasized. FIG. 16C illustrates a subset of the plurality of sharpening filters associated with the plurality of reference angles. The illustrated filters are associated with the reference angles 0, 7.5, 15, 22.5, 30, and 37.5 degrees. Each sharpening filter is slightly rotated to match its corresponding reference angle. The sharpening filters may each have a Gaussian distribution of filtering values along their filter axis. In embodiments, sixty evenly spaced sharpening filters may be used.

[0145] In an operation 1510, the sharpened images are combined. As discussed above, the plurality of sharpening filters, each corresponding to a reference angle, are applied to the gradient angle magnitude data to sharpen object edges. Some or all of the resulting plurality of sharpened gradient angle magnitude images may then be combined, e.g., as shown in FIGS. 16D-16F. FIG. 16D illustrates the entire scene with object edges sharpened. FIG. 16E illustrates a zoomed-in portion of the image from prior to image sharpening. FIG. 16F illustrates the same zoomed-in portion from after image sharpening. It can be seen, in the zoomed in portion, that the edges are more solid with less variation in the post-sharpening image.

[0146] Returning now to FIG. 8, in an operation 809, the system 1100 may be configured to generate a depth layer (e.g., a fourth layer) of the cost map according to the image depth of objects in the scene. In the depth layer, the cost values are determined according to depth separation between neighboring image data units. Depth separation represents a computed depth difference between an image data point and its neighboring image data units. The system may determine a maximum separation depth according to the expected types of objects. In embodiments, the maximum separation depth may be equal to or smaller than the maximum thickness of the expected objects. For thin expected objects (e.g., envelopes), the maximum separation depth may be equal to the expected maximum thickness. For thicker objects (e.g., boxes), the maximum separation depth may be a fraction (e.g., ‘A, 2 / 3, %, etc.) of the expected thickness. Image data units may be assigned cost values according to a comparison between their depth separation value and the maximum separation depth. Where the depth separation value equals or excess the maximum separation depth, a maximum cost value (e.g., 1) may be assigned. Where the depth separation value is zero, a minimum cost value (e.g., 0) may beassigned. Depth separation values between zero and the maximum separation depth may be interpolated, for example, in a linear or other suitable fashion. In embodiments, the cost value for any one image data unit in the depth layer may be calculated as a function of the depth separation value of itself and the 8 neighboring image data units, for example, by first averaging the depth separation values between an image data unit and its neighbors and then comparing to the maximum separation depth.

[0147] In an operation 810, the cost map is generated by the system 1100. The system 1100 may generate the multilayer cost map by combining the previously calculated cost values, e.g., from the object edges layer, the angle magnitude layer, the angle differences layer, and a depth layer. As discussed above, the system 1100 may be configured to use any combination of one or more of the described layers (e.g., partial cost maps). In further embodiments, the system 1100 may be configured to combine any combination of one or more of the described partial cost maps with other cost maps obtained through alternate techniques. To combine the layers the costs of each may be added together, or the system 1100 may apply weighting factors to favor certain cost determinations over others. The system 1100 may normalize the combined costs or may use the weighted summed values without normalization. In embodiments, the background may either be ignored or may have a high cost value assigned to it. As discussed above, the cost map represents a map of values where the values correspond to the individual image units of the original image information (e.g., 2D or 3D). The values in the cost map may represent a likelihood that the corresponding image data units are associated with an object boundary or edge. Low cost values indicate a likelihood that the image data units are associated with an object surface while high cost values indicate a likelihood that the image data units are associated with an object edge or boundary.

[0148] Returning now to FIG. 5, the subprocess 506 for cost map segmentation may be executed. The subprocess 506 may be performed to segment the cost map according to a comparison between cost map values and a reference cost threshold. The subprocess 506, as shown in FIG. 17, is an operation 1702 for dividing the cost map, an operation 1704 for identifying representative cost values, an operation 1706 for selecting a reference cost threshold, an operation 1708 for determining candidate segments, an operation 1710 for repeating operations 1706 and 1708, and an operation 1712 for determining polygons. The subprocess 506 is not the sole way in which a cost map may be segmented. Other suitable algorithms that identify object boundaries according to the cost map values may be employed, including, forexample, those described in U.S. Patent Publication 2023 / 0286165, published September 14, 2023, and incorporated herein by reference.

[0149] In an operation 1702, the cost map (e.g., the cost map as constructed at subprocess 504), may be divided. The system may divide the cost map into discrete units or portions. This division may be performed, for example, by the application of grid lines. In other embodiments, each section may be selected according to having image data units with similar cost values. However, in embodiments, the cost map may be divided in non-equivalent ways. For example, certain portions of the image map may have smaller units, such as one side or a center. In embodiments, the size of the segments may be selected based on available computing power, with smaller segments being selected to correspond to greater computing power. In embodiments, the size of the segments may also be selected to be small enough to represent the desired potential detail of the cost map. FIG. 18A illustrates a divided cost map 1825, showing a grid of divided units 1827.

[0150] In an operation 1704, within each unit of the divided cost map, the system 1100 may determine which value to select as representative of other nearby values within the divided unit. For example, the system 1100 may select the cost value of the image units closest (e.g., as determined from the 3D image data) to the imaging device to represent the unit or portion. As another example, the system 1100 may select the cost value of the image units having the lowest combined cost value to represent the unit or portion. Selecting a single representative value for each unit or portion may help to reduce the computing overhead. Representative values 1826 for each divided unit 1827 are illustrated in FIG. 18A

[0151] In an operation 1706, the system 1100 may select a reference cost threshold. For instance, in embodiments, cost values of the cost map may be normalized to range from 0.0 to 1.0. A reference cost threshold between 0 and 1.0 (or between a minimum value and a maximum value of the cost map) may be selected. The system 1100 may be configured to avoid selecting cost thresholds that are less than and / or more than a certain value, e.g., cost values outside of a rang of cost values that appear in the cost map.

[0152] In an operation 1708, system 1100 may determine at least one candidate segment. The candidate segment may be determined based on having on one or more data point (e.g., representative cost value of an individual unit or portion) that satisfies the selected reference cost threshold. The system 1100 may begin at the determined representative value within a divided unit of the cost map and proceed to search for other representative values that are equalto or lower than the selected reference threshold. Once the system determines that a representative value in a certain direction is too high (e.g. above the reference threshold), the system may cease searching in that direction. Values above the reference threshold may represent edges.

[0153] Thus, the system would stop searching in a particular direction when such a value is reached. The group of searched values that are equal to or less than the selected reference cost threshold may generally correspond to an area of suitable values as determined by the system 1100 and may be surrounded by image units having cost map values greater than the reference cost threshold. This technique leads to the identification of one or more candidate segments, e.g., lines surrounding the area around the representative value for which each image unit satisfies the reference threshold. The candidate segments are lines that are identified as potentially corresponding to the edges of the physical objects represented by the image information.

[0154] The operation 1708, wherein an area of below -threshold cost values is “grown” from a reference cost value for a divided unit may be performed until all the entire cost map is segmented in this fashion. Some reference cost values may be encompassed by the below- threshold area grown from a different reference cost value. In some embodiments, the system 1100 may count such reference cost values as already dealt with. In some embodiments, the system 1100 may grow a new below -threshold area from the such reference cost values and combine or compare the resulting new below-threshold area to the previously grown below- threshold area.

[0155] In embodiments, the system 1100 may be configured to link or join neighboring areas or portions having cost values below the threshold. Such neighboring areas or portions may be separated by image data units having cost values above the threshold. In embodiments, the system may be configured to link or join such areas to generate larger portions. In embodiments, linking or joining may be based on the above-threshold separating points being close to the threshold, having values close to the neighboring below-threshold image data units, and / or being few in number.

[0156] In an operation 1710, a determination to repeat operations 1706-1708 may be performed. In embodiments, operations 1706-1708 may be repeated according to a determination made based on one or more criteria. In some embodiments, the system 1100 may determine to repeat operations 1706-1708 with an adjusted reference cost threshold.

[0157] In embodiments, the determination to repeat segmentation operations 1706-1708 may be based on a criterion related to candidate segment parameters. The determination may be performed based parameters associated with one or more identified candidate segment. If tire reference cost threshold is too low, the system 1100 may be unable to identify a large enough area of suitable values to create viable candidate segment. In such a situation, the system 1100 may select a higher reference cost threshold and repeat operations 1706-1708 to identify a new candidate segment. The candidate segment parameters on which tire determination is made may be related to the size of the candidate segment. For example, the system 1100 may identify one or more candidate segments that potentially represent an edge of an object or a portion thereof. If the system 1100 determines that the length of the candidate segments (or the length of a perimeter of the candidate segment) is too short relative to a length threshold, the system 1100 may determine that a larger reference cost threshold is required.

[0158] In embodiments, tire determination to repeat segmentation operations 1706-1708 may be based on a criterion related to the size of below-threshold value areas. The criterion may include comparing the size of one or more below-threshold areas to a minimum size, a maximum size, or both. In embodiments, the system 1100 may have stored the minimum size and maximum size and these may be based on the expected sizes of object. For example, the minimum size may be based on a smallest expected object and the maximum size may be based on a largest expected object. Because of the possibility of object occlusion, the minimum size may be predetermined as a percentage or fraction of the size of a smallest expected object. If there are an excess number of below-threshold areas that are below the minimum size and none that are within the min-max range, then the cost threshold is too low and the determination to repeat segmentation operations 1706-1708 may be made with a raised cost threshold. If there are below-threshold areas that are larger than the max range, then the threshold is too high, and the detennination to repeat segmentation operations 1706-1708 may be made with a lowered cost threshold. Operations 1706-1708 may be reiterated with adjusted cost thresholds until a target number of below-threshold areas within the range between the minimum size and maximum size is achieved.

[0159] In embodiments, the system 1100 may select an alternative reference cost threshold and perform operations 1706-1708 again for segmenting the entire cost map. In some embodiments, the system 1100 may select an alternative reference cost threshold for only the divided units and representative values for which the identified candidate segment parameters failed to meet the appropriate size thresholds.

[0160] In an operation 1712, the system 1100 may combine one or more candidate segments determined in the operations 1706-1708 to establish one or more reference polygons from the one or more candidate segments. An example reference polygon 1835 is represented in FIG. 18B.

[0161] In embodiments, candidate segments may be combined as follows to establish a reference polygon. The system 1100 may operate to add additional candidate segments to perform the combining. In some cases, this may result in the closure of and completion of polygons while, in other cases, it may result in the splitting of potential polygons, which may result in the generation of additional polygons.

[0162] In embodiments, the plurality of candidate segments may be altered based on a system determination that the spacing between two candidate segments (or ends thereof) is too short (e.g. below a length threshold). In such a case, the system 1100 may determine that an additional line segment should be drawn to bridge this short gap. This may result in a more solid polygon border. It may also result in a split of potential polygons. For example, in situations where a reference cost threshold may have been too low, a continuous edge may fail to be properly identified. The system 1100 may be unable to determine the edges corresponding to the entire length of the edge of an object. These uneven edges may result in a broken line representing the edge that would allow two potential reference polygons to fuse together. Joining the close candidate segments may therefore split the potential reference polygon.

[0163] As illustrated in FIG. 18C, more than one reference polygon may correspond to or represent a single object. For example, the physical object 1836 as shown in FIG. 18C corresponding to the reference polygon 1835 is larger than the reference polygon 1835. This result may occur based on one or more factors. First, as discussed above, the cost map is divided into units from which representative values are taken. A polygon established from the representative value of one divided unit may be separate from a polygon established from the representative value of a second divided unit. Second, as discussed above, different reference cost thresholds may be used when the polygons are established. This may lead to multiple different polygons corresponding to a single object. If multiple reference polygons are identified, polygons overlap, and / or a first polygon encapsulates a second polygon, the system 1100 may operate to select only one of these for future use. Among the multiple reference polygons, the selected polygon may be determined based on its size, e.g., within theminimum / maximum size range, close to an enveloping rectangle (e.g., as discussed below) and / or shape, e.g., whether the shape is similar to an enveloping rectangle.

[0164] The system 1100 may also be configured to generate enveloping rectangles for the one or more reference polygons. The enveloping rectangles are rectangles sized and shaped to encompass the entirety of a reference polygon. As shown in FIG. 18D, reference polygon 1835 is encompassed by enveloping rectangle 1841. The enveloping rectangles 1841, also referred to as bounding boxes, represent an estimate by the system 1100 of the object size and shape that may be represented by the identified reference polygon. In embodiments, the reference polygon may be selected as the smallest potential rectangle that can fully envelop the reference polygon. In embodiments, the reference polygon may include a buffer space such that it is slightly larger than the smallest potential rectangle.

[0165] In some embodiments, the processes 1706-1708 may be repeated when one or more reference polygons below a size threshold are identified or determined at operation 1710. To address this, a new, higher reference cost threshold may be selected and one or more new reference polygons may be established. Expanding the reference cost threshold allows the inclusion of more cost map values, thereby expanding the polygon.

[0166] In embodiments, the process 1712 may be repeated to continue identifying reference polygons. The system 1100 may operate to identify a plurality of reference polygons. In embodiments, the process 1712 may operate on all the identified below -threshold areas or candidate segments to identify a plurality of reference polygons. FIG. 18D illustrates image information for which a plurality of reference polygons have been identified.

[0167] In embodiments, the system 1100 may further be configured to determine several additional factors related to the identified reference polygons. The system 1 100 may be configured to identify a confidence parameter for one or more of the identified reference polygons. The confidence parameter may represent a confidence level (e.g., from 0 to 1) indicating how sure the system is that an identified reference polygon corresponds to the actual object it is intended to represent. The confidence parameter may be based on, for example, size parameters (including, e.g., whether the reference polygon has an area size in the minimum / maximum range, whether the reference polygon has an area size similar to the enveloping rectangle, whether the reference polygon has edge sizes similar to the enveloping rectangle, etc.) and shape parameters (including, e.g., whether the reference polygon has a shape similar to the enveloping rectangle).

[0168] In embodiments, the system 1100 may generate polygons that correspond to objects but do not correspond to the physical edges of the objects. For example, the generated polygons may be smaller than the physical objects. In embodiments, this may occur due to high cost values determined for object edges and / or due to the addition of a buffer zone by the system 1100. In embodiments, the system 1100 may generate the reference polygons with a buffer zone between the polygon and the identified candidate segments (e.g., representing potential object edges) so as to ensure that the motion planning aspects of the system 1100 do not attempt to pick or grasp an object to close to its edge, which might be interfered with by a neighboring occluded or occluding object.

[0169] Returning now to FIG. 5, the subprocess 508 for occlusion checking may be executed. The subprocess 508, as shown in FIG. 19, includes an operation 1902 for selecting a reference polygon, an operation 1904 for determining a region of interest, an operation 1906 for determining a gradient change, an operation 1708 for estimating an occlusion status, and an operation 1910 for repeating operations 1902 and 1908. FIGS. 20A-I help illustrate the subprocess 508 for occlusion checking.

[0170] In an operation 1902, for selecting a reference polygon, the system may select a reference polygon for which to perform an occlusion check. FIG. 20A illustrates an example reference polygon 1835 to illustrate the occlusion checking process.

[0171] In an operation 1904 for determining a region of interest, the system 1100 may select a segment 1852 of the reference polygon 1835 and determine a region of interest 1851 or a “search box,” around the selected segment 1852. The length of the region of interest 1851 may be aligned with the segment 1852 and may have a length approximately the same as the segment 1852. The width of the region of interest may be predetermined. For example, the width may be selected to have a sufficient width to accommodate potential errors (e.g., misalignment, segment not representing the physical edge of the object, etc.) without encompassing too many potential data points. The region of interest may have its center aligned with the segment or may be aligned to have one edge closer to the segment than the opposing edge.

[0172] In an operation 1906 for determining a data change, the system 1100 may be configured to identify a data change. Within the determined region of interest, the system may identify a change in data. For example, the system may determine a data change based on changes in cost values in one or more layers of the cost map, e.g., the object edge layer, theangle difference layer, the angle magnitude layer, and / or the depth layer. The data changes may be representative of a direction of curvature in the region of interest.

[0173] In the illustrated example of FIG. 20B, the system 1100 is configured to determine a change in the gradient angle of the normal vector data. More specifically, the system 1100 is configured to determine whether the gradient angle (or other curvature measure) is increasing or decreasing. As discussed above, gradient angle represents a direction of curvature. Increasing changes in gradient angle are illustrated by a thick solid line and decreasing changes in gradient angle magnitude are by a thick dashed line. The system 1100 may be configured to determine the changes in data, here, the changes in gradient angle in a direction along the width of the region of interest, e.g., in a direction substantially perpendicular to the segment on which the region of interest selection is based.

[0174] In an operation 1908 for estimating an occlusion status, the system 1100 may estimate an occlusion status based on the changes in data identified at operation 1906. The occlusion status may represent whether the polygon (and accordingly the object it represents) is occluding another object or is being occluded by another object. The system 1100 is configured to represent whether the object is overlapping another object or is being overlapped by an object. The system 1100 may be configured to estimate occlusion status according to whether the data change represents an outward facing corner or an inward facing comer. An outward facing corner would indicate that the identified corner belongs to the object being analyzed. An inward facing comer would indicate that the identified corner belongs to a neighboring object, e.g., an occluding object.

[0175] Identifying occlusion is an important factor in robotic motion planning. When picking certain items or types of items, it is preferable to pick an object that is not being occluded, even if the end effector can access it. Picking an overlapped object can result in the overlapping object being shifted or thrown by the overlapped object when the overlapped object is being moved. In the current example, occlusion status is determined according to the identified the change in data (e.g., the increasing / decreasing of the gradient angle magnitude). More specifically, the system 1100 is configured to determine the direction of the increasing and decreasing change in gradient angle.

[0176] The sy stem 1100 may select a point within the reference polygon 1835 and within the region of interest 1851. The system 1100 may then trace a path towards the edge of the reference polygon 1835 and opposite side of the region of interest 1851 and may analyze thegradient angle magnitude at each image unit along the path. In the present example, the gradient angle may first decrease (dotted line) and then increase (solid line). This decrease / increase pattern may represent that the candidate line segment corresponds to a positive (outwardly facing) corner. As the edge is approached, the angle of the surface normal vectors change from perpendicular to the image plane to parallel to the image plane, pointing in the direction of search. When measured in the direction of search, this change represents a decrease in the gradient angle. As the edge is passed, the surface normal vectors change from parallel back to perpendicular. In the direction of search, this results in changes in gradient angle that are opposite the changes from perpendicular to parallel, or increases in the gradient angle. This change, where the surface normal vectors go from perpendicular to parallel in the direction of search and back to perpendicular is indicative of a positive corner. A positive comer indicates that the object corresponding to the polygon 1835 is occluding another object (e.g., the next object in the direction of the traced path).

[0177] In embodiments, a depth discontinuity (e.g., based on depth separation values computed for the depth layer of the cost map) may optionally be used for detecting the type of comer or edge between two obj ects. The depth separation values on either side of the identified edge may provide information as to which object is on top of the other (and therefore occluding the other). In embodiments, depth discontinuity may be employed with the gradient angle magnitudes in combination. In embodiments, depth discontinuity may be employed by itself to determine occlusion status.

[0178] In contrast, in the region of interest 1853, the pattern is reversed and shows an increase (solid line) and then a decrease (dotted line) in gradient angle. This indicates that the surface normal vectors go from perpendicular to parallel in the opposite of the direction of search and back to perpendicular, which represents a negative or inward comer. This pattern indicates that the identified line segment represents a negative (inwardly projecting) comer. A negative comer indicates that the next object in the traced path is on top of, or occluding, the current object.

[0179] In an operation 1910 for repeating operations 1902 and 1908, the system may operate to determine the occlusion status of all or some of the identified reference polygons in the image information. By determining the occlusion status of the reference polygons, tire system 1100 may be configured to determine which associated objects are not occluded by other objects. This may permit the system to determine which item or items can be picked without disturbing the other items.

[0180] Another example of operations 1902-1910 is provided with reference to FIGS. 20E-20I. In operations 1902 and 1904, a line segment is identified and an associated region of interest is selected. As illustrated in FIG. 20E, the line segment 2036 of the reference polygon 2037 is identified. The region of interest 2035 is selected. In this example, the line segment 2036 is not centered with the identified region of interest. At operation 1906, changes in data are identified. In the example shown, the majority of the line segment 2036 within the region of interest 2035 is neither occluded nor occluding - there are no other neighboring objects. The majority of the line segment 2036 represents an outer edge of the stack of objects. If the background of the scene is removed, there is a portion of the edge that will only have a decreasing change in the gradient angle of the normal vector. Accordingly, only a dashed line is shown for the majority of the line segment 2036. However, there does remain a portion of this line segment 2036 for which the occlusion status should be determined.

[0181] FIG. 20F illustrates the same portion of the scene rotated 90°. Starting from reference point within reference polygon 2037, at the right hand side where the line segment 2036 does not border the background, the system 1100 may be configured to identify changes in data. The data indicates that the change in gradient angle first decreases (indicated by the dashed line) and then increases (indicated by the solid line). Accordingly, the system 1100 may determine that the object associated with the reference polygon 2037 is occluding the object overlapped by the upper-right corner of the region of interest, e.g., at operation 1908.

[0182] FIGS. 20G and 20H illustrate the operations 1902-1908 as conducted on another line segment of the reference polygon 2037. In FIG. 20G, the line segment is identified and a region of interest is selected (operations 1902, 1904). In FIG. 20H, the direction of change of data is identified, illustrated as an increase / decrease pattern. This pattern indicates that the object represented by the reference polygon 2037 is occluded by the object to its upper right in the image of FIG. 20H. The region of interest 2035 may be used in the determination of occlusion for the reference polygon 2037.

[0183] In some embodiments, the system 1100 may be configured to employ detection masks with respect to the identified reference polygons. Such masks may be used for display purposes and / or for further processing purposes. FIG. 201 represents such a detection mask for the reference polygon 2037. The system may apply the detection mask to represent the detected area of the object associated with a reference polygon (e.g., reference polygon 2037). Alternatively, the detection mask may correspond to an area of data values that satisfied the selected reference cost threshold.

[0184] As discussed above, operation 1910 may involve the repetition of the aspects of the occlusion checking subprocess. Such repetition may result in the generation of an occlusion information map containing information about the occlusion state of some, most, or all of the objects represented by the image information. In some embodiments, tire system 1100 may operate to identify occlusion states of all objects within the scene. The occlusion information map may be employed by the motion planning module 1129 to assist in robotic motion planning for object picking. In embodiments, the occlusion information map may be updated after each picking operation. In embodiments, updating may be performed by carrying out the appropriate steps of the occlusion determination process 500 with respect to only the scene area where the recently removed object was located. In some embodiments, the system 1100 may begin the occlusion determination process and may halt when an unoccluded object is identified. The system 1100 may then cause the picking of the unoccluded object before repeating the occlusion determination process 500 until another unoccluded object is identified.

[0185] Motion planning may include planning robotic motion, e.g., plotting trajectories, for a robot 3300 to carry out to retrieve the object 5012. Trajectories may be plotted so as to account for and avoid the identified obstacles. Motion execution may include sending commands related to the motion planning to a robot 3300 or robotic control system to cause the robot to execute the planned motion. Motion planning may be performed based on the occlusion map. In embodiments, motion planning may involve picking the least occluded objects spaced apart by a minimum distance. Picking a first object may be assumed to cause changes in the occlusion status of neighboring objects. The minimum distance may be selected based on object size and may be determined based on an expected distance required to minimize or eliminate disturbances to objects from the previous picking operation. Thus, the motion planning may be configured to pick several objects from different parts of the scene before reassessing the occlusion status of the objects.

[0186] It will be apparent to one of ordinary skill in the relevant arts that other suitable modifications and adaptations to the methods and applications described herein can be made without departing from the scope of any of the embodiments. The embodiments described above are illustrative examples and it should not be construed that the present disclosure is limited to these particular embodiments. It should be understood that various embodiments disclosed herein may be combined in different combinations than the combinations specifically presented in the description and accompanying drawings. It should also be understood that, depending on the example, certain acts or events of any of the processes or methods describedherein may be performed in a different sequence, may be added, merged, or left out altogether (e.g., all described acts or events may not be necessary to carry out the methods or processes). In addition, while certain features of embodiments hereof are described as being performed by a single component, module, or unit for purposes of clarity, it should be understood that the features and functions described herein may be performed by any combination of components, units, or modules. Thus, various changes and modifications may be affected by one skilled in the art without departing from the spirit or scope of the invention as defined in the appended claims.

[0187] Further embodiments include:

[0188] Embodiment 1 is a computing system configured to identify object occlusions within a scene comprising: a control system configured to communicate with a robot and to communicate with a camera; at least one processing circuit configured for: obtaining image information captured by the camera of a plurality of objects in the scene; identifying and removing a background portion of the image information; generating a cost map based on the image information, the cost map being indicative of surface variations among the objects in the scene; segmenting the cost map according to a comparison between cost map values and a reference cost threshold; and identifying object occlusions between segmented polygons of the cost map.

[0189] Embodiment 2 is the computing system of embodiment 1, wherein the at least one processing circuit is further configured to generate a multilayer cost map, wherein a first layer of the cost map is determined from identified object edges and surfaces, a second layer of the cost map is determined according to gradient angle magnitudes, and a third layer of the cost map is determined according to gradient angles.

[0190] Embodiment 3 is the computing system of embodiment 2, wherein the gradient angle magnitudes represent changing magnitudes of surface normal vectors of objects in the scene.

[0191] Embodiment 4 is the computing system of embodiment 2, wherein the gradient angles represent changing angles of surface normal vectors of objects in the scene.

[0192] Embodiment 5 is the computing system of embodiment 2, wherein the cost map is generated according to weighting factors applied to the first layer, the second layer, and the third layer.

[0193] Embodiment 6 is the computing system of embodiment 2, wherein the gradient angle magnitudes are modified according to weights determined from the gradient angles.

[0194] Embodiment 7 is the computing system of embodiment 1, wherein the at least one processing circuit is further configured to segment the cost map by: dividing the cost map into a plurality of units; identifying representative cost map values within each of the plurality of units; determining candidate line segments corresponding to each of the representative cost map values according to a reference cost threshold.

[0195] Embodiment 8 is the computing system of embodiment 2, wherein the at least one processing circuit is further configured to segment the cost map by: combining candidate line segments to generate reference polygons.

[0196] Embodiment 9 is the computing system of embodiment 1, wherein the at least one processing circuit is further configured to identify occlusions by identifying a region of interest encompassing a line segment; and identifying gradient changes in the region of interest.

[0197] Embodiment 10 is the computing system of embodiment 10, wherein the at least one processing circuit is further configured to identify occlusions by: identifying a directionality of the gradient changes.

[0198] Embodiment 11 is a method of identifying object occlusions within a scene performed by a control system having at least one processing circuit and being configured to communicate with a robot and to communicate with a camera, the method comprising: obtaining image information captured by the camera of a plurality of objects in the scene; identifying and removing a background portion of the image information; generating a cost map based on the image information, the cost map being indicative of surface variations among the objects in the scene; segmenting the cost map according to a comparison between cost map values and a reference cost threshold; and identifying object occlusions between segmented polygons of the cost map.

[0199] Embodiment 12 is the method of embodiment 11, further comprising: generating a multilayer cost map, wherein a first layer of the cost map is determined from identified object edges and surfaces, a second layer of the cost map is determined according to gradient angle magnitudes, and a third layer of the cost map is determined according to gradient angles.

[0200] Embodiment 13 is the method of embodiment 12, wherein the gradient angle magnitudes represent changing magnitudes of surface normal vectors of objects in the scene.

[0201] Embodiment 14 is the method of embodiment 12, wherein the gradient angles represent changing angles of surface normal vectors of objects in the scene.

[0202] Embodiment 15 is the method of embodiment 12, wherein the cost map is generated according to weighting factors applied to the first layer, the second layer, and the third layer.

[0203] Embodiment 16 is the method of embodiment 12, wherein the gradient angle magnitudes are modified according to weights determined from the gradient angles.

[0204] Embodiment 17 is the method of embodiment 11, further comprising segmenting the cost map by: dividing the cost map into a plurality of units; identifying representative cost map values within each of the plurality of units; determining candidate line segments corresponding to each of the representative cost map values according to a reference cost threshold.

[0205] Embodiment 18 is the method of embodiment 12, further comprising segmenting the cost map by: combining candidate line segments to generate reference polygons.

[0206] Embodiment 19 is the c method of embodiment 11, further comprising identifying occlusions by identifying a region of interest encompassing a line segment; and identifying gradient changes in the region of interest.

[0207] Embodiment 20 is the method of embodiment 20, further comprising identifying occlusions by: identifying a directionality of the gradient changes.

Claims

Claims:

1. A computing system configured to identify object occlusions within a scene comprising: a control system configured to communicate with a robot and to communicate with a camera; at least one processing circuit configured for: obtaining image information captured by the camera of a plurality of objects in the scene; identifying and removing a background portion of the image information; generating a cost map based on the image information, the cost map being indicative of surface variations among the objects in the scene; segmenting the cost map according to a comparison between cost map values and a reference cost threshold; and identifying object occlusions between segmented polygons of the cost map.

2. The computing system of claim 1, wherein the cost map includes a plurality of layers, wherein a first layer of the cost map is determined from identified object edges and surfaces, a second layer of the cost map is determined according to a curvature measure, and a third layer of the cost map is determined according to depth values.

3. The computing system of claim 2, wherein the curvature measure includes gradient angle magnitudes representing changing angles between surface normal vectors of objects in the scene.

4. The computing system of claim 2, wherein the cost map further includes a fourth layer determined from gradient angles representing angles of gradients between surface normal vectors of objects in the scene.

5. The computing system of claim 2, wherein the cost map is generated according to weighting factors applied to the first layer, the second layer, and the third layer.

6. The computing system of claim 2, wherein the curvature measure is modified according to weights determined from gradient angles.

7. The computing system of claim 1, wherein the at least one processing circuit is further configured to segment the cost map by: dividing the cost map into a plurality of units; identifying representative cost map values within each of the plurality of units; determining candidate line segments corresponding to each of the representative cost map values according to a reference cost threshold.

8. The computing system of claim 2, wherein the at least one processing circuit is further configured to segment the cost map by: combining candidate line segments to generate reference polygons.

9. The computing system of claim 1, wherein the at least one processing circuit is further configured to identify occlusions by identifying a region of interest encompassing a line segment; and identifying data changes in the region of interest.

10. The computing system of claim 9, wherein the at least one processing circuit is further configured to identify occlusions by: identifying a direction of curvature according to the data changes.

11. A method of identifying object occlusions within a scene performed by a control system having at least one processing circuit and being configured to communicate with a robot and to communicate with a camera, the method comprising: obtaining image information captured by the camera of a plurality of objects in the scene; identifying and removing a background portion of the image information; generating a cost map based on the image information, the cost map being indicative of surface variations among the objects in the scene; segmenting the cost map according to a comparison between cost map values and a reference cost threshold; and identifying object occlusions between segmented polygons of the cost map.

12. The method of claim 11, further comprising: generating a multilayer cost map, wherein the cost map includes a plurality of layers, wherein a first layer of the cost map is determined from identified object edges and surfaces, a second layer of the cost map is determined according to a curvature measure, and a third layer of the cost map is determined according to depth values.

13. The method of claim 12, wherein the curvature measure includes gradient angle magnitudes representing changing angles between surface normal vectors of objects in the scene.

14. The method of claim 12, wherein the cost map further includes a fourth layer determined from gradient angles representing angles of gradients between surface normal vectors of objects in the scene.

15. The method of claim 12, wherein the cost map is generated according to weighting factors applied to the first layer, the second layer, and tire third layer.

16. The method of claim 12, wherein the curvature measure is modified according to weights determined from gradient angles.

17. The method of claim 11, further comprising segmenting the cost map by: dividing the cost map into a plurality of units; identifying representative cost map values within each of the plurality of units; determining candidate line segments corresponding to each of the representative cost map values according to a reference cost threshold.

18. The method of claim 12, further comprising segmenting the cost map by: combining candidate line segments to generate reference polygons.

19. The method of claim 11, further comprising identifying occlusions by identifying a region of interest encompassing a line segment; and identifying data changes in the region of interest.

20. The method of claim 19, further comprising identifying occlusions by : identifying a directionality of the data changes.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, and program

    US20190364224A1

  • Systems and methods for robotic system with object handling

    US20230286165A1

  • Stylized glyphs using generative ai

    US20240127510A1