Method and computing system for performing or facilitating physical edge detection

The computing system addresses the challenge of detecting physical edges by using a combination of 2D and 3D image analysis to evaluate candidate edges, enhancing detection accuracy and enabling more precise robot interactions.

JP7694885B2Active Publication Date: 2025-06-18MUJIN INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022118399
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-27
Filing Date
2022-07-26
Publication Date
2025-06-18
Estimated Expiration
2041-06-03

AI Technical Summary

Technical Problem

Existing technologies face challenges in accurately detecting physical edges of objects in complex environments, particularly in distinguishing between actual physical edges and false edges.

Method used

A computing system that communicates with a robot and a camera, processing image information to identify candidate edges and determining their confidence level by evaluating darkness conditions defined by these edges, using both 2D and 3D image information to enhance detection accuracy.

Benefits of technology

The system effectively distinguishes physical edges from false edges, improving the accuracy of object recognition and robot interaction by using a combination of 2D and 3D image analysis to compensate for limitations in each type of image information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694885000006
    Figure 0007694885000006
  • Figure 0007694885000007
    Figure 0007694885000007
  • Figure 0007694885000008
    Figure 0007694885000008
Patent Text Reader

Abstract

To provide a computing system that uses image information representing a group of objects to detect or otherwise identify physical edges of the group of objects. A computing system includes a processing circuit configured to receive image information representing a group of objects within a camera field of view and identify from the image information a plurality of candidate edges associated with the group of objects. If the plurality of candidate edges includes a first candidate edge formed based on a boundary between a first image region and a darker second image region, the computing system determines whether the image information satisfies a darkness condition defined by the first candidate edge. A subset of the plurality of candidate edges is selected to form a selected subset of candidate edges for representing a physical edge of the group of objects. The computing system determines whether to retain the first candidate edge as a candidate for representing at least one of the physical edges of the group of objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the priority of U.S. Patent Application No. 17 / 331,878, filed on May 27, 2021, entitled "METHOD AND COMPUTING SYSTEM FOR PERFORMING OR FACILITATING PHYSICAL EDGE DETECTION", which claims the priority of U.S. Provisional Patent Application No. 63 / 034,403, filed on June 4, 2020, entitled "ROBOTIC SYSTEM WITH VISION MECHANISM", and the entire content of which is incorporated herein by reference.

[0002] The present disclosure relates to computing systems and methods for performing or facilitating physical edge detection.

Background Art

[0003] As automation becomes more common, robots are used in more environments, such as warehouse storage and retail settings. For example, robots can be used to interact with objects within a warehouse. The operation of the robot may be constant or based on inputs such as information generated by sensors within the warehouse.

Summary of the Invention

[0004] One aspect of the present disclosure relates to a computing system, or a method performed by a computing system. The computing system may include a communication interface and at least one processing circuit. The communication interface may be configured to communicate with a robot and a camera having a camera field of view. The at least one processing circuit is configured to receive, when a group of objects is within the camera field of view, image information representing the group of objects generated by the camera, and to identify, from the image information, a plurality of candidate edges associated with the group of objects, the plurality of candidate edges being respective sets of image positions or physical positions that form respective candidates for representing physical edges of the group of objects, or including them. When the plurality of candidate edges includes a first candidate edge formed based on a boundary between a first image region and a second image region, determining whether the image information satisfies a darkness condition defined by the first candidate edge, wherein the first image region is darker than the second image region, and the first image region and the second image region are respective regions described by the image information, and selecting a subset of the plurality of candidate edges to form a selected subset of candidate edges for representing physical edges of the group of objects, the selecting being based on whether the image information satisfies the darkness condition defined by the first candidate edge, and including the first candidate edge in the selected subset of candidate edges to determine whether to retain the first candidate edge as a candidate representing at least one of the physical edges of the group of objects.

Brief Description of the Drawings

[0005]

Figure 1A

Figure 1B

Figure 1C

Figure 1D

[0006]

Figure 2A

Figure 2B

Figure 2C

Figure 2D

[0007]

Figure 3A

Figure 3B

Figure 3C

[0008]

Figure 4

[0009]

Figure 5A

Figure 5B

[0010]

Figure 6A

Figure 6B

[0011]

Figure 7A

Figure 7B

Figure 7C

[0012]

Figure 8

[0013]

Figure 9A

Figure 9B

Figure 9C

Figure 9D

[0014]

Figure 10A

Figure 10B

Figure 10C

[0015]

Figure 11A

Figure 11B

Figure 11C

Figure 11D

[0016]

Figure 12A

Figure 12B

Figure 12C

[0017]

Figure 13A

Figure 13B

Figure 13C

Mode for Carrying Out the Invention

[0018] One aspect of the present disclosure relates to detecting or otherwise identifying physical edges of a group of objects using image information representing the group of objects. For example, a 2D image may represent a group of boxes and may include candidate edges that potentially represent physical edges of the group of boxes. A computing system may use the candidate edges within the image information to distinguish individual objects represented within the image information. In some examples, the computing system may use information identifying individual boxes to control robotic interactions involving the individual boxes. For example, robotic interactions may include an operation of picking up an object from a pallet, where an end effector device of a robot approaches one of the objects, picks up the object, and moves the object to a destination location.

[0019] In some scenarios, a 2D image or other image information may include candidate edges that are false edges, which may be candidate edges that do not correspond to actual physical edges of objects within a camera field of view. Accordingly, one aspect of the present disclosure relates to evaluating candidate edges to determine a confidence level that a candidate edge corresponds to an actual physical edge as opposed to being a false edge. In embodiments, such determination may be based on an expectation or prediction regarding how a particular physical edge is likely to appear in an image. More specifically, such determination may be based on the expectation that when a physical edge is associated with a physical gap between objects (e.g., the physical edge forms one side of the physical gap), such physical gap may appear very dark in the image and / or may have an image intensity profile characterized by a spike decrease in image intensity of an image region corresponding to the physical gap. Accordingly, the method or computing system of the present disclosure may operate based on the expectation that physical gaps between objects, particularly narrow physical gaps, may be represented by an image having specific characteristics related to how dark the physical gap is within the image. Such features or characteristics of the image may be referred to as dark priors, and the present disclosure may be related to detecting dark priors, and the presence of a dark prior may increase the confidence level regarding whether a candidate edge corresponds to an actual physical edge.

[0020] In an embodiment, the method or system of the present disclosure may determine whether an image meets darkness conditions defined by candidate edges, and the defined darkness conditions may be related to detecting a dark prior. More specifically, the defined darkness conditions may be defined by a darkness threshold criterion, and / or a spike intensity profile criterion, which will be discussed in more detail below. In this embodiment, when a computing system or method determines that an image meets the darkness conditions defined by candidate edges, the candidate edges may have a higher level of confidence corresponding to an actual physical edge, such as a physical edge forming one side of a physical gap between two objects. In some examples, when an image does not meet the darkness conditions defined by candidate edges, the candidate edges may be more likely to be false edges.

[0021] One aspect of the present disclosure relates to using 2D image information to compensate for limitations in 3D image information, and vice versa. For example, when multiple objects, such as two or more boxes, are placed in close proximity to each other and separated by a narrow physical gap, the 3D image information may not have a high enough resolution to capture or otherwise represent the physical gap. Thus, the 3D image information may have limitations in its ability to be used to distinguish individual objects of a group of objects, particularly when the multiple objects have the same depth relative to the camera that generates the 3D image information. In such examples, the physical gap between the multiple objects may be represented in the 2D image information. More specifically, the physical gap may be represented by an image region that meets defined darkness conditions. Thus, candidate edges associated with such image regions may represent physical edges of the objects with a high level of reliability. In such situations, candidate edges in the 2D image information may be useful for distinguishing individual objects of a group of objects. Thus, 2D image information can enhance the ability to distinguish individual objects in certain situations.

[0022] In certain situations, 3D image information may compensate for limitations of 2D image information. For example, a 2D image may not satisfy darkness conditions defined by certain candidate edges in the 2D image. In such examples, the candidate edges may have a low confidence level corresponding to any actual physical edge object within the camera's field of view. 3D image information may be used to compensate for this limitation in 2D image information when candidate edges in the 2D image information correspond to candidate edges in the 3D image information. More specifically, candidate edges in the 2D image information may be mapped to a position or set of positions in the 3D image information where there is a sharp change in depth. In such situations, 3D image information may be used to increase the confidence level that candidate edges in the 2D image information correspond to actual physical edges.

[0023] In an embodiment, 3D image information may be used to identify the surface of an object (e.g., the top surface), and candidate edges may be identified based on positions where there is a transition between two surfaces. For example, a surface may be identified based on a set of positions having respective depth values in the 3D image information that do not deviate from each other beyond a defined measurement variance threshold. The defined measurement variance threshold may account for the effects of imaging noise, manufacturing tolerances, or other factors that may introduce random variations in the depth measurements of the 3D image information. The identified surfaces may be associated with a depth value that is the average of the respective depth values. In some implementations, candidate edges may be detected in the 3D image information based on identifying a depth transition between two surfaces identified in the 3D image information that exceeds a defined depth difference threshold.

[0024] FIG. 1A shows a system 1000 for performing or facilitating physical edge detection that may involve using image information representing one or more objects to detect or otherwise identify physical edges of the one or more objects. More particularly, system 1000 may include a computing system 1100 and a camera 1200. In this example, camera 1200 may be configured to generate image information depicting or otherwise representing the environment in which camera 1200 is located, or more specifically, the environment within the field of view of camera 1200 (also referred to as the camera field of view). The environment may be, for example, a warehouse, a manufacturing facility, a retail space, or some other facility. In such examples, the image information may represent objects located in such facilities, such as boxes, bins, cases, wooden frames, or other containers. System 1000 may be configured to generate, receive, and / or process the image information, such as to distinguish individual objects within the camera field of view, perform object recognition or object registration based on the image information, and / or perform robot motion planning based on the image information (the terms “and / or” and “or” are used interchangeably in this disclosure). Robot motion planning may be used, for example, to control a robot in a facility to facilitate robot interaction with containers or other objects. Computing system 1100 and camera 1200 may be located in the same facility or may be located remotely from each other. For example, computing system 1100 may be part of a cloud computing platform hosted at a data center remote from a warehouse or retail space and may communicate with camera 1200 via a network connection.

[0025] In an embodiment, the camera 1200 (also referred to as an image sensing device) can be a 2D camera and / or a 3D camera. For example, FIG. 1B shows a system 1000A (which can be an embodiment of the system 1000) that includes a computing system 1100, as well as cameras 1200A and 1200B (both of which can be embodiments of the camera 1200). In this example, the camera 1200A can be a 2D camera configured to generate 2D image information that includes or forms a 2D image that depicts the visual appearance of the environment within the camera's field of view. The camera 1200B can be a 3D camera (also referred to as a spatial structure sensing camera or a spatial structure sensing device) configured to generate 3D image information that includes or forms spatial structure information regarding the environment within the camera's field of view. The spatial structure information may include depth information (e.g., a depth map) that describes the respective depth values of various positions with respect to the camera 1200B, such as positions on the surfaces of various objects within the field of view of the camera. These positions of the camera's field of view or the surface of an object may also be referred to as physical positions. The depth information in this example can be used to estimate how objects are spatially arranged within a three-dimensional (3D) space. In some examples, the spatial structure information may include or be used to generate a point cloud that describes positions on one or more surfaces of an object within the field of view of the camera 1200B. More specifically, the spatial structure information can describe various positions on the structure of an object (also referred to as the object structure).

[0026] In an embodiment, the system 1000 can be a robot operation system for facilitating robot interactions between a robot and various objects in the environment of the camera 1200. For example, FIG. 1C shows a robot operation system 1000B that can be an embodiment of the system 1000 / 1000A of FIGS. 1A and 1B. The robot operation system 1000B may include a computing system 1100, a camera 1200, and a robot 1300. As described above, the robot 1300 can be used to interact with one or more objects in the environment of the camera 1200, such as boxes, wooden frames, bins, or other containers. For example, the robot 1300 can be configured to pick up containers from one location and move them to another location. In some cases, the robot 1300 can be used to perform an operation of unloading a group of containers or other objects from a pallet, for example, where the group is lowered and moved, for example, onto a conveyor belt. In some implementations, the camera 1200 may be attached to the robot 1300, such as to a robot arm of the robot 1300. In some implementations, the camera 1200 can be separated from the robot 1300. For example, the camera 1200 may be mounted on the ceiling of a warehouse or other structure and may remain stationary with respect to the structure.

[0027] In an embodiment, the computing system 1100 of FIGS. 1A-1C may form, or may be part of, a robot control system (also referred to as a robot controller) that is part of the robot operating system 1000B. The robot control system may be a system configured to generate commands for the robot 1300, such as robot interaction movement commands for controlling the robot interaction between the robot 1300 and a container or other object. In such embodiments, the computing system 1100 may be configured to generate such commands based on, for example, the image information generated by the cameras 1200 / 1200A / 1200B. For example, the computing system 1100 may be configured to determine a motion plan based on the image information, and the motion plan may be intended to, for example, grasp an object or pick it up in some other way. The computing system 1100 may generate one or more robot interaction movement commands to execute the motion plan.

[0028] In an embodiment, the computing system 1100 may form, or may be part of, a vision system. The vision system may be a system configured to generate vision information that describes the environment in which the robot 1300 is located, or more specifically, the environment in which the camera 1200 is located. The vision information may include the 3D image information, and / or 2D image information, or some other image information discussed above. In some scenarios, when the computing system 1100 forms the vision system, the vision system may be part of the robot control system discussed above, or may be separated from the robot control system. When the vision system is separated from the robot control system, the vision system may be configured to output information that describes the environment in which the robot 1300 is located. The information may be output to a robot control system that can receive such information from the vision system and, based on the information, execute a motion plan and / or generate robot interaction movement commands.

[0029] In an embodiment, the computing system 1100 can communicate with the camera 1200 and / or the robot 1300 by direct connection, such as via a dedicated wired communication interface such as an RS-232 interface, a Universal Serial Bus (USB) interface, and / or a connection provided via a local computer bus such as a Peripheral Component Interconnect (PCI) bus. In an embodiment, the computing system 1100 can communicate with the camera 1200 and / or the robot 1300 via a network. The network can be any type and / or form of network, such as a Personal Area Network (PAN), a Local Area Network (LAN), such as an intranet, a Metropolitan Area Network (MAN), a Wide Area Network (WAN), or the Internet. The network can utilize different technologies, and layers or stacks of protocols, including, for example, Ethernet protocol, Internet protocol suite (TCP / IP), ATM (Asynchronous Transfer Mode) technology, SONET (Synchronous Optical Networking) protocol, or SDH (Synchronous Digital Hierarchy) protocol.

[0030] In an embodiment, the computing system 1100 may communicate directly with the camera 1200 and / or the robot 1300, or may communicate via an intermediate storage device, or more generally, an intermediate non-transitory computer-readable medium. For example, FIG. 1D may illustrate an embodiment of the system 1000 / 1000A / 1000B that includes a non-transitory computer-readable medium 1400 that may be external to the computing system 1100, for example, a system 1000C that may act as an external buffer or repository for storing image information generated by the camera 1200. In such an example, the computing system 1100 may retrieve or otherwise receive the image information from the non-transitory computer-readable medium 1400. Examples of non-transitory computer-readable media include electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. The non-transitory computer-readable medium may form, for example, a computer diskette, a hard disk drive (HDD), a solid state drive (SDD), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), and / or a memory stick.

[0031] As described above, camera 1200 can be a 3D camera and / or a 2D camera. The 2D camera can be configured to generate a 2D image, such as a color image or a grayscale image. The 3D camera can be, for example, a depth sensing camera such as a time-of-flight (TOF) camera or a structured light camera, or any other type of 3D camera. In some cases, the 2D camera and / or the 3D camera can include an image sensor, such as a charge-coupled device (CCD) sensor and / or a complementary metal-oxide semiconductor (CMOS) sensor. In an embodiment, the 3D camera can include a laser, a LIDAR device, an infrared device, a light / dark sensor, a motion sensor, a microwave detector, an ultrasonic detector, a radar detector, or any other device configured to capture depth information or spatial structure information.

[0032] As described above, the image information can be processed by the computing system 1100. In an embodiment, the computing system 1100 can include, or be configured as, a server (e.g., having one or more server blades, processors, etc.), a personal computer (e.g., a desktop computer, a laptop computer, etc.), a smartphone, a tablet computing device, and / or any other arbitrary computing system. In an embodiment, all of the functionality of the computing system 1100 may be performed as part of a cloud computing platform. The computing system 1100 can be a single computer device (e.g., a desktop computer) or can include multiple computer devices.

[0033] FIG. 2A provides a block diagram showing an embodiment of a computing system 1100. The computing system 1100 includes at least one processing circuit 1110 and a non-transitory computer-readable medium (or media) 1120. In an embodiment, the processing circuit 1110 includes one or more processors, one or more processing cores, a programmable logic controller (“PLC”), an application specific integrated circuit (“ASIC”), a programmable gate array (“PGA”), a field programmable gate array (“FPGA”), any combination thereof, or any other processing circuit.

[0034] In an embodiment, a non-transitory computer-readable medium 1120, which is part of the computing system 1100, may be an alternative or addition to the intermediate non-transitory computer-readable medium 1400 discussed above. The non-transitory computer-readable medium 1120 may be a storage device such as an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, for example, a computer diskette, a hard disk drive (HDD), a solid state drive (SSD), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, any combination thereof, or any other storage device, etc. In some examples, the non-transitory computer-readable medium 1120 may include a plurality of storage devices. In a particular implementation, the non-transitory computer-readable medium 1120 is configured to store image information generated by the camera 1200 and received by the computing system 1100. In some examples, the non-transitory computer-readable medium 1120 may store one or more object recognition templates used to perform an object recognition operation. When executed by the processing circuit 1110, the non-transitory computer-readable medium 1120 may alternatively or additionally store computer-readable program instructions that cause the processing circuit 1110 to perform one or more techniques described herein, such as the operations described with respect to FIG. 4.

[0035] FIG. 2B depicts a computing system 1100A, which is an embodiment of the computing system 1100 and includes a communication interface 1130. The communication interface 1130 can be configured to receive, for example, image information generated by the cameras 1200 of FIGS. 1A - 1D. The image information can be received via the intermediate non - transitory computer - readable medium 1400 or network discussed above, or via a more direct connection between the camera 1200 and the computing system 1100 / 1100A. In an embodiment, the communication interface 1130 can be configured to communicate with the robot 1300 of FIG. 1C. If the computing system 1100 is external to the robot control system, the communication interface 1130 of the computing system 1100 can be configured to communicate with the robot control system. The communication interface 1130 may also be referred to as a communication component or communication circuit and may include, for example, a communication circuit configured to communicate by means of a wired or wireless protocol. As an example, the communication circuit may include an RS - 232 port controller, a USB controller, an Ethernet controller, a Bluetooth® controller, a PCI bus controller, any other communication circuit, or a combination thereof.

[0036] In an embodiment, in FIG. 2C, the non-transitory computer-readable medium 1120 may store edge detection information 1126 that can describe a plurality of candidate edges identified from the image information generated by the camera 1200. As discussed in more detail below, when the image information represents a group of objects, each of the candidate edges may be a candidate to represent at least one of the plurality of physical edges of the group of objects, or may form a candidate. In some examples, the computing system 1100 / 1100A / 1100B may determine whether to use a particular candidate edge within the edge detection information 1126 to represent at least one of the physical edges of the group of objects. Such a determination may involve evaluating a confidence level related to whether the candidate edge actually represents a physical edge, as opposed to a false edge. In one example, such an evaluation may be based on whether the candidate edge is associated with image characteristics resulting from representing a physical edge. Such features may be associated with image features referred to as dark priors, which are discussed in more detail below. In some scenarios, the computing system 1100 may select a subset of candidate edges having a sufficiently high confidence level that actually represent the physical edges of the group of objects from the plurality of candidate edges, while candidate edges excluded from the subset may not have a sufficiently high confidence level to represent the physical edges of the group of objects. Thus, if the computing system 1100 / 1100A / 1100B determines to use a particular candidate edge to represent at least one of the physical edges, the computing system may include the candidate edge in the subset. If the computing system 1100 / 1100A / 1100B determines not to use a particular candidate edge to represent at least one of the physical edges, the computing system may determine not to include the candidate edge in the subset. Candidate edges not included in the subset may be removed from the edge detection information 1126, or more broadly, may be excluded from further consideration as candidates to represent at least one of the physical edges of the group of objects.

[0037] In an embodiment, the processing circuit 1110 can be programmed by one or more computer-readable program instructions stored in the non-transitory computer-readable medium 1120. For example, FIG. 2D shows a computing system 1100C, an embodiment of the computing systems 1100 / 1100A / 1100B, in which the processing circuit 1110 is programmed by one or more modules including a physical edge detection module 1125, an object recognition / registration module 1128, and / or a motion planning module 1129.

[0038] In an embodiment, the physical edge detection module 1125 can be configured to determine which candidate edges to use to represent the physical edges of a group of objects from among a plurality of candidate edges that appear in the image information representing the group of objects. In some implementations, the physical edge detection module 1125 can perform such determination based on whether defined darkness conditions are met and / or whether depth discontinuity conditions are met, as discussed in more detail below. In some examples, the physical edge detection module 1125 can also be configured to identify a plurality of candidate edges from the image information. In some examples, the physical edge detection module 1125 can be configured to perform image segmentation (e.g., point cloud segmentation) that may involve distinguishing individual objects represented by the image information. For example, the module 1125 can extract or otherwise identify an image segment (also referred to as an image portion) of the image information representing one object of the group of objects. In some implementations, the image segmentation may be performed, for example, based on candidate edges that the module 1125 has determined should be used to represent the physical edges of the group of objects.

[0039] In an embodiment, the object recognition / registration module 1128 may be configured to execute an object recognition operation or an object registration module based on the results from the physical edge detection module 1125. For example, when the physical edge detection module 1125 identifies an image segment representing one object of a group of objects, the object recognition / registration module 1128 may be configured to, for example, determine whether the image segment sufficiently matches an object recognition template and / or generate a new object recognition template based on the image segment.

[0040] In an embodiment, the motion planning module 1129 may be configured to execute a robot motion plan based on the results of the physical edge detection module 1125 and / or based on the results of the object recognition / registration module 1128. As described above, the robot motion plan may be for robot interaction between a robot (e.g., 1300) and at least one object of a group of objects. In some examples, the robot motion plan may involve, for example, the movement of a component of the robot (e.g., an end effector device) to pick up an object and / or the determination of the trajectory of subsequent components after picking up the object.

[0041] In various embodiments, the terms "computer-readable instructions" and "computer-readable program instructions" are used to describe software instructions or computer code configured to perform various tasks and operations. In various embodiments, the term "module" broadly refers to a collection of software instructions or code configured to cause a processing circuit 1110 to perform one or more functional tasks. Modules and computer-readable instructions may be described as performing various operations or tasks when a processing circuit or other hardware component is executing the module or computer-readable instructions.

[0042] Figures 3A - 3C show the processing of candidate edges, or more specifically, an exemplary environment in which physical edge detection can be performed. More specifically, FIG. 3A depicts a system 3000 (which can be an embodiment of the system 1000 / 1000A / 1000B / 1000C of FIGS. 1A - 1D) that includes a computing system 1100, a robot 3300, and a camera 3200. The camera 3200 may be an embodiment of the camera 1200 and is configured to generate image information representing a scene within the camera's field of view 3210, or more specifically, objects within the camera's field of view 3210 such as objects 3510, 3520, 3530, 3540, and 3550. In one example, each of the objects 3510 - 3540 may be a container such as a box or a wooden frame, while the object 3550 may be, for example, a pallet on which the containers are placed.

[0043] Objects 3510 - 3540 are shown in FIG. 3B, which more specifically shows the physical edges of the objects. More specifically, the figure shows physical edge portions 3510A - 3510D of the upper surface of object 3510, physical edge portions 3520A - 3520D of the upper surface of object 3520, physical edge portions 3530A - 3530D of the upper surface of object 3530, and physical edge portions 3540A - 3540D of the upper surface of object 3540. The physical edges in FIG. 3B (e.g., 3510A - 3510D, 3520A - 3520D, 3530A - 3530D, and 3540A - 3540D) can be the outer edges of the respective upper surfaces of objects 3510 - 3540. In some examples, the physical edges of the surface of an object (e.g., 3510A - 3510D) can define the contour of the surface. When an object forms a polyhedron (e.g., a cube) having a plurality of non - coplanar surfaces (also referred to as a plurality of faces), the physical edge of one surface can form a boundary where the surface intersects another surface of the object.

[0044] In an embodiment, an object within the camera's field of view may have visual details (also referred to as visible details), such as visual markings, on the outer surface of the object. For example, in FIGS. 3A and 3B, objects 3510, 3520, 3530, 3540 may have visual markings 3512, 3522, 3532, 3542 that are printed or otherwise disposed on the respective outer surfaces (e.g., upper surfaces) of objects 3510 - 3540. As an example, the visual markings may include visible shapes such as visible lines (e.g., straight or curved), polygons, visual patterns, or other visual markings. In some scenarios, the visual markings (e.g., visible lines) may form or be part of a symbol or drawing displayed on the outer surface of the object. The symbol may include, for example, a logo or characters (e.g., alphanumeric). In some scenarios, the visual details on the outer surface of a container or other object may be formed by the outline of a layer of material (e.g., a strip of packaging tape or a sheet of mailing label) disposed on the outer surface of the container.

[0045] In an embodiment, system 3000 of FIG. 3A may include one or more light sources, such as light source 3600. Light source 3600 may be, for example, a light emitting diode (LED), a halogen lamp, or any other light source, and may be configured to emit visible light, infrared light, or any other form of light towards the surfaces of objects 3510 - 3550. In some embodiments, computing system 1100 may be configured to communicate with light source 3600 to control when light source 3600 is activated. In other implementations, light source 3600 may operate independently of computing system 1100.

[0046] In an embodiment, as shown in FIG. 3C, system 3000 may include a camera 3200A (which may be an embodiment of camera 1200A) having a camera field of view 3210A and a camera 3200B (which may be an embodiment of camera 1200B) having a camera field of view 3210B, and may include a plurality of cameras. Camera 3200A may be, for example, a 2D camera configured to generate a 2D image or other 2D image information, while camera 3200B may be, for example, a 3D camera configured to generate 3D image information. The 2D image (e.g., a color image or a grayscale image) may depict the appearance of one or more objects such as objects 3510 - 3550 in camera field of view 3210 / 3210A. For example, the 2D image may capture visual details such as visual markings 3512 - 3542 disposed on the outer surface (e.g., the upper surface) of objects 3510 - 3540 and / or the contours of their outer surfaces, or may represent them in other ways. In an embodiment, the 3D image information may describe the structure of one or more of objects 3510 - 3550, which may also be referred to as the structure of the object or the physical structure of the object. For example, the 3D image information may include a depth map, and more generally, may include depth information that describes the respective depth values of various positions in camera field of view 3210 / 3210B relative to camera 3200B or some other reference point. The positions corresponding to the respective depth values may be positions on various surfaces (also referred to as physical positions) in camera field of view 3210 / 3210B, such as positions on the upper surface of each of objects 3510 - 3550. In some examples, the 3D image information may include a point cloud that includes a plurality of 3D coordinates that describe various positions on one or more outer surfaces of objects 3510 - 3550 or some other objects within camera field of view 3210 / 3210B.

[0047] In the embodiments of FIGS. 3A and 3B, a robot 3300 (which may be an embodiment of robot 1300) can include a robot arm 3320 having one end attached to a robot base 3310 and another end attached to or formed by an end effector device 3330 such as a robot gripper. The robot base 3310 can be used to mount the robot arm 3320, but the robot arm 3320, and more specifically, the end effector device 3330, can be used to interact with one or more objects (e.g., 3510 / 3520 / 3530 / 3540) in the environment of the robot 3300. The interaction (also referred to as robot interaction) can include, for example, grasping or otherwise picking up at least one of the objects 3510 - 3540. For example, the robot interaction can be part of an operation of picking up objects 3510 - 3540 (e.g., boxes) from an object 3550 (e.g., a pallet or other platform) by the robot 3300 and moving the objects 3510 - 3540 to a destination location.

[0048] As discussed above, one aspect of the present disclosure relates to performing or facilitating the detection of one or more physical edges of a group of objects, such as a group of boxes, based on image information representing one or more objects. FIG. 4 shows a flowchart of an exemplary method 4000 for performing or facilitating physical edge detection, or more specifically, for determining whether at least one of the physical edges of a group of objects should be represented using candidate edges. More specifically, the method may involve receiving image information that has candidate edges that can represent physical edges or that could be false edges. A false edge can be, for example, a candidate edge that represents a visible line or other visual marking displayed on one of the surfaces of a group of objects. The visual marking may have an appearance that resembles a physical edge but does not actually correspond to any physical edge. Thus, method 4000 can be used, in embodiments, to evaluate a confidence level or likelihood as to whether a candidate edge corresponds to an actual physical edge or whether the candidate edge is likely to be a false edge. If the candidate edge is likely to be a false edge and / or does not have a sufficiently high confidence level corresponding to an actual physical edge, method 4000 can, in embodiments, remove or more broadly exclude the candidate edge from further consideration for representing any physical edge of the group of objects.

[0049] In an embodiment, method 4000 may be performed, for example, by computing system 1100 of FIGS. 2A-2D, or FIGS. 3A or 3C, or more specifically, by at least one processing circuit 1110 of computing system 1100. In some scenarios, at least one processing circuit 1100 may perform method 4000 by executing instructions stored on a non-transitory computer-readable medium (e.g., 1120). For example, the instructions may cause processing circuit 1110 to execute one or more of the modules shown in FIG. 2D capable of performing method 4000. As an example, one or more of steps 4002-4008 discussed below may be performed by physical edge detection module 1125. If method 4000 includes steps for performing object recognition and / or object registration, the steps may be performed, for example, by object recognition / registration module 1128. If method 4000 involves planning robot interaction or generating robot interaction movement commands, such steps may be performed, for example, by motion planning module 1129. In an embodiment, method 4000 may be performed in an environment where computing system 1100 is communicating with robots and cameras such as robots 3300 and cameras 3200 / 3200A / 3200B of FIGS. 3A and 3C, or any other camera or robot discussed in this disclosure. In some scenarios as shown in FIGS. 3A and 3C, a camera (e.g., 3200) may be mounted on a stationary structure (e.g., a ceiling of a room). In other scenarios, the camera may be mounted on a robot arm (e.g., 3320), or more specifically, on an end effector device (e.g., 3330) of a robot (e.g., 3300).

[0050] In an embodiment, one or more steps of method 4000 may be performed when a group of objects (e.g., 3510 - 3550) is currently within the camera field of view (e.g., 3210 / 3210A / 3210B) of a camera (e.g., 3200 / 3200A / 3200B). For example, one or more steps of method 4000 may be performed immediately after a group of objects enters the camera field of view (e.g., 3210 / 3210A / 3210B), or more generally, while a group of objects is within the camera field of view. In some scenarios, one or more steps of method 4000 may be performed when a group of objects is within the camera field of view. For example, when a group of objects is within the camera field of view (e.g., 3210 / 3210A / 3210B), the camera (e.g., 3200 / 3200A / 3200B) may generate image information representing the group of objects and communicate the image information to a computing system (e.g., 1100). The computing system may perform one or more steps of method 4000 based on the image information while the group of objects is still within the camera field of view or even when the group of objects is no longer within the camera field of view.

[0051] In an embodiment, method 4000 may start from step 4002 where computing system 1100 receives image information representing a group of objects within the camera field of view (e.g., 3210 / 3210A / 3210B) of a camera (e.g., 3200 / 3200A / 3200B), or alternatively may include step 4002. The image information may be generated by a camera (e.g., 3200 / 3200A / 3200B) when the group of objects is (or was) within the camera field of view, and may include, for example, 2D image information and / or 3D image information. For example, FIG. 5A shows a 2D image 5600, or more specifically, a 2D image generated by camera 3200 / 3200A and representing the objects 3510-3550 of FIGS. 3A and 3C. More specifically, the 2D image 5600 (e.g., grayscale or color image) may depict the appearance of the objects 3510-3550 from the viewpoint of camera 3200 / 3200A. In an embodiment, the 2D image 5600 may correspond to a single color channel (e.g., red, green, or blue channel) of a color image. When camera 3200 / 3200A is disposed above the objects 3510-3550, the 2D image 5600 may represent the appearance of the respective upper surfaces of the objects 3510-3550. In the example of FIG. 5A, the 2D image 5600 may include respective portions 5610, 5620, 5630, 5640, and 5650 (also referred to as image portions), each representing a respective surface (e.g., upper surface) of the objects 3510-3550. In FIG. 5A, each of the image portions 5610-5650 of the 2D image 5600 may be an image region, or more specifically, a pixel region (if the image is formed by pixels). More specifically, the image region may be a region of the image, and the pixel region may be a region of pixels. One or more of the image portions 5610-5550 may capture or otherwise represent visual markings or other visual details that are visible or appear on the surface of the object. For example, image portion 5610 may represent the visual marking 3612 of FIG. 3B that may be printed or otherwise disposed on the upper surface of object 3610.

[0052] 5B illustrates an example in which the image information of step 4002 includes 3D image information 5700. More specifically, the 3D image information 5700 may include, for example, a depth map or point cloud indicating respective depth values ​​of various locations on one or more surfaces (e.g., top surfaces, or other exterior surfaces) of the objects 3510-3550. For example, the 3D image information 5700 may include a set of locations 57101-5710 on the surface of the object 3510. n a first portion 5710 (also called image portion) showing respective depth values ​​of the first and second positions 57201-57202 (also called physical positions) on the surface of the object 3520; n and a set of positions 57301-5730 on the surface of the object 3530. n and a set of positions 57401-5740 on the surface of the object 3540. n and a fourth portion 5740 indicating the depth values ​​of each of the positions 57501-5750 on the surface of the object 3550. n and a fifth portion 5750 indicating a depth value of each of the objects 3510-3550. The respective depth values ​​may be relative to the camera (e.g., 3200 / 3200B) generating the 3D image information, or may be relative to some other reference point. In some implementations, the 3D image information may include a point cloud including respective coordinates for various positions on the structure of the object within the camera field of view (e.g., 3210 / 3210B). In the example of FIG. 5B, the point cloud may include respective sets of coordinates describing positions on the surface of each of the objects 3510-3550. The coordinates may be 3D coordinates, such as [XYZ] coordinates, and may have values ​​relative to the camera coordinate system, or some other coordinate system. As an example, the camera coordinate system is defined by X, Y, Z, as shown in FIGS. 3A, 3C, and 5B.

[0053] In an embodiment, step 4002 may involve receiving both 2D image information and 3D image information. In some examples, the computing system 1100 may use the 2D image information to compensate for limitations in the 3D image information, and vice versa. For example, if multiple objects within the camera field of view are arranged in close proximity to each other and have substantially equal depths relative to the camera (e.g., 3200B), the 3D image information (e.g., 5700) may describe multiple positions having substantially equal depth values. In particular, when the spacing between objects is too narrow relative to the resolution of the 3D image information, the 3D image information may lack the details necessary to distinguish the individual objects represented. In some examples, the 3D image information may have errors or missing information due to noise or other sources of error, which may further increase the difficulty of distinguishing individual objects. In this example, the 2D image information may compensate for this lack of detail by capturing or otherwise representing the physical edges between individual objects. However, in some examples, the 2D image information may include spurious edges that are candidate edges that do not correspond to actual physical edges, as discussed below. In some implementations, the computing system 1100 may evaluate the likelihood that a candidate edge is a spurious edge by determining whether the 2D image information meets defined darkness conditions for candidate edges, as discussed below with respect to step 4006. In some implementations, the computing system 1100 may determine whether a candidate edge corresponds to a physical edge within the 3D image information, such as when the candidate edge corresponds to a physical location where the 3D image information describes a sharp change in depth. In such scenarios, the 3D image information may be used to check whether a candidate edge is a spurious edge, complementing or replacing the use of the defined darkness conditions to provide a more robust method of determining whether a candidate edge is a spurious edge or whether the candidate edge corresponds to an actual physical edge.

[0054] Returning to FIG. 4, in one embodiment, method 4000 may include step 4004 where computing system 1100 identifies a plurality of candidate edges associated with a group of objects (e.g., 3510 - 3550) from the image information of step 4002. In an embodiment, a candidate edge may be or may include a set of image positions or physical positions that form a candidate for representing a physical edge of an object or group of objects. In one example, where the image information includes a 2D image representing one or more objects, a candidate edge may refer to a set of pixel positions (e.g., pixel positions [u1v1] - [u k v k ). The set of pixel positions may correspond to a set of pixels that are collectively similar to a physical edge. For example, FIG. 6A shows an example where computing system 1100 identifies candidate edges 56011, 56012, 56013, 56014, 56015, 56016,... 5601 n from 2D image 5600. Each candidate edge of candidate edges 56011 - 5601 n may include or may be formed by each respective set of pixel positions that, for example, define a line or line segment where the 2D image has a sharp change in image intensity. A sharp change in image intensity may occur, for example, between two image regions where one image region is darker than the other and the two are directly adjacent to each other. As discussed in more detail below, candidate edges may be formed based on a boundary between two image regions. In such embodiments, the boundary may be formed by the lines or line segments discussed above. Candidate edges identified from a 2D image (e.g., 5600) may be referred to as 2D candidate edges or 2D edges.

[0055] In an embodiment, the image information may include several candidate edges corresponding to actual physical edges and may also include several candidate edges that are false edges. For example, the candidate edges 56011, 56012, 56015, 56016 in FIG. 6A may correspond to the actual physical edges of the object groups 3510 to 3550, or more specifically, may correspond to the physical object 3510. On the other hand, the candidate edges 56013 and 56014 may be false edges. The candidate edges 56013 and 56014 may represent, for example, visible lines or other visual markings displayed on the surface of the object 3510. These visible lines may be similar to physical edges but do not correspond to the actual physical edges of the objects 3510 to 3550. Therefore, as discussed below with respect to step 4008, the method 4000 may involve, in one embodiment, determining whether to use a specific candidate edge to represent at least one of the physical edges of the object group.

[0056] In one example, when the image information includes 3D information, the candidate edge may refer to a set of image positions or a set of physical positions. As an example, when the image positions are pixel positions, they may correspond to a set of pixels that appear like physical edges. In another example, when the 3D image information includes a depth map, the candidate edge may include a set of pixel positions that form a boundary with a sharp change in depth in the depth map, for example, defining a line or a line segment. When the 3D image information describes the 3D coordinates of physical positions on the surface of an object (e.g., via a point cloud), the candidate edge may include a set of physical positions that form a boundary with a sharp change in depth in the point cloud or other 3D image information, for example, defining a virtual line or a line segment. For example, FIG. 6B shows an example where the computing system 1100 identifies candidate edges 57011, 57012, 57013, 5701 n Each of the candidate edges 57011 to 5701 n defines, for example, a boundary where a sharp change in depth occurs, physical positions [X1Y1Z1] to [X p Y pIt may include a set of Z1]. Candidate edges identified from the 3D image information may be referred to as 3D candidate edges or 3D edges.

[0057] In an embodiment, when the computing system 1100 identifies both 2D candidate edges and 3D candidate edges, the computing system 1100 may be configured to determine whether any of the 2D candidate edges (e.g., 56015) represent a common physical edge with one of the 3D candidate edges (e.g., 57011), or vice versa. In other words, the computing system 1100 may determine whether any of the 2D candidate edges map to one of the 3D candidate edges, or vice versa. The mapping may be based on, for example, converting the coordinates of the two-dimensional candidate edge from being represented in the coordinate system of the 2D image information to being represented in the coordinate system of the 3D image information, or converting the coordinates of the three-dimensional candidate edge from being represented in the coordinate system of the 3D image information to being represented in the coordinate system of the 2D image information. The mapping from 2D candidate edges to 3D candidate edges is discussed in more detail in U.S. Patent Application No. 16 / 791,024, entitled "METHOD AND COMPUTING SYSTEM FOR PROCESSING CANDIDATE EDGES" (Attorney Docket No. MJ0049-US / 0077-0009US1), the entire content of which is incorporated herein by reference.

[0058] As described above, the computing system 1100 may identify candidate edges from a 2D image or other 2D image information by identifying image positions (e.g., pixel positions) within the 2D image information where there are abrupt changes in image intensity (e.g., pixel intensity). In some implementations, the abrupt change may occur at a boundary between two image regions where one image region is darker than another image region. For example, the two image regions may include a first image region and a second image region. The first image region may be a region of the 2D image that is darker than one or more directly adjacent regions that may include the second image region. The darkness of an image region may indicate how much reflected light was detected from the corresponding physical region by a camera (e.g., 3200 / 3200A) that generates the image information (e.g., 5600), or was otherwise sensed. More specifically, a dark image region may indicate that the camera sensed a relatively small amount of reflected light (or no reflected light) from the corresponding physical region. In some embodiments, the darkness of an image region may indicate how close the image intensity within the image region is to the minimum possible image intensity value (e.g., zero). In these implementations, a dark image region may indicate that the image intensity value(s) of the image region are close to zero, and an image region that is not darker may indicate that the image intensity value(s) of the image region are close to the maximum possible image intensity value.

[0059] In an embodiment, the second image region may have an elongated shape such as a rectangular band or a line or a line segment. As an example, FIG. 7A shows examples of image regions 56031, 56032, 56033, 56034, 56035, 56036 that are darker than the directly adjacent image regions 56051, 56052, 56053, 56054, 56055, 56056, respectively. If the image 5600 includes pixels, the image regions in FIG. 7A may be pixel regions. Each of the image regions 56031 to 56034 may be a band of pixels such as a rectangular band and may have the width of a plurality of pixels, and each of the image regions 56035 and 56036 may be a line of pixels having the width of one pixel or may be formed. As described above, the candidate edge may be formed by or based on the boundary between a first image region (for example, one of the image regions 56031 to 56036) and a second image region (for example, one of the image regions 56051 to 56056), the first image region may be directly adjacent to the second image region, may be darker than the second image region, and a sharp change in image intensity may occur at the boundary between the two image regions. For example, FIG. 7B shows candidate edges 56011 to 56016 defined or formed by the respective boundaries between the image regions 56031 to 56036 and the directly adjacent corresponding image regions 56051 to 56056. As an example, the computing system 1100 may identify the candidate edge 56011 as a set of pixel positions defining the boundary between one image region 56051 and another dark image region 56031. As an additional example, the computing system 1100 may identify the candidate edge 56015 as a set of pixel positions defining the boundary between one image region 56055 and the dark image region 56035. In some examples, the pixel positions of the candidate edge 56015 may be located in the dark image region 56035. More specifically, the candidate edge 56015 may be or may coincide with the image region 56035, which may be, for example, a line of pixels having the width of a single pixel.

[0060] In an embodiment, the computing system 1100 can detect or otherwise identify candidate edges, such as one of candidate edges 56011-56016, based on, for example, image edge detection techniques that can detect abrupt changes in image intensity. For example, the computing system 1100 can be configured to detect candidate edges or other image information in a 2D image by applying a Sobel operator, a Prewitt operator, or other techniques for determining intensity gradients in a 2D image and / or by applying a Canny edge detector or other edge detection techniques.

[0061] In an embodiment, when the computing system 1100 identifies an image region in a 2D image that is a band of pixels darker than one or more immediately adjacent image regions, the image region can be wide enough in some environments to form more than a candidate edge. For example, FIG. 7C shows the computing system 1100 identifying additional candidate edges 56017-5601 10 based on image regions 56031-56034. More specifically, the additional candidate edges 56017-5601 10can be respective sets of pixel positions that define respective boundaries between the image regions 56031 - 56034 and the immediately adjacent image regions 56071 - 56074. As a more specific example, the image region 56032 of this embodiment can be wide enough such that image edge detection techniques identify a candidate edge 56012 formed by a boundary between the region 56052 immediately adjacent to one side (e.g., the right side) of the image region 56032, as shown in FIG. 7B, and another candidate edge 56018 formed by a boundary between the region 56072 immediately adjacent to the opposite side (e.g., the left side) of the image region 56032, as shown in FIG. 7C. In an embodiment, the image region can be very narrow such that the image edge detection technique can identify only a single candidate edge from the image region. In such embodiments, the image region can have a width of a single pixel or a few pixels. For example, as described above, the image region 56035 can form a line of pixels and can have a width of one pixel. In this example, the computing system 1100 can identify only a single candidate edge 56015 based on the image region 56035, and the candidate edge 56015 can coincide with the image region 56035 such that, for example, the image candidate edge 56015 can be a line of pixels forming the image region 56035 or can overlap it.

[0062] Returning to FIG. 4, the method 4000 can include, in an embodiment, step 4006, which can be performed when a plurality of candidate edges in step 4004 include a first candidate edge formed by a boundary between a first image region and a second image region, where the first image region can be darker than the second image region in image intensity and can be immediately adjacent to the second image region. In this example, the first image region and the second image region can be regions described by a 2D image (e.g., 5600) or other image information. For example, FIGS. 7A - 7C provide examples, and the first candidate edge of step 4006 is among the plurality of candidate edges 56011 - 5601 n and the first candidate edge of step 4006 is among the plurality of candidate edges 56011 - 5601 nIt can be any one of them. As described above, the first candidate edge may be formed by or may be formed based on the boundary between the first image region and the second bright image region. As an example, when the first candidate edge is candidate edge 56011, the first image region may be image region 56031, and the second image region may be 56051. As another example, when the first candidate edge is candidate edge 56012, the first image region may be image region 56032, and the second image region may be image region 56052.

[0063] In step 4006, the computing system 1100 may determine whether the image information (e.g., 2D image 5600) satisfies the darkness condition defined by the first candidate edge (e.g., 56012). Such a determination may more specifically include, for example, determining whether the first image region (e.g., 56032) satisfies the defined darkness condition. In an embodiment, the defined darkness condition can be used to determine whether the first candidate edge (e.g., 56012) corresponds to the actual physical edge of an object (e.g., 3510) in the camera field of view (e.g., 3210 / 3210A), or whether the first candidate edge is a false edge.

[0064] In an embodiment, the defined darkness condition can be used to detect an image prior, or more specifically, a dark prior. An image prior can refer to an image feature that has a likelihood of appearing in an image or can be expected within an image in a particular situation. More specifically, an image prior can correspond to an expectation, anticipation, or prediction of what image feature(s) there might be in an image generated in such a situation, such as a situation where the image is generated to represent a group of boxes or other objects placed adjacent to each other in the camera's field of view. In some instances, a dark prior can refer to an image feature having a high level of darkness and / or having a spike-shaped image intensity profile (e.g., pixel intensity profile). A spike-shaped image intensity profile can be accompanied by a spike increase in darkness and / or a spike decrease in image intensity. A dark prior can correspond to a situation where a group of boxes or other objects in the camera's field of view are placed sufficiently close to each other such that only a narrow, physical gap exists between some or all of the objects. More specifically, a dark prior can correspond to an expectation, anticipation, or prediction that when an image is generated in such a situation to represent a group of objects, the physical gap will appear very dark within the image. More specifically, a dark prior may also correspond to an expectation or prediction that the image region in the image representing the physical gap has a high level of darkness and / or may have a spike-shaped image intensity profile, as discussed in more detail below. In some implementations, the dark prior can be used to determine whether a candidate edge corresponds to a physical edge by evaluating whether the image region associated with the candidate edge corresponds to a physical gap between two objects.

[0065] In an embodiment, in some scenarios, defined darkness conditions, which can be conditions for detecting a dark prior, may be based on a model of how the physical gap between two objects (e.g., 3510 and 3520 in FIGS. 3A - 3C) is configured. Particularly when the physical gap is narrow (e.g., less than 5 mm or less than 10 mm), it needs to be displayed, or is likely to be displayed, in a 2D image. For example, the defined darkness conditions may be based on the Lambert model of diffuse reflectance. Such a model of reflectance can estimate how light is reflected from one or more surfaces or regions, particularly surfaces or regions that cause diffuse reflection of incident light. Thus, the model may estimate the intensity of the reflected light from the surface or surface region, and may indicate how bright or how dark the surface or surface region is in an image generated by a camera (e.g., 3200 / 3200A) that senses the reflected light.

[0066] As an example of how the Lambert model is applied to a group of objects (e.g., 3510 - 3550 in FIGS. 3A - 3C), FIG. 8 depicts a camera 3200A configured to generate an image (e.g., 5600) representing at least objects 3510 and 3520 by sensing the reflected light from various surfaces of objects 3510, 3520. In some scenarios, the reflected light can be the reflection of the emitted light from a light source 3600. More specifically, the light source 3600 can emit light along at least a vector towards the objects 3510, 3520.

Number

Number

Number

Number

Number

[0067] In some situations, the physical gap may appear darker in the middle than at its periphery. That is, when some reflected light leaves the physical gap, more reflected light may come from the outer periphery of the physical gap than from the middle of the physical gap. The peripheral portion may refer to a position within the physical gap that is close to, for example, the physical edge portion 3520D or the physical edge portion 3510B. In some scenarios, a peak level of darkness may occur in the middle of the physical gap. Accordingly, the image region representing the physical gap may have a spike-shaped image intensity profile (e.g., a pixel intensity profile) in which the image intensity profile has a spike increase in darkness or a spike decrease in image intensity within the image region. Accordingly, the darkness condition defined in step 4006 may, in some situations, include a defined spike intensity profile criterion to evaluate whether the image region has, for example, a spike-shaped image intensity profile (as opposed to, e.g., a step-shaped image intensity profile).

[0068] In embodiments, the defined darkness condition may be defined by, for example, one or more rules, criteria, or other information stored in the non-transitory computer-readable medium 1120 or elsewhere. For example, the information may define whether the darkness condition is met only by meeting a darkness threshold criterion, only by meeting a spike intensity profile criterion, by meeting both criteria, or by meeting either the darkness threshold criterion or the spike intensity profile criterion. In some examples, the information may be predefined manually or otherwise so that the defined darkness condition may be a predefined darkness condition and may be stored in the non-transitory computer-readable medium 1120. Depending on the example, the information about the darkness condition may be defined dynamically.

[0069] In an embodiment, the defined darkness threshold criterion and / or the defined spike intensity profile criterion may be defined by information stored in the non-transitory computer-readable medium 1120 or elsewhere. The information may be predefined such that the defined darkness threshold criterion and / or the defined spike intensity profile criterion can be a predetermined criterion (s). In an embodiment, various predetermined thresholds or other predetermined values of the present disclosure may be manually defined as stored values in the non-transitory computer-readable medium 1120 or elsewhere. For example, the defined darkness threshold or the defined depth difference threshold discussed below may be a value stored on the computer-readable medium 1120. These may be predetermined values or may be defined dynamically.

[0070] Figures 9A-9C show embodiments for evaluating whether image information (e.g., 5600) meets the darkness conditions defined at candidate edge 56012, or more specifically, whether image region 56032 meets the darkness conditions defined. Image region 56032 may represent a physical gap between a first object such as object 3510 in FIG. 8 and a second object such as object 3520. In an embodiment, candidate edge 56012 may represent the physical edge 5610B of object 3510 and may be formed by or based on the boundary between image region 56032 and an image region 56052 that is directly adjacent to it. In this example, image region 56032 may be a first image region and image region 56052 may be a second image region. More specifically, image region 56032 may be a first pixel region that forms a band of pixels, while image region 56052 may be a second pixel region, and thus candidate edge 56012 may include or be formed by, for example, a set of pixel positions that define the boundary between the first pixel region and the second pixel region. As discussed above, computing system 1100 may identify candidate edge 56012 by detecting, for example, a sharp change in image intensity (e.g., pixel intensity) between image regions 56032, 56052.

[0071] In an embodiment, a candidate edge is formed based on a boundary between a first image region and a second image region, and when the first image region is darker than the second image region, the computing system 1100 may determine that the defined darkness condition is satisfied if the first image region meets a defined spike intensity profile criterion. More specifically, the computing system may determine that the first image region (e.g., 56032) meets the defined spike intensity profile criterion if the first image region has a particular shape, such as a shape, for its image intensity profile (e.g., pixel intensity profile) where the image intensity increases towards a peak level of darkness at a location within the first image region in the darkness within the first image region and then decreases within the darkness. Such a criterion may be consistent with a spike-shaped intensity profile where the image intensity profile has a spike increase in darkness within the image region or a spike decrease in intensity within the image region. Such a criterion may be associated with detecting a dark prior where any physical gap appearing in the image is expected to appear darker in the middle of the gap than on the outer periphery of the gap.

[0072] FIG. 9B shows an image intensity profile 9001, or more specifically, a pixel intensity profile, that can meet the defined spike intensity profile criteria. More specifically, the image intensity profile can include information describing how the image intensity, or more specifically, the pixel intensity, changes as a function of an image position such as a pixel position. In some embodiments, the image intensity profile can be represented by a curve or graph that describes the values of the image intensity as a function of positions within the image. For example, FIG. 9B depicts, as the image intensity profile 9001, a curve or graph that describes the values of the image intensity, or more specifically, the pixel intensity values, as a function of the pixel positions in a particular direction along axis 5609. Axis 5609 can be an axis that extends along and across the width dimension of the image region 56032, and the direction along axis 5609 can be a particular direction along axis 5609. In the example of FIG. 9B, the width dimension can be aligned, for example, with the coordinate axis u of the image 5600 of FIG. 5A, and the direction along axis 5609 can be the positive direction in which the pixel coordinates [u, v] along that direction have increasing values of u.

[0073] In some embodiments, the computing system 1100 can determine whether the image region 56032 meets the defined spike intensity profile criteria by determining whether the pixel intensity profile (e.g., 9001) has a first profile portion (e.g., 9011) in which the image intensity (e.g., pixel intensity) increases in darkness within a first image region as a function of positions along a first direction (e.g., the positive direction along axis 5609) and reaches a peak level of darkness (e.g., 9002) at a position u1 within the first image region, and then (ii) a second profile portion (e.g., 9012) in which the image intensity decreases in darkness within the first image region away from the peak level of darkness as a function of positions along the same direction (e.g., the positive direction). The image intensity profile 9001 of FIG. 9B can more specifically be a spike-shaped intensity profile having a spike decrease in the image intensity within the image region 95032.

[0074] In some implementations, an image intensity profile in which the darkness increases may correspond to an image intensity profile having values where the image intensity is decreasing. For example, an image (e.g., 5600) may have pixel intensity values within a range from a minimum possible pixel intensity value (e.g., zero) to a maximum possible pixel intensity value (e.g., 255 for pixel intensity values coded in 8 bits). In this example, when the pixel intensity value is low, the brightness level is low, so the darkness level may be high. On the other hand, when the pixel intensity value is high, the brightness level is high, so the darkness level may be low. Further in this example, the peak darkness level (e.g., 9002) of the image intensity profile may correspond to the minimum image intensity value of the image intensity profile (e.g., 9001).

[0075] In the above example, the computing system 1100 may determine whether an image region meets a defined spike intensity profile criterion by determining whether the image intensity profile has a shape such that the image intensity value (e.g., pixel intensity value) starts to decrease in image intensity towards the minimum image intensity value and then switches to increasing in image intensity away from the minimum image intensity value. For example, the image intensity profile 9001 of FIG. 9B may describe the respective pixel intensity values for a series of pixels extending across the width dimension of the image region 56032. The computing system 1100 may determine whether the image region 56032 meets the spike intensity profile criterion by determining whether the image intensity profile has a shape such that each pixel intensity value decreases towards the minimum pixel intensity value within the image region 56032 and then switches to increasing within the image region 56032 away from the minimum pixel intensity value. In this example, the minimum pixel intensity value may correspond to the peak darkness level 9002 in the image intensity profile 9001.

[0076] In an embodiment, a candidate edge is formed based on a boundary between a first image region and a second image region. When the first image region is darker than the second image region, meeting a defined darkness threshold criterion may involve a comparison with a defined darkness threshold. Such a criterion may correspond to detecting a dark prior where any physical gap present in the image is expected to be very dark in appearance. FIG. 9C shows another image intensity profile 9003 for image region 56032 and the immediately adjacent image regions. In this example, the first image region may be the image region 56032 that is immediately adjacent to a second, brighter image region (e.g., 56052), while the dark image region 56032 can be the first image region. As described above, the image region 56032 may form a band of pixels. In FIG. 9C, the computing system 1100 may determine whether the image region 56032 meets the defined darkness threshold criterion by determining whether the image region 56032 has at least one portion that is darker in image intensity than a defined darkness threshold τ dark_prior In some cases, a higher level of darkness may correspond to a lower image intensity value. In such instances, the computing system 1100 may determine whether the image region 56032 has an image intensity profile with an image intensity value (e.g., pixel intensity value) that is less than τ dark_prior of the defined darkness threshold. In some situations, the computing system 1100 may more specifically determine whether the minimum intensity value of the image intensity profile 9003 is at or below the defined darkness threshold τ dark_prior where the minimum intensity value may correspond to the peak level 9004 of darkness of the intensity profile 9003. In an embodiment, when the image region 56032 has the image intensity profile 9003, it can meet both the defined darkness threshold criterion and the defined spike intensity profile criterion.

[0077] In an embodiment, the computing system 1100 may determine that a defined darkness condition is satisfied for a candidate edge and / or an image region if at least one of a defined darkness threshold criterion or a defined spike intensity profile criterion is satisfied such that any one of the above criteria can be used to satisfy the defined darkness condition. In an embodiment, the computing system 1100 may determine that the defined darkness condition is satisfied only in response to a determination that the defined spike intensity profile criterion is satisfied (regardless of whether the defined darkness threshold criterion is satisfied), only in response to a determination that the defined darkness threshold criterion is satisfied (regardless of whether the defined spike intensity profile criterion is satisfied), or only in response to a determination that both the defined darkness threshold criterion and the defined spike intensity profile criterion are satisfied.

[0078] In an embodiment, the computing system 1100 may identify a candidate edge (e.g., 56012) based on 2D image information such as a 2D image 5600 and determine whether the candidate edge satisfies a defined darkness condition. As described above, if the computing system 1100 receives both 2D image information and 3D image information, the computing system 1100 may use the 2D image information to compensate for limitations in the 3D image information or to compensate for a lack of 3D image information, and vice versa. For example, when a camera (e.g., 3200B) generates 3D image information to represent a group of objects, the 3D image information may lack information for distinguishing individual objects within the group, particularly if the group of objects has equal depth values relative to the camera. More specifically, the 3D image information may lack information for detecting narrow physical gaps between objects and may therefore have limited usefulness in identifying physical edges associated with the physical gaps.

[0079] As an example, FIG. 9D shows depth values associated with a portion 5715 of the 3D image information 5700 of FIG. 5B. More specifically, portion 5715 is a physical location 5720 a ~5720 a+5, and the position 5710 on the upper surface of the object 3510 b ~5710 b+4 Each depth value for can be described. These physical positions may be mapped to or otherwise correspond to image positions within or around the image region 56032, or image positions around the candidate edge 56012. As shown in FIG. 8, the image region 56032 may represent the physical gap g between the objects 3510, 3520. As described above, the computing system 1100 may attempt to detect one or more positions where there is a sharp change in depth using the 3D image information. However, the physical gap in FIG. 8 may be too narrow or otherwise too small relative to the resolution of the 3D image information captured by the 3D image information 5700. Thus, in the embodiment of FIG. 9D, the computing system 1100 determines that there is no sharp change in depth at the positions 5720 a ~5720 a+5 and 5710 b ~5710 b+4 and thus may determine that the 3D image information does not show any candidate edges at those positions. Further, in some situations, the 3D image information is at the positions 5720 a ~5720 a+5 and 5710 b ~5710 b+4Depth information may be missing for a portion thereof, or more specifically, for one or more positions corresponding to candidate edge 56011. In some situations, a portion of the 3D image information that maps or otherwise corresponds to candidate edge 56011 may be affected by an imaging noise level greater than a defined noise tolerance threshold that may be a value defined in non-transitory computer-readable medium 1120. In the above embodiments, the 2D image information may include candidate edge 56011 that represents physical edge 3510B of object 3510, and may also include image region 56032 that does not represent a physical gap between object 3510 and object 3520, such that the 2D image information may compensate for these limitations of the 3D image information. In the above embodiments that include limitations of the 3D image information, computing system 1100 may further use defined darkness conditions to determine whether to use candidate edge 56012 to represent one of the physical edges (e.g., 3510B) of a group of objects.

[0080] Figures 10A-10C show an example for determining whether candidate edge 56014 and / or image region 56034 meet defined darkness conditions. Image region 56034 may represent visual marking 3512 of FIG. 3B, such as a visible line printed on the upper surface of object 3510. In this example, image region 56034 may be darker than immediately adjacent image regions, such as image regions 56054 and 56074. Candidate edge 56014 may be formed based on the boundary between dark image region 56034 and immediately adjacent region 56054.

[0081] FIG. 10B shows an image region 56034 having an image intensity profile 10001 that does not meet the defined spike intensity profile criteria because the image intensity profile 10001 varies within the image region 56034 in such a way that the darkness increases as a function of position towards the peak level of darkness and then decreases away from the peak level of darkness. The image intensity profile 10001 of this embodiment can describe pixel intensity values as a function of pixel position along axis 5608, which can be aligned with the width dimension of the image region 56034. As described above, the spike intensity profile criteria can accommodate the expectation that the physical gap between the physical edges of an object may appear darker in the middle of the physical gap than on the outer periphery of the physical gap. Thus, an image region representing a physical gap varies as a function of image position within the image region, and more specifically, the darkness increases as a function of image position along a particular direction from the image positions corresponding to the periphery of the physical gap to the image positions corresponding to the center of the physical gap, and then the darkness decreases as a function of position along the same direction. More specifically, the image intensity profile can have a spike increase in darkness or a spike decrease in image intensity in the image region representing the physical gap. In an embodiment, instead, an image region representing a visual line or other visual marking may lack such an image intensity profile and instead may have a more uniform level of darkness within the image region. Thus, as shown in FIG. 10B, the image region 56034 representing a part of the visual marking 3512 can have an image intensity profile 10001 that is more uniform within the image region 56034 such that the profile 10001 does not substantially vary within the image region 56034. Further, the image intensity profile 10001 can have a stepped change in image intensity from the image intensity associated with a brighter adjacent image region (e.g., 56054) at the boundary of the image region 56034 to the uniform image intensity within the image region 56034.Therefore, the image intensity profile 10001 does not have a shape that begins by increasing the darkness along a specific direction across the image region 56034 towards the peak level of darkness and then switches to decreasing the darkness along that direction. More specifically, the image intensity profile 10001 does not exhibit a spike decrease in image intensity. Thus, the computing system 1100 of this embodiment can determine that the image region 56034, and thus the 2D image 5600, does not meet the defined spike intensity profile criteria, which could lead to a determination that the image region 56034 does not meet the darkness conditions defined by the candidate edge 56014 and / or the image region 56034.

[0082] FIG. 10C depicts an image region 56034 having an image intensity profile 10003, which may not meet the defined darkness threshold criteria because the image intensity profile 10003 may indicate that the image region 56034 is not dark enough. More specifically, the computing system 1100 may determine that most or all of the pixel intensity values within the image intensity profile 10003 of the image region 56034 exceed τ dark_prior of the defined darkness threshold. Thus, the computing system 1100 of FIG. 10C can determine that the image region 56034 does not meet the defined darkness threshold criteria, which could lead to a determination that the image region 56034, and thus the image 5600, does not meet the darkness conditions defined by the candidate edge 56014 and / or the image region 56034. The image intensity profile 10003 may also not meet the defined spike intensity profile criteria, as discussed above for FIG. 10B.

[0083] In an embodiment, the image region may have a width that is too small to make a reliable assessment as to whether the image region meets the defined spike intensity profile criteria. For example, the image region may have a width of only a single pixel or only a few pixels. In some examples, the computing system 1100 may determine that such an image region does not meet the defined darkness conditions. In other examples, the computing system 1100 may determine whether the defined darkness conditions for the image region are met based on whether the image region meets the defined darkness threshold criteria. In some examples, the computing system 1100 may decide not to evaluate such an image region or associated candidate edges with respect to the defined darkness conditions.

[0084] As described above, one aspect of the present disclosure relates to a situation where the computing system 1100 identifies a plurality of candidate edges based on at least 2D image information, such as the 2D image 5600. In such embodiments, the plurality of candidate edges may include at least a first candidate edge (e.g., 56011 / 56012 / 56013 / 56014) identified based on the 2D image. For example, the first candidate edge may be formed based on the boundary between two image regions of the 2D image. In some examples, the computing system 1100 may identify a plurality of candidate edges based on 2D image information and 3D image information. In such examples, the plurality of candidate edges may, as described above, include the first candidate edge from the 2D image information and may further include a second candidate edge (e.g., 57011 in FIG. 6B) identified based on the 3D image information.

[0085] As an example, FIGS. 11A-11B show a computing system 1100 that identifies a candidate edge 57011 as a second candidate edge among a plurality of candidate edges based on 3D image information 5700. In this example, the computing system 1100 may identify the candidate edge 57011 based on detecting a sharp change in depth at the candidate edge 57011 between a first portion 5710A and a second portion 5750A of the 3D image information 5700. The first portion 5710A may represent, for example, a region of positions on the upper surface of the object 3510 of FIGS. 3A-3C, while the second portion 5750A may represent, for example, a region of positions on the upper surface of the object 3550. More specifically, as shown in FIG. 11B, the first portion 5710A of the 3D image information may include respective depth values for positions 5710 c ~5710 c+5 on the upper surface of the object 3510, while the second portion 5750A may include respective depth values for positions 5750 d ~5750 d+4 on the upper surface of the object 3550. In this example, the positions 5710 c ~5710 c+5 and 5750 d ~5750 d+4 may be a series of positions aligned along the Y-axis of FIG. 11A.

[0086] In an embodiment, the computing system 1100 may identify the candidate edge 57011 of FIG. 11B based on detecting a sharp change in depth between two consecutive or otherwise adjacent positions of a series of positions described by the 3D image information 5700. Such a sharp change may be referred to as a depth discontinuity. The sharp change may be detected, for example, when the difference between the respective depth values of two positions exceeds a defined depth difference threshold. For example, the computing system 1100 may determine that the difference between the depth value of position 5710 c+5 and the depth value of position 5750 d exceeds a defined depth difference threshold. As a result, the computing system 1100 may identify these two positions 5710 c+5 、5750 dBased on this, candidate edge 57011 can be identified. For example, candidate edge 57011 can be identified to include positions between positions 5710 c+5 , 5750 d on the Y-axis.

[0087] In an embodiment, the computing system 1100 can identify candidate edges by identifying two surfaces having a depth difference exceeding a defined depth difference threshold based on 3D image information. For example, as shown in FIG. 11C, the computing system 1100 may identify a first surface of object groups 3510-3550 based on a first set of positions described by 3D image information 5700, and the first set of positions has respective depth values that do not deviate from each other beyond a defined measurement variance threshold. Similarly, the computing system 1100 may identify a second surface of object groups 3510-3550 based on a second set of positions described by 3D image information 5700, and the second set of positions has respective depth values that do not deviate from each other beyond a defined measurement variance threshold. In the example of FIG. 11C, the first set of positions may include positions 5710 c ~5710 c+5 that may represent the upper surface of object 5710, while the second set of positions may include positions 5750 d ~5750 d+4 that may represent the upper surface of object 5750.

[0088] In this embodiment, the defined measurement variance threshold can describe the effects of imaging noise, manufacturing tolerances, or other factors that can introduce random variations in the depth measurements taken by a camera (e.g., 3200B). Such sources of random variation introduce some natural variance in the depth values of different positions, even if those different positions are part of a common surface and actually have the same depth relative to the camera. In some examples, the defined measurement variance threshold may be equal to or based on a nominal standard deviation that is used to describe the expected random variation in the depth measurements, or more generally, how sensitive the camera is to noise or other sources of error. The nominal standard deviation can describe a baseline standard deviation or other form of variance expected in the depth values or other depth information generated by the camera. The nominal standard deviation, or more generally, the defined measurement variance threshold, may be a value stored, for example, in the non - transitory computer - readable medium 1120, and can be a predetermined value or a dynamically defined value. In an embodiment, if a set of positions have respective depth values that do not deviate from each other beyond the defined measurement variance threshold, the computing system 1100 may determine that the set of positions is part of a common surface. In a more specific embodiment, the computing system 1100 may determine that the set of positions is part of a common surface if the standard deviation (e.g., Std 5710 or Std 5750 ) of their respective depth values is less than the defined measurement variance threshold.

[0089] In the above embodiment, the computing system 1100 can identify candidate edges from the 3D image information based on two surfaces having a sufficient depth difference. For example, the first set of positions in FIG. 11C that describe the upper surface of the object 5710 can have an average depth value Avg 5710 or be otherwise associated with it. Similarly, the second set of positions that describe the upper surface of the object 5750 can have an average depth value Avg 5750 or be otherwise associated with it. The computing system 1100 can compare Avg 5710 with Avg 5750It can be determined whether the difference from is greater than or equal to a defined depth difference threshold. In some examples, the defined depth difference threshold can be determined as a multiple of the defined measurement variance threshold (e.g., 2 times the defined measurement variance threshold, or 5 times the defined measurement variance threshold). Avg associated with two surfaces 5710 and Avg 5750 If the difference between and is greater than or equal to the defined depth difference threshold, the computing system 1100 may determine that the depth discontinuity condition is satisfied. More specifically, the computing system 1100 may determine that a candidate edge (e.g., 57011) exists at a position between two surfaces, and more specifically, may identify the candidate edge based on a position where there is a transition between two surfaces.

[0090] As described above, one aspect of the present disclosure relates to compensating 2D image information and 3D image information for each other such that the 3D image information can compensate for limitations of the 2D image information (and vice versa). In some examples, physical edges detected from the 3D image information may be associated with a higher level of confidence than physical edges detected only from the 2D image information. In some cases, when a physical edge (e.g., 3510A in FIG. 3B) is represented in both the 2D image information and the 3D image information, the computing system 1100 can identify a candidate edge (e.g., 56015) representing the physical edge in the 2D image information and a corresponding candidate edge (e.g., 57011) representing the physical edge in the 3D image information. As described above, the corresponding candidate edges can map to each other. For example, a candidate edge (e.g., 57011) in the 3D image information can be mapped to a candidate edge (e.g., 56015) in the 2D image information. A candidate edge (e.g., 56015) can be formed based on a boundary between two image regions such as 56055 and 5650. However, in some situations, the computing system 1100 may not be able to determine with a high level of confidence whether a candidate edge (e.g., 56015) from the 2D image information corresponds to an actual physical edge. For example, in FIG. 11D, the 2D image 5600 may have a step-shaped change in image intensity at the candidate edge 56015. In this example, the computing system 1100 may determine that the 2D image 5600 does not meet the defined darkness conditions at the candidate edge 56015, or more specifically, in the two image regions 56055 and 5650. Thus, the computing system 1100 may determine that there is not a sufficiently high level of confidence associated with the candidate edge 56015 representing the physical edge.

[0091] In such a situation, the computing system 1100 may provide additional input using the 3D image information. More specifically, as described above with respect to FIGS. 11A-11C, the computing system 1100 may identify candidate edge 57011 based on the 3D image information, and candidate edge 56015 in the 2D image information may be mapped to candidate edge 57011 in the 3D image information, or otherwise correspond thereto. In some examples, since candidate edge 57011 is identified based on depth information, the computing system 1100 may determine that there is a reasonably high likelihood that a physical edge, i.e., candidate edge 57011 representing physical edge 3510A in FIG. 3B, exists. Thus, while the 2D image information may not lead to the detection of physical edge 3510A, or may lead to the detection of physical edge 3510A with a low confidence level, the 3D image information may be used by the computing system 1100 to detect physical edge 3510A with a higher confidence level.

[0092] Returning to FIG. 4, in one embodiment, method 4000 includes the computing system 1100 identifying a plurality of candidate edges (e.g., a plurality of candidate edges 56011-5601 nSelect a subset of the subset (e.g., of the subset of 3510-3540) to form a selected subset of candidate edges for representing the physical edges of the group of objects (e.g., 3510-3540), which may include step 4008. In an embodiment, this step may involve excluding from the subset one or more candidate edges that are each likely to be a false edge. One or more candidate edges that are likely to be false edges may be removed from the subset of candidate edges, or more generally, may be ignored from further consideration for representing the physical edges of the group of objects (e.g., 3510-3540). In one example, the computing system 1100 may select a subset of the plurality of candidate edges by determining the candidate edge(s) to remove from the plurality of candidate edges, and after the removal, the plurality of candidate edges form the resulting subset. In one example, when the plurality of candidate edges are represented or described by the edge detection information 1126 of FIG. 2C, removing a candidate edge may involve deleting the information regarding that candidate edge from the edge detection information 1126.

[0093] As described above, a plurality of candidate edges (e.g., 56011-5601 n or 56011-5601 n and 57011-5701 nIt may include at least a first candidate edge (e.g., 56011 or 56014) formed based on the boundary between the first image region and the second image region darker than the first image region. Further, the first candidate edge can be identified from the 2D image information. In an embodiment, step 4008 may involve determining whether to include the first candidate edge in a subset (also referred to as a subset of candidate edges). By including the first candidate edge (e.g., 56011) in the subset, it may be possible to use the first candidate edge to represent at least one physical edge (e.g., 3510B) of a group of objects within the camera's field of view. More specifically, when the first candidate edge (e.g., 56011) is included in the subset, such inclusion may be an indication that the first candidate edge (e.g., 56011) is a candidate that remains under consideration for representing at least one of the physical edges of the group of objects. In other words, the computing system 1100 may determine whether to retain the first candidate edge as a candidate for representing at least one physical edge. If the computing system 1100 determines to retain the first candidate edge as such a candidate, it may include the first candidate edge within a subset (which can also be referred to as a selected subset of candidate edges). This determination may be part of the step of selecting a subset of multiple candidate edges and may be made based on whether the image meets the darkness conditions defined by the first candidate edge. In some examples, the inclusion of the first candidate edge within the subset may be an indication that the likelihood of the first candidate edge being a false edge is sufficiently low. In some cases, including the first candidate edge in the subset means that the computing system 1100 uses the first candidate edge to represent at least one physical edge of the group of objects or continues to consider at least the first candidate edge to represent at least one physical edge of the group of objects, indicating that the first candidate edge (e.g., 56011) has a sufficiently high confidence level corresponding to the actual physical edge of the group of objects.If the computing system 1100 determines not to include a first candidate edge (e.g., 56014) in a subset such that the first candidate edge (e.g., 56014) is filtered or otherwise excluded from the subset, such exclusion may be an indication that the first candidate edge (e.g., 56014) is no longer a candidate for representing at least one of the physical edges of the group of objects. In some instances, the exclusion of the first candidate edge from the subset may be an indication that the first candidate edge (e.g., 56014) is likely a false edge.

[0094] In an embodiment, the determination of whether the selected subset of candidate edges includes the first candidate edge may be based on whether the image information (e.g., 5600) satisfies the darkness condition defined by the first candidate edge, as described above. In some implementations, if the image information satisfies the darkness condition defined by the first candidate edge, such a result may indicate that the likelihood that the first candidate edge is a false edge is sufficiently low. This is because the first candidate edge in such a situation is likely to be associated with an image region representing a physical gap between two objects. Therefore, the first candidate edge may represent a physical edge forming one side of the physical gap. In such a situation, the computing system 1100 may determine to include the first candidate edge in the selected subset. In some examples, if the image information does not satisfy the darkness condition defined by the first candidate edge, the computing system 1100 may determine not to include the first candidate edge in the selected subset. In some examples, if the computing system determines that the 2D image information does not satisfy the darkness condition defined by the first candidate edge, the computing system 1100 may further evaluate the first candidate edge using the 3D image information. For example, if the computing system 1100 determines that the 2D image 5600 does not satisfy the darkness condition defined by the candidate edge 56015, the computing system 1100 can determine whether the candidate edge 56015 is mapped to the candidate edge 57011 described by the 3D image information, and whether the candidate edge 57011 of the 3D image information exhibits a depth change greater than the defined depth difference threshold, as discussed above with respect to FIGS. 11A - 11D.

[0095] In an embodiment, the method 4000 may execute steps 4006 and / or 4008 multiple times (e.g., through multiple iterations) to determine whether the image information satisfies the darkness conditions defined by multiple candidate edges, and select the subset discussed above based on these determinations. As an example, the multiple candidate edges may include at least candidate edges 56011 - 5601 nWhen including, the computing system 1100 executes step 4006 multiple times so that the 2D image 5600, for example, the candidate edges 56011 to 5601 n can determine whether it satisfies the defined darkness condition. The computing system 1100 further executes step 4008 multiple times to determine which of these candidate edges are included in the subset and remain candidates for representing the physical edge, and which of these candidate edges are excluded from the subset and thus cease to be candidates for representing the physical edge. For example, the computing system 1100 can determine that the subset includes candidate edges 56011 and 56012 because the 2D image 5600 satisfies the darkness condition defined by those candidate edges, and that the subset does not include candidate edges 56013 and 56014 because the 2D image does not satisfy the darkness condition defined by those candidate edges. In some situations, the computing system 1100 can determine not to include candidate edge 56015 in the subset because the 2D image 5600 cannot satisfy the darkness condition defined by the candidate edges. In some situations, the computing system 1100 can determine to still include candidate edge 56015 in the subset if candidate edge 56015 is mapped to candidate edge 57011 of the 3D image information indicating a depth change exceeding the depth difference threshold.

[0096] In an embodiment, the method 4000 can include a step where the computing system 1100 outputs a robot interaction movement command. The robot interaction movement command can be used for robot interaction between a robot (e.g., 3300) and at least one object of a group of objects (e.g., 3510 to 3550). The robot interaction can involve, for example, the robot (e.g., 3300) picking up an object (e.g., a box) from a pallet, moving the object to a destination position, or performing other operations such as lowering it from the pallet.

[0097] In an embodiment, the robot interaction movement command may be generated based on a selected subset of the candidate edges of step 4008. For example, the computing system 1100 may use the selected subset of candidate edges to distinguish individual objects from a group of objects described by the image information. In some examples, the computing system 1100 may perform segmentation of the image information using the selected subset. For example, when the image information includes a point cloud, the computing system may perform point cloud segmentation, which may involve using the selected subset of candidate edges to identify a portion of the point cloud corresponding to an individual object within a group of objects. The point cloud segmentation is described in U.S. Patent Application No. 16 / 791,024 (Attorney Docket No. MJ0049-US / 0077-0009US1), which is hereby incorporated by reference in its entirety. In one example, when the image information includes 2D image information, the computing system 1100 may use the selected subset of candidate edges to separate a portion of the 2D image information corresponding to an individual object from a group of objects. The separated portion may be used as a target image or a target image portion for performing, for example, an object recognition operation or an object registration operation (e.g., by module 1128). Object registration and object recognition are discussed in more detail in U.S. Patent Application No. 16 / 991,466 (Attorney Docket No. MJ0054-US / 0077-0012US1) and U.S. Patent Application No. 17 / 193,253 (Attorney Docket No. MJ0060-US / 0077-0017US1), the entire contents of which are hereby incorporated by reference. In such examples, the robot interaction movement command may be generated based on the results of an object recognition operation or an object registration operation. For example, the object recognition operation may generate a detection hypothesis that is an estimate of which object or object type is represented by the image information, or a portion thereof. In some examples, the detection hypothesis may be associated with an object recognition template that may include information describing, for example, the physical structure of one of objects 3510 - 3540.This information can be used by the computing system 1100 to plan the movement of a robot (e.g., 3300) to retrieve and move an object (e.g., via module 1129).

[0098] While the above steps of method 4000 are shown with respect to objects 3510 - 3550 of FIGS. 3A - 3C, FIG. 12A shows the above steps with respect to object 12510, while FIG. 13A shows the above steps with respect to objects 13510 - 13520. In an embodiment, object 12510 of FIG. 12A can be a box having an upper surface with a first physical region 12512 that is darker than a second, directly adjacent physical region 12514. For example, the first physical region 12512 can have more ink printed thereon as compared to the second physical region 12514. Object 12510 can be placed on an object 12520, which can be a pallet on which the box is placed. FIG. 12B shows a 2D image 12600 that can be generated to represent object 12510. More specifically, 2D image 12600 can include a first image region 12603 representing the first physical region 12512 and can include a second image region 12605 representing the second physical region 12514. The computing system 1100 of this embodiment can identify a first candidate edge 126011 based on the boundary between the first image region 12603 and the second image region 12605.

[0099] In an embodiment, the computing system 1100 may determine that the 2D image 12600 does not meet the darkness condition defined by the first candidate edge 126011. For example, the computing system 1100 may determine that the 2D image 12600 has an image intensity profile 12001 with a change in the step shape of the image intensity at the first candidate edge 126011. The image intensity profile may be measured along an axis 12609 extending along the u-axis of the image. In some implementations, the computing system 1100 may determine that the image intensity profile 12001, or more specifically, the image regions 12603 and 12605, do not meet the spike intensity profile criteria. The computing system 1100 may further determine that the defined darkness condition is not met at the first candidate edge 126011. As a result, the computing system 1100 may remove the first candidate edge 126011 from the edge detection information 1126.

[0100] In the embodiment of FIG. 13A, the objects 13510 and 13520 may each be boxes and may be placed on an object 13530 that can be a pallet or other platform. In this embodiment, the object 13510 may be darker than the object 13520 (e.g., as a result of being made of darker cardboard or other material). Further, the two objects may be separated by a narrow physical gap g. FIG. 13B shows a 2D image 13600 including a first image region 13603 representing the first object 13510 and a second image region 13605 representing the second object 13530. The computing system 1100 of this embodiment may identify a candidate edge 136011 based on the boundary between the two image regions 13605, 13605.

[0101] As shown in FIGS. 12B and 13B, images 12600 and 13600 may have a similar appearance. However, as shown in FIG. 13C, image 13600 may have an image intensity profile that includes a spike decrease in image intensity. More specifically, image region 13603 of image 13600 may, more specifically, include an image region 136031 for representing a physical gap g between objects 13510, 13520, and may include an image region 136032 for representing object 13520. In this embodiment, image region 136031 may include a spike decrease in image intensity and may have a minimum pixel intensity value that is less than a defined darkness threshold. Accordingly, computing system 1100 may determine that image region 136031 meets a defined spike intensity profile criterion and / or a defined darkness threshold criterion. As a result, computing system 1100 may determine that image 13600 meets a darkness condition defined by first candidate edge 136031. Accordingly, computing system 1100 may use first candidate edge 136031 to determine that it represents one of the physical edges of objects 13510, 13520.

[0102] Additional Considerations Regarding Various Embodiments:

[0103] Embodiment 1 includes a computing system or a method implemented by a computing system. The computing system may include a communication interface and at least one processing circuit. The communication interface may be configured to communicate with a robot and a camera having a camera field of view. The at least one processing circuit, when a group of objects is within the camera field of view, receives image information representing the group of objects generated by the camera, and identifies from the image information a plurality of candidate edges associated with the group of objects, where the plurality of candidate edges form respective sets of image positions or physical positions that each represent a respective candidate for a physical edge of the group of objects, or include them. When the plurality of candidate edges includes a first candidate edge formed based on a boundary between a first image region and a second image region, determining whether the image information satisfies a darkness condition defined by the first candidate edge, where the first image region is darker than the second image region, and the first image region and the second image region are respective regions described by the image information. Selecting a subset of the plurality of candidate edges to form a selected subset of candidate edges representing a physical edge of the group of objects, where the selection is based on whether the image information satisfies the darkness condition defined by the first candidate edge, and includes determining whether to retain the first candidate edge as a candidate representing at least one of the physical edges of the group of objects by including the first candidate edge within the selected subset of candidate edges. Outputting a robot interaction movement command, where the robot interaction movement command is for a robot interaction between the robot and at least one object of the group of objects and is generated based on the selected subset of candidate edges, and may be configured to execute. In this embodiment, the at least one processing circuit is configured to determine that the image information satisfies the darkness condition defined by the first candidate edge in response to a determination that the first image region satisfies at least one of a defined darkness threshold criterion or a defined spike intensity profile criterion.Furthermore, in the present embodiment, at least one processing circuit is configured to determine whether the first image region meets the defined darkness threshold criterion by determining whether the first image region has at least one portion with an image intensity darker than the defined darkness threshold. Furthermore, in the present embodiment, at least one processing circuit is configured to determine whether the first image region meets the spike intensity profile criterion by determining whether the first image region has an image intensity profile that includes (i) a first profile portion where the image intensity increases as a function of position and reaches a peak level of darkness at a position within the first image region, and subsequently (ii) a second profile portion where the image intensity decreases as a function of position away from the peak level of darkness within the first image region.

[0104] Embodiment 2 includes the computing system described in Embodiment 1, wherein the first image region is a first pixel region that forms a band of pixels representing a physical gap between a first object and a second object of a group of objects, and the second image region is a second pixel region that is directly adjacent to the first pixel region such that a boundary forming a first candidate edge is between the first pixel region and the second pixel region.

[0105] Embodiment 3 includes the computing system described in Embodiment 2, wherein at least one processing circuit is configured to determine whether the first image region meets the defined darkness threshold criterion by determining whether the first image region has a pixel intensity value smaller than the defined darkness threshold.

[0106] Embodiment 4 includes the computing system of Embodiment 2 or 3, where the image intensity profile of the first image region describes the pixel intensity value of each pixel in a series of pixels extending across the width dimension of the first image region, and at least one processing circuit determines whether the first image region meets the spike intensity profile criterion by determining whether the image intensity profile has a shape such that each pixel intensity value decreases towards the minimum pixel intensity value in the first image region and then switches to increase away from the minimum pixel intensity value, and the minimum pixel intensity value is related to the peak level of darkness in the first image region.

[0107] Embodiment 5 includes the computing system described in any one of Embodiments 1 to 4, and at least one processing circuit is configured to determine that the first image region meets the defined darkness condition only in response to a determination that the first image region meets the spike intensity profile criterion.

[0108] Embodiment 6 includes the computing system described in any one of Embodiments 1 to 5, and at least one processing circuit is configured to determine that the first image region meets the defined darkness condition only in response to a determination that the first image region meets the defined darkness threshold criterion.

[0109] Embodiment 7 includes the computing system described in Embodiment 1, and at least one processing circuit is configured to determine that the first image region meets the defined darkness condition only in response to a determination that the first image region meets both the defined darkness threshold criterion and the defined spike intensity profile criterion.

[0110] Embodiment 8 includes the computing system described in any one of Embodiments 1 to 7, and at least one processing circuit is configured to identify a first candidate edge formed based on the boundary between the first image region and the second image region based on 2D image information when the image information includes 2D image information and 3D image information, and the 3D image information includes depth information of positions within the camera's field of view.

[0111] Embodiment 9 includes the computing system of Embodiment 8, and at least one processing circuit is configured to determine whether to retain a first candidate edge as a candidate for representing at least one of the physical edges of a group of objects when (i) depth information for one or more positions corresponding to the first candidate edge is missing in the 3D image information, and (ii) a part of the 3D image information corresponding to the first candidate image is affected by an imaging noise level greater than a defined noise tolerance threshold.

[0112] Embodiment 10 includes the computing system of Embodiment 8 or 9, and at least one processing circuit is configured to determine whether to retain the first candidate edge as a candidate for representing at least one of the physical edges of a group of objects when the 3D image information does not satisfy a defined depth discontinuity condition at one or more positions corresponding to the first candidate edge.

[0113] Embodiment 11 includes the computing system of Embodiment 10, and at least one processing circuit is configured to determine that the 3D image information does not satisfy a defined depth discontinuity condition at one or more positions corresponding to the first candidate edge in response to a determination that the 3D image information does not describe a depth change at one or more positions exceeding a defined depth difference threshold.

[0114] Embodiment 12 includes the computing system of any one of Embodiments 8 to 11, and at least one processing circuit is configured to identify a second candidate edge among a plurality of candidate edges based on the 3D image information.

[0115] Embodiment 13 includes the computing system according to claim 12, and at least one processing circuit identifies a first surface of a group of objects based on a first set of positions described by 3D image information, each having a depth value that does not deviate from each other beyond a defined measurement variance threshold, and identifies a second surface of the group of objects based on a second set of positions described by 3D image information, each having a depth value within the defined measurement variance threshold, determines an average depth value associated with the first surface as a first average depth value, determines an average depth value associated with the second surface as a second average depth value, and in response to a determination that the difference between the first average depth value and the second average depth value exceeds a defined depth difference threshold, identifies a second candidate edge based on the position where there is a transition between the first surface and the second surface, and is configured to identify a second candidate edge based on 3D image information.

[0116] Embodiment 14 includes the computing system of Embodiment 12 or 13, and at least one processing circuit is configured to identify a second candidate edge based on 3D image information when the second candidate edge is mapped to a candidate edge formed based on a boundary between two image regions that are in the 2D image information and do not meet a defined darkness condition.

[0117] Embodiment 15 includes the computing system of Embodiments 1 to 14, and at least one processing circuit is configured to perform an object recognition operation or an object registration operation based on a selected subset of candidate edges, and a robot interaction movement command is generated based on the result of the object recognition operation or the object registration operation.

[0118] Embodiment 16 includes the computing system according to any one of Embodiments 1 to 15, and at least one processing circuit is configured to select a subset of a plurality of candidate edges by determining which candidate edges to filter from the plurality of candidate edges, and after the plurality of candidate edges are filtered, form a subset of candidate edges.

[0119] It will be apparent to those skilled in the relevant art that other suitable modifications and adaptations to the methods and uses described herein can be made without departing from the scope of any of the embodiments. The embodiments described above are illustrative examples and the present invention should not be construed as being limited to these specific embodiments. It should be understood that the various embodiments disclosed herein may be combined in combinations different from those specifically presented in the description and the accompanying figures. By way of example, any particular act or event of any of the processes or methods described herein may be performed in a different sequence and may be added, incorporated, or completely omitted (e.g., all acts or events described may not be necessary to carry out the method or process). In some examples, method 4000 may be modified to omit step 4002. The various embodiments described above relate to steps 4002-4008 of method 4000, but another method of the present disclosure may include identifying candidate edges based on 3D image information, as discussed with respect to FIGS. 11B or 11C, and may omit steps 4002-4008. Additionally, although certain features of the embodiments herein are described as being performed by a single component, module, or unit for clarity, it should be understood that the features and functions described herein may be performed by any combination of components, modules, or units. Accordingly, those skilled in the art can make various changes and modifications without departing from the spirit or scope of the invention as defined by the appended claims.

Claims

1. A computing system, comprising a communication interface configured to communicate with a robot and a camera having a camera field of view, and at least one processing circuit, wherein the at least one processing circuit, when a group of objects is within the camera field of view, receives image information representing the group of objects generated by the camera, identifies candidate edges from the image information, determines whether a portion of the image information adjacent to the candidate edge satisfies an intensity profile criterion based on whether the portion of the image information includes (i) a first profile portion where the image intensity increases in darkness and subsequently (ii) a second profile portion where the image intensity decreases in darkness, outputs a robot interaction movement command, and is configured to execute, wherein the robot interaction movement command is for robot interaction between the robot and at least one object of the group of objects and is based on the candidate edge, the computing system.

2. The computing system according to claim 1, wherein when the intensity profile criterion is satisfied, the portion of the image information adjacent to the candidate edge is a first pixel region forming a band of pixels representing a physical gap between a first object and a second object of the group of objects.

3. The image intensity profile of the portion of the image information adjacent to the candidate edge describes the pixel intensity value of each of a series of pixels extending across the width dimension of the portion, The at least one processing circuit is configured to determine whether a portion of the image information adjacent to the candidate edge satisfies a spike intensity profile criterion by determining whether the image intensity profile has a shape in which the respective pixel intensity values decrease towards the minimum pixel intensity value and then switch to increase away from the minimum pixel intensity value, wherein the minimum pixel intensity value is related to a peak level of darkness in the portion of the image information adjacent to the candidate edge. The computing system according to claim 1.

4. The at least one processing circuit is only in response to a determination that a portion of the image information adjacent to the candidate edge satisfies the spike intensity profile criterion, configured to determine that a portion of the image information adjacent to the candidate edge satisfies the intensity profile criterion. The computing system according to claim 3.

5. The at least one processing circuit is when the image information includes 2D image information and 3D image information, configured to identify the candidate edge based on the 2D image information, wherein the 3D image information includes depth information of positions within the camera field of view. The computing system according to claim 1.

6. The at least one processing circuit is (i) whether depth information of one or more positions corresponding to the candidate edge is missing from the 3D image information, or (ii) when a portion of the 3D image information corresponding to the candidate edge is affected by an imaging noise level greater than a defined noise tolerance threshold, configured to determine whether to retain the candidate edge as a candidate representing at least one physical edge of the group of objects. The computing system according to claim 5.

7. The at least one processing circuit is When the 3D image information does not satisfy a defined depth discontinuity state at one or more positions corresponding to the candidate edge, The computing system according to claim 5, configured to determine whether to retain the candidate edge as a candidate for representing at least one physical edge of the group of objects.

8. The at least one processing circuit, In response to a determination that the 3D image information does not describe a depth change at one or more positions exceeding a defined depth difference threshold, The computing system according to claim 7, configured to determine that the 3D image information does not satisfy the defined depth discontinuity state at the one or more positions corresponding to the candidate edge.

9. The at least one processing circuit is configured to identify a second candidate edge based on the 3D image information, according to claim 5 of the computing system.

10. The at least one processing circuit, Identifying a first surface of the group of objects based on a first set of positions described by the 3D image information, each having a depth value that does not deviate from each other beyond a defined measurement variance threshold; Identifying a second surface of the group of objects based on a second set of positions described by the 3D image information, each having a depth value within the defined measurement variance threshold; Determining an average depth value associated with the first surface as a first average depth value; Determining an average depth value associated with the second surface as a second average depth value; In response to a determination that the difference between the first average depth value and the second average depth value exceeds a defined depth difference threshold, identifying the second candidate edge based on a position where there is a transition between the first surface and the second surface; The computing system according to claim 9, configured to identify the second candidate edge based on the 3D image information.

11. The at least one processing circuit When the second candidate edge is mapped to a candidate, which is formed based on a boundary between two image regions that are in the 2D image information and do not satisfy the intensity profile criterion The computing system according to claim 9, configured to identify the second candidate edge based on the 3D image information.

12. The at least one processing circuit is configured to perform an object recognition operation or an object registration operation based on the candidate edge, the computing system according to claim 1.

13. A non-transitory computer-readable medium having instructions The instructions, when executed by at least one processing circuit of a computing system, cause the at least one processing circuit to Receive image information by the at least one processing circuit of the computing system, wherein the computing system is configured to communicate with (i) a robot and (ii) a camera having a camera field of view, and the image information is for representing a group of objects within the camera field of view and is generated by the camera Identify candidate edges from the image information Determine whether a portion of the image information adjacent to the candidate edge satisfies an intensity profile criterion, based on whether the portion of the image information includes (i) a first profile portion where the image intensity increases in darkness and subsequently (ii) a second profile portion where the image intensity decreases in darkness Output a robot interaction movement command and cause to execute A non - transitory computer - readable medium, wherein the robot interaction movement command is for robot interaction between the robot and at least one object of the group of objects, and is based on the candidate edge. **Claim 14** The non - transitory computer - readable medium according to claim 13, wherein when the intensity profile criterion is satisfied, the portion of the image information adjacent to the candidate edge forms a first pixel region that is a band of pixels representing a physical gap between a first object and a second object of the group of objects. **Claim 15** A method performed by a computing system, comprising: Receiving, by the computing system, image information, wherein the computing system is configured to communicate with (i) a robot and (ii) a camera having a camera field of view, the image information being for representing a group of objects within the camera field of view and being generated by the camera; Identifying, from the image information, candidate edges; Determining whether a portion of the image information adjacent to the candidate edge satisfies an intensity profile criterion, based on whether the portion of the image information includes (i) a first profile portion where the image intensity increases in darkness and subsequently (ii) a second profile portion where the image intensity decreases in darkness; Outputting a robot interaction movement command; wherein the robot interaction movement command is for robot interaction between the robot and at least one object of the group of objects and is based on the candidate edge.

Citation Information

Patent Citations

  • Edge detecting method of object to be welded applicable to robot

    JP1996118021A

  • Method and device for determining boundary position, program for making computer function as boundary position determination device, and recording medium

    JP2007298376A

  • Information processing apparatus, information processing method, and program

    JP2016197287A

  • Contour extraction device and contour extraction method

    JP2019174931A

  • Information processor, robot system, information processing method and program

    JP2019211903A