Method and system for performing image classification for object recognition

The computing system classifies image portions using bitmaps to distinguish texture from textureless objects, enhancing object recognition and robotic interaction accuracy in environments like warehouses and factories.

JP7709153B2Active Publication Date: 2025-07-16MUJIN INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021022439
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-12
Filing Date
2021-02-16
Publication Date
2025-07-16
Estimated Expiration
2040-10-28

AI Technical Summary

Technical Problem

Existing image classification systems struggle to accurately distinguish between textured and textureless objects, which affects the reliability and efficiency of robotic interaction and object recognition in environments like warehouses and factories.

Method used

A computing system that classifies image portions as having texture or no texture by generating bitmaps based on visual features, such as descriptors and edges, and adjusts for lighting effects, using a fused bitmap to determine the texture classification, which informs robotic interaction plans.

Benefits of technology

Enhances the accuracy of object recognition and robotic interaction by differentiating between textured and textureless objects, improving the reliability and efficiency of automated tasks in environments like warehouses and factories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007709153000001
    Figure 0007709153000001
  • Figure 0007709153000002
    Figure 0007709153000002
  • Figure 0007709153000003
    Figure 0007709153000003
Patent Text Reader

Abstract

To facilitate tasks such as automatic tracking of a package, inventory management, or interaction between an object and a robot, by images.SOLUTION: A system for classifying at least portions of an image into a textured one and a nontextured one, receives an image produced by an image capture device. The image represents one or more objects in a visual field of the image capture device. The system generates one or more bitmaps based on at least one image portion of the image. The system describes, by one or more bitmaps, whether one or more features for feature detection are present in at least one image portion, whether at least one or more visual features for the feature detection are present in at least one image portion, or whether a variation in intensity is present across at least one image portion. The system determines whether at least one image portion is classified into the textured or nontextured one, based on one or more bitmaps.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 959,182, filed on January 10, 2020, entitled "Robot System with Object Detection", the entire content of which is incorporated herein by reference.

[0002] This disclosure relates to computing systems and methods for image classification. In particular, embodiments herein relate to classifying an image or a portion thereof, with or without texture.

Background Art

[0003] As automation becomes more common, images representing objects may be used to automatically extract information about objects such as boxes or other packages in a warehouse, factory, or retail space. The images can facilitate tasks such as automated tracking of packages, inventory management, or robotic interaction with objects.

Summary of the Invention

[0004] In an embodiment, a computing system is provided that includes a non-transitory computer-readable medium and a processing circuit. The processing circuit is configured to perform the following methods, namely, receiving an image by the computing system, wherein the computing system is configured to communicate with an image capture device, and the image is generated by the image capture device and is for representing one or more objects within the field of view of the image capture device; generating, by the computing system, one or more bitmaps based on at least one image portion of the image, wherein the one or more bitmaps and the at least one image portion are associated with a first object among the one or more objects, and the one or more bitmaps describe whether one or more visual features for feature detection are present in the at least one image portion or describe whether there are intensity variations across the at least one image portion. Additionally, the method includes determining, based on the one or more bitmaps, whether to classify the at least one image portion as having texture or no texture, and executing an operation plan for robotic interaction with the one or more objects based on whether the at least one image portion is classified as having texture or no texture. In an embodiment, the method may be performed by executing instructions on a non-transitory computer-readable medium.

Brief Description of the Drawings

[0005]

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 1E

Figure 1F

[0006]

Figure 2A

Figure 2B

Figure 2C

[0007]

Figure 3

[0008]

Figure 4A

Figure 4B

Figure 4C

Figure 4D

[0009]

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 5E

[0010]

Figure 6

[0011]

Figure 7

[0012]

Figure 8A

Figure 8B

Figure 8C

DETAILED DESCRIPTION OF THE INVENTION

[0013] The present disclosure relates to systems and methods for classifying at least a portion of an image as either with texture or without texture. In some cases, the classification may be part of an object registration process for determining the characteristics of a group of one or more objects, such as boxes or other packages arriving at a warehouse or retail space. These characteristics can be determined, for example, to facilitate the automatic handling or other interactions of the group of objects, or other objects having substantially the same design as the group of objects. In embodiments, a portion of an image (also referred to as an image portion) that may be generated by a camera or other image capture device may represent one of the one or more objects and provide an indication of whether there is any visual detail present on the surface of the object, whether there is at least a certain amount or quality of visual detail on the surface of the object, and / or whether there is at least a certain amount of variation in the visual detail. In some cases, the image portion may be used to generate a template for object recognition. In such cases, the classification of whether the image or image portion forms a textured template or a non-textured template may be involved. The template may describe, for example, the appearance of the object (also referred to as the object appearance) and / or the size of the object (also referred to as the object size). In embodiments, the template may be used, for example, to identify objects having a matching object appearance or, more broadly, any other object that matches the template. Such a match may indicate that two objects belong to the same object design and, more specifically, may indicate that they have other characteristics such as the same or substantially the same object size. In some scenarios, when a particular object has an appearance that matches an existing template, such a match may facilitate robotic interaction. For example, the match may indicate that the object has the object size (e.g., object dimensions or surface area) described by the template. The object size can be used to plan how a robot can pick up the object or interact with the object in other ways.

[0014] In an embodiment, classifying whether at least an image portion has texture or not may involve generating one or more bitmaps (also called one or more masks) based on the image portion. In some cases, some or all of the one or more bitmaps may act as heatmaps indicating the probability or strength of certain characteristics across various positions of the image portion. In some cases, some or all of the one or more bitmaps may be for describing whether the image portion has one or more visual features for object recognition. If the image portion has one or more such visual features, the one or more bitmaps may describe where the one or more features are located within the image portion. By way of example, the one or more bitmaps may include a descriptor bitmap and / or an edge bitmap. The descriptor bitmap may describe whether there is a descriptor in the image portion, or may describe where one or more descriptors are located within the image portion (the terms "or" or "or" in this disclosure may refer to "and / or" or "and / or"). The edge bitmap may describe whether an edge is detected within the image portion, or may describe where one or more edges are located within the image portion.

[0015] In an embodiment, some or all of the one or more bitmaps may be for describing whether there are intensity variations across the image portion. For example, such variations (which may also be called spatial variations) may indicate whether there are variations among the pixel values of the image portion. In some cases, the spatial variation may be described by a standard deviation bitmap that may describe the local standard deviation among the pixel values of the image portion.

[0016] In an embodiment, the classification as to whether at least an image portion has texture or not may involve information from a single bitmap, or may involve information from a fused bitmap that combines a plurality of bitmaps. For example, the fused bitmap may be based on a combination of a descriptor bitmap, an edge bitmap, and / or a standard deviation bitmap. In some cases, a texture bitmap may be generated using the fused bitmap to identify, for example, whether the image portion has one or more textured regions and whether the image portion has one or more textureless regions. In some cases, the texture bitmap may be used to describe the total area or total size occupied by one or more textured regions or one or more textureless regions.

[0017] In an embodiment, the fused bitmap may be generated to correct for the effects of conditions such as excessive light reflected from a shiny object surface that causes glare in the image portion, or light blocked by an object surface that causes a shadow in the image portion. The effects of the illumination condition may be described, for example, by a highlight bitmap and / or a shadow bitmap. In some implementations, the fused bitmap may be further generated based on the highlight bitmap and / or the shadow bitmap.

[0018] In an embodiment, the classification as to whether at least an image portion has texture or not may be based on the information provided by a descriptor bitmap, an edge bitmap, a standard deviation bitmap, a highlight bitmap, a shadow bitmap, a fused bitmap, and / or a texture bitmap. For example, the classification may be performed based on the number of descriptors (if present) detected in the image portion, the total area occupied by textured regions (if present) in the image portion, the total area occupied by textureless regions (if present) in the image portion, and / or the standard deviation associated with the image portion or the fused bitmap.

[0019] In an embodiment, the classification of whether a template, or more broadly, an image portion, has texture or not can affect the way object recognition is performed based on the template. Object recognition based on such classification is discussed in more detail in U.S. Patent Application No. ______ (Attorney Docket No. MJ0054-US / 0077-0012US1), filed on the same day as this specification and entitled "METHOD AND COMPUTING SYSTEM FOR OBJECT RECOGNITION OR OBJECT REGISTRATION BASED ON IMAGE CLASSIFICATION", the entire content of which is incorporated herein by reference. In some cases, the classification may affect the confidence level associated with the result of object recognition. For example, the result of object recognition can be assigned a relatively high confidence level if the object recognition is based on a textured template, and a relatively low confidence level if the object recognition is based on a non-textured template. In some cases, the confidence level associated with the result of object recognition can affect whether the object recognition should be performed again (e.g., using another object recognition technique), and / or how to plan a robotic interaction with a particular object. For example, if the object recognition for an object is based on a non-textured template, the robotic interaction with that object can be controlled to proceed more carefully or more slowly. In some cases, if the object recognition process determines that a particular image portion does not match any existing template, an object registration process can be performed to generate and store a new template based on the image portion.

[0020] Figure 1A shows a system 100 for classifying an image or a portion thereof. The system 100 may include a computing system 101 and an image capture device 141 (also referred to as an image sensing device). The image capture device 141 (e.g., a camera) may be configured to capture or otherwise generate an image representing an environment within the field of view of the image capture device 141. In some cases, the environment may be, for example, a warehouse or a factory. In such cases, the image may represent one or more objects within the warehouse or factory, such as one or more boxes that are subject to robotic interaction. The computing system 101 may receive the image directly or indirectly from the image capture device 141 and process the image to perform, for example, object recognition. As will be discussed in more detail below, the processing may involve classifying whether the image or a portion thereof is with or without texture. In some examples, the computing system 101 and the image capture device 141 may be located within the same facility, such as a warehouse or a factory. In some examples, the computing system 101 and the image capture device 141 may be remote from each other. For example, the computing system 101 may be located in a data center that provides a cloud computing platform.

[0021] In an embodiment, the computing system 101 may receive the image from the image capture device 141 via a data storage device (which may also be referred to as a storage device) or via a network. For example, FIG. 1B may depict a system 100A that is an embodiment of the system 100 of FIG. 1A, including the computing system 101, the image capture device 141, and further including a data storage device 198 (or any other type of non-transitory computer-readable medium). The data storage device 198 may be part of the image capture device 141 or may be separate from the image capture device 141. In this embodiment, the computing system 101 may be configured to access the image by reading (or more generally, receiving) the image from the data storage device 198.

[0022] In FIG. 1B, the memory device 198 may include any type of non-transitory computer-readable medium (or media), which may also be referred to as a non-transitory computer-readable storage device. Such non-transitory computer-readable media or storage devices may be configured to store data and provide access to the data. Examples of non-transitory computer-readable media or storage devices include, for example, computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), solid state drives, static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile discs (DVD), and / or memory sticks, etc., electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof, but are not limited thereto.

[0023] FIG. 1C depicts a system 100B that may be an embodiment of the systems 100 / 100A of FIGS. 1A and 1B and includes a network 199. More specifically, the computing system 101 may receive an image generated by the image capture device 141 via the network 199. The network 199 may provide individual network connections or a series of network connections such that the computing system 101 can receive image data consistent with the embodiments herein. In an embodiment, the network 199 may be connected via a wired or wireless link. Wired links may include digital subscriber line (DSL), coaxial cable lines, or fiber optic lines. Wireless links may include Bluetooth®, Bluetooth Low Energy (BLE), ANT / ANT+, ZigBee, Z-Wave, Thread, Wi-Fi®, Worldwide Interoperability for Microwave Access (WiMAX®), Mobile WiMAX®, WiMAX®-Advanced, NFC, SigFox, LoRa, Random Phase Multiple Access (RPMA), Weightless-N / P / W, infrared channels, or satellite bands. Wireless links may also include any cellular network standard for communicating between mobile devices, including standards qualified for 2G, 3G, 4G, or 5G. The wireless standard may use various channel access methods such as, for example, FDMA, TDMA, CDMA, or SDMA. Network communication may be implemented by any suitable protocol including, for example, http, tcp / ip, udp, Ethernet, ATM, etc.

[0024] In an embodiment, the network 199 can be any type of network. The geographical scope of the network can vary widely, and the network 199 can be a Body Area Network (BAN), a Personal Area Network (PAN), a Local Area Network (LAN) such as an intranet, a Metropolitan Area Network (MAN), a Wide Area Network (WAN), or the Internet. The topology of the network 199 can be in any form and can include, for example, any of the following: point-to-point, bus, star, ring, mesh, or tree. The network 199 can consist of any such network topology known to those skilled in the art that can support the operations described herein. The network 199 can utilize different technologies, and layers or stacks of protocols, including, for example, Ethernet protocol, Internet Protocol Suite (TCP / IP), ATM (Asynchronous Transfer Mode) technology, SONET (Synchronous Optical Networking) protocol, or SDH (Synchronous Digital Hierarchy) protocol. The network 199 can be a type of broadcast network, a telecommunications network, a data communication network, or a computer network.

[0025] In an embodiment, the computing system 101 and the image capture device 141 may communicate by direct connection rather than network connection. For example, the computing system 101 in such embodiments may be configured to receive images from the image capture device 141 via a dedicated communication interface such as an RS-232 interface, a Universal Serial Bus (USB) interface, and / or a local computer bus such as a Peripheral Component Interconnect (PCI) bus.

[0026] In an embodiment, the computing system 101 may be configured to communicate with a spatial structure sensing device. For example, FIG. 1D shows a system 100C (which may be an embodiment of system 100 / 100A / 100B) that includes a computing system 101, an image capture device 141, and further includes a spatial structure sensing device 142. The spatial structure sensing device 142 may be configured to sense the 3D structure of an object within its field of view. For example, the spatial structure sensing device 142 may be a depth sensing camera (e.g., a time-of-flight (TOF) camera or a structured light camera) configured to generate spatial structure information such as a point cloud that describes how the structure of the object is arranged in 3D space. More specifically, the spatial structure information may include depth information such as a set of depth values that describe the depth at various positions on the surface of the object. The depth may be relative to the spatial structure sensing device 142 or some other reference frame.

[0027] In an embodiment, the image generated by the image capture device 141 may be used to facilitate the control of the robot. For example, FIG. 1E shows a robot operation system 100D (which is an embodiment of system 100) that includes a computing system 101, an image capture device 141, and a robot 161. The image capture device 141 may be configured to generate, for example, an image representing an object in a warehouse or other environment, and the robot 161 may be controlled to interact with the object based on the image. For example, the computing system 101 may be configured to receive the image and perform object recognition based on the image. Object recognition may involve, for example, determining the size or shape of the object. In this example, the interaction of the robot 161 with the object may be controlled based on the determined size or shape of the object.

[0028] In an embodiment, the computing system 101 may form, or be part of, a robot control system (also referred to as a robot controller) configured to control the movement or other operations of the robot 161. For example, the computing system 101 in such an embodiment may be configured to execute an operation plan for the robot 161 based on an image generated by the image capture device 141 and generate one or more movement commands (e.g., motion commands) based on the operation plan. The computing system 101 in such an example may output one or more movement commands to the robot 161 to control the movement of the robot 161.

[0029] In an embodiment, the computing system 101 may be separated from the robot control system and may be configured to transmit information to the robot control system to enable the robot control system to control the robot. For example, FIG. 1F depicts a robot operation system 100E (an embodiment of the system 100 in FIG. 1A) including the computing system 101 and a robot control system 162 separated from the computing system 101. The computing system 101 and the image capture device 141 in this example may form a vision system 150 configured to provide information about the environment of the robot 161, and more specifically, about the objects in that environment, to the robot control system 162. The computing system 101 may function as a vision controller configured to process an image generated by the image capture device 141 and determine information about the environment of the robot 161. The computing system 101 may be configured to transmit the determined information to the robot control system 162, and the robot control system 162 may be configured to execute an operation plan for the robot 161 based on the information received from the computing system 101.

[0030] As described above, the image capture device 141 of FIGS. 1A - 1F can be configured to capture an image representing one or more objects within the environment of the image capture device 141 or to generate image data that forms an image. More specifically, the image capture device 141 may have a device field of view and may be configured to generate an image representing one or more objects within the device field of view. As used herein, image data refers to any type of data (also referred to as information) that describes the appearance of one or more physical objects (also referred to as one or more objects). In an embodiment, the image capture device 141 may be a camera, such as a camera configured to generate a two - dimensional (2D) image, or may include a camera. The 2D image may be, for example, a grayscale image or a color image.

[0031] As further mentioned above, the image generated by the image capture device 141 may be processed by the computing system 101. In an embodiment, the computing system 101 may include, or be configured as, a server (e.g., having one or more server blades, processors, etc.), a personal computer (e.g., a desktop computer, a laptop computer, etc.), a smartphone, a tablet computing device, and / or any other computing system. In an embodiment, all of the functionality of the computing system 101 may be performed as part of a cloud computing platform. The computing system 101 may be a single computing device (e.g., a desktop computer or a server) or may include multiple computing devices.

[0032] FIG. 2A provides a block diagram illustrating an embodiment of a computing system 101. The computing system 101 includes at least one processing circuit 110 and a non-transitory computer-readable medium (or media) 120. In an embodiment, the processing circuit 110 includes one or more processors, one or more processing cores, a programmable logic controller ("PLC"), an application specific integrated circuit ("ASIC"), a programmable gate array ("PGA"), a field programmable gate array ("FPGA"), any combination thereof, or any other processing circuit.

[0033] In an embodiment, the non-transitory computer-readable medium 120 is a storage device such as an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, for example, a computer diskette, a hard disk, a solid state drive (SSD), a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, any combination thereof, or any other storage device. In some examples, the non-transitory computer-readable medium 120 may include a plurality of storage devices. In a particular case, the non-transitory computer-readable medium 120 is configured to store image data received from an image capture device 141. In a particular case, the non-transitory computer-readable medium 120 further stores computer-readable program instructions that, when executed by the processing circuit 110, cause the processing circuit 110 to perform one or more methods described herein, such as the methods described in connection with FIG. 3.

[0034] FIG. 2B depicts a computing system 101A, which is an embodiment of the computing system 101 and includes a communication interface 130. The communication interface 130 may be configured to receive, for example, an image, or more broadly, image data, from the image capture device 141, via the storage device 198 of FIG. 1B, the network 199 of FIG. 1C, or by a more direct connection, etc. In an embodiment, the communication interface 130 may be configured to communicate with the robot 161 of FIG. 1D or the robot control system 162 of FIG. 1E. The communication interface 130 may include, for example, a communication circuit configured to communicate by a wired or wireless protocol. By way of example, the communication circuit may include an RS-232 port controller, a USB controller, an Ethernet controller, a Bluetooth® controller, a PCI bus controller, any other communication circuit, or a combination thereof.

[0035] In an embodiment, the processing circuit 110 may be programmed by one or more computer-readable program instructions stored in the non-transitory computer-readable medium 120. For example, FIG. 2C shows a computing system 101B, which is an embodiment of the computing system 101, in which the processing circuit 110 is programmed by, or configured to execute, an image access module 202, an image classification module 204, an object recognition module 206, an object registration module 207, and a motion planning module 208. It will be understood that the functionality of the various modules discussed herein is representative and not limiting.

[0036] In an embodiment, the image access module 202 may be a software protocol operating on the computing system 101B and may be configured to access (e.g., receive, read, store) an image, or more broadly, image data. For example, the image access module 202 may be configured to access image data stored in the non-transitory computer-readable medium 120 or 198, or via the network 199 and / or the communication interface 130 of FIG. 2B. In some cases, the image access module 202 may be configured to receive image data directly or indirectly from the image capture device 141. The image data may be for representing one or more objects within the field of view of the image capture device 141. In an embodiment, the image classification module 204 may be configured to classify an image or an image portion, with or without texture, as will be discussed in more detail below, and the image may be represented by the image data accessed by the image access module 202.

[0037] In an embodiment, the object recognition module may be configured to perform object recognition based on the appearance of an object. As described above, object recognition may be based on one or more templates, such as template 210 in FIG. 2C. These templates may be stored on the computing system 101B as depicted in FIG. 2C, or may be stored elsewhere, such as in a database hosted by another device or group of devices of the apparatus. In some cases, each of the templates may include, or be based on, respective image portions received by the image access module 202 and classified with or without texture by the image classification module 204. The object recognition module 206 may use the templates, for example, to perform object recognition on an object appearing in another image portion. If the object recognition module 206 determines that an image portion does not match any existing template in the template storage space (e.g., the non-transitory computer-readable medium 120, or the database discussed above), or if there is no template in the template storage space, the object registration module 207 may, in some instances, be configured to generate and store a new template based on that image portion. In an embodiment, the motion planning module 208 may be configured to execute a motion plan for controlling the robot's interaction with the object, based on, for example, the classification performed by the image classification module 204 and / or the results of the object recognition module 206, as discussed in more detail below.

[0038] In various embodiments, the terms "software protocol," "software instruction," "computer instruction," "computer-readable instruction," and "computer-readable program instruction" are used to describe software instructions or computer code configured to perform various tasks and operations. As used herein, the term "module" broadly refers to a collection of software instructions or code configured to cause the processing circuitry 110 to perform one or more functional tasks. For convenience, in practice, when programming a hardware processor to perform various operations and tasks by various modules, computer instructions, and software protocols, the modules, management units, computer instructions, and software protocols will be described as performing their operations or tasks. Although described as "software" in various places, it is understood that the functionality performed by "modules," "software protocols," and "computer instructions" may more broadly be implemented as firmware, software, hardware, or any combination thereof. Further, embodiments of the present specification describe with respect to method steps, functional steps, and other types of occurrences. In an embodiment, these actions occur in accordance with computer instructions or software protocols executed by the processing circuitry 110 of the computing system 101.

[0039] FIG. 3 is a flowchart illustrating an exemplary method 300 for classifying an image or an image portion with or without texture. The image can represent, for example, one or more objects in a warehouse, a retail space, or other facilities. For example, FIG. 4A depicts an environment in which method 300 can be performed. More specifically, FIG. 4A depicts a system 400 that includes a computing system 101, a robot 461 (which can be an embodiment of robot 161), and an image capture device 441 (which can be an embodiment of image capture device 141) having a device field of view 443. The image capture device 441 can be configured to generate an image representing the appearance of a scene within the device field of view 443. For example, when objects 401, 402, 403, 404 are within the device field of view 443, the image capture device 441 can be configured to generate an image representing the appearance of objects 401-404, that is, more specifically, the appearance of objects 401-404. In one example, objects 401-404 can be stacked boxes or other packages that are unloaded from a pallet by robot 461. The appearance of objects 401-404, if present, can include visual markings printed or otherwise disposed on one or more surfaces of objects 401-404. The visual markings can form, or include, for example, letters, logos, or other visual designs or patterns, or a pattern on one or more surfaces of objects 401-404. For example, objects 401, 404 can be boxes each having a pattern 401A / 404A printed on the respective upper surface of box 401 / 404. If boxes 401 / 404 are used to hold merchandise, the pattern 401A / 404A or other visual markings can identify, for example, a brand name or company associated with the merchandise, and / or identify the merchandise itself or other contents of the box. In some situations, the appearance of objects 401-404, if present, can include the outline of a physical item attached to one or more surfaces of objects 401-404. For example, object 403 can have a piece of tape 403A on its upper surface.In some cases, there may be sufficient contrast between one piece of tape 403A and the peripheral area of the object 403 such that the edge of the tape 403A can appear in the image of the object 403.

[0040] In some cases, some or all of the objects (e.g., 401 - 404) within the field of view of the image capture device (e.g., 443) may have a matching appearance, or a substantially matching appearance. More specifically, those objects may each include the same or substantially the same visual markings, such as the same pattern. For example, the pattern 401A printed on the upper surface of the object 401 may be the same as, or substantially the same as, the pattern 404A printed on the upper surface of the object 404. In some cases, the objects (e.g., 401 - 404) may all be examples of a common object design and thus may have a matching appearance. For example, the object design may be a box design for creating a box that holds a particular product or type of product. Such a box design may be accompanied by a specific size and / or a specific visual design or other visual markings. Thus, objects having the same object design may have a matching appearance and / or a matching size (e.g., matching dimensions).

[0041] In an embodiment, the method 300 of FIG. 3 can be performed by the computing system 101 of FIGS. 2A - 2C, and more specifically, by the processing circuit 110. The method 300 may be performed, for example, when an image representing one or more objects (e.g., objects 401 - 404) is stored in a non - transitory computer - readable medium (e.g., 120 of FIGS. 2A - 2C), or when the image is generated by an image capture device (e.g., 441 of FIG. 4A). In an embodiment, the non - transitory computer - readable medium (e.g., 120) may further store a plurality of instructions (e.g., computer program instructions) that, when executed by the processing circuit 110, cause the processing circuit 110 to implement the method 300.

[0042] In an embodiment, method 300 of FIG. 3 may begin with or may include step 302, in which processing circuit 110 of computing system 101 receives an image generated by an image capture device (e.g., 141 / 441) that represents one or more objects (e.g., 401-404) within the device field of view (e.g., 443) of the image capture device (e.g., 141 / 441). For example, FIG. 4B shows an image 420 that represents objects 401-404 of FIG. 4A. Image 420 may be generated by image capture device 441, which may be positioned directly above objects 401-404 in this example. Thus, image 420 may represent the appearance of the respective upper surfaces of objects 401-404, i.e., more specifically, the unobscured portions of the upper surfaces. In other words, image 420 of this example may represent a top perspective view that captures the upper surfaces of objects 401-404. In an embodiment, the received image may be used to create one or more templates for performing object recognition, as discussed in more detail below.

[0043] In some cases, the image received in step 302 may represent multiple objects, such as a stack of boxes. For example, as depicted in FIG. 4B, the entire received image 420 may represent multiple objects, namely, objects 401-404. In this example, each of objects 401-404 may be represented by a specific portion of image 420 (also referred to as an image portion). For example, as shown in FIG. 4C, object 401 may be represented by image portion 421 of image 420. Image portion 421 may be, for example, a rectangular region (e.g., a square region) or other region of image 420. In such examples, method 300 may involve extraction of an image portion (e.g., 421) associated with a particular object (e.g., 401) from the received image (e.g., 420). The particular object, which may also be referred to as a target object, may be an individual object (e.g., 401), such as an individual box identified by computing system 101. The identified object may be a target for performing object recognition or object registration and / or a target for performing robot interaction (e.g., being removed from a pallet).

[0044] In an embodiment, the extraction of the image portion 421 from the image 420 may be based on the identification of the position (also referred to as the image position) within the image 420 where the edge of the object 401 appears and the extraction of the region of the image 420 surrounded by the image position. In some cases, when one or more objects 401-404 are also within the field of view of the spatial structure sensing device (e.g., 142 in FIG. 1D), the computing system 101 may be configured to receive the spatial structure information generated by the spatial structure sensing device (e.g., 142) and extract the image portion 421 with the aid of the spatial structure information. For example, the spatial structure information may include depth information, and the computing system 101 may be configured to determine the position of the edge of the object 401 (also referred to as the edge position) based on the depth information, such as by detecting positions where there are abrupt changes in depth. In this example, the computing system 101 may be configured to map the edge position sensed by the spatial structure sensing device (e.g., 142) to an image position within the image 420 and extract the region surrounded by the image position, and the extracted region may be the image portion (e.g., 421).

[0045] In an embodiment, the image portion 421 may be used, in some cases, to generate a template for performing object recognition, and the template may be classified with or without texture as discussed below with respect to step 308. The template may represent a specific object design, that is, more specifically, the appearance of the object and / or the structure of the object associated with the object design. The structure of the object may describe the object size, such as the length of the object, the width of the object, the height of the object, and / or any other object dimension, or a combination thereof. Object recognition may involve, for example, comparing the appearance of another object with the template, that is, more specifically, with the appearance of the object described by the template. For example, object recognition may include comparing the respective appearances of each of the objects 402-404 to determine which object (if any) has an appearance that matches the template created from the image portion 421. In some cases, the appearance of each of the objects 402-404 may be represented by the corresponding image portion of the images 420 in FIGS. 4B and 4C. As an example, the computing system 101 may determine that the image portion representing the object 404 matches the template created from the image portion 421 and the object 401 (e.g., by the object recognition module 206 in FIG. 2C). Such a match may indicate, for example, that the object 404 has the same object design as the object 401, and more specifically, the same object design as represented by the template. More specifically, the match may indicate that the object 404 has the same object size (e.g., object dimensions) as the object 401 and the object size associated with the object design represented by the template.

[0046] As described above, in some cases, the image 420 can represent multiple objects. In other cases, the image received at step 302 may represent only one object (e.g., only one box). For example, before being received by the computing system 101, the image may represent only a specific object (e.g., object 401), and if present, any portion representing other objects may be removed by the image capture device (e.g., 141 / 441) or by another device (e.g., cropped) to remove any portion representing other objects within the field of view (e.g., 443) of the image capture device (e.g., 141 / 441). In such examples, the image received at step 302 may represent only that specific object (e.g., object 401).

[0047] In an embodiment, step 302 may be performed by the image access module 202 of FIG. 2C. In an embodiment, the image (e.g., 420 of FIG. 4B) may be stored on a non - transitory computer - readable medium (e.g., 120 of FIG. 2C), and receiving the image at step 302 may involve reading (or more broadly, receiving) the image (e.g., 420) from the non - transitory computer - readable medium (e.g., 120) or from any other device. In some situations, the image (e.g., 420) may be received by the computing system 101 from the image capture device (e.g., 141 / 441) via the communication interface 130 of FIG. 2B and may be stored on a non - transitory computer - readable medium (e.g., 120) that can provide a temporary buffer or long - term storage for the image (e.g., 420). For example, the image (e.g., 420) may be received from the image capture device (e.g., 141 / 441 of FIG. 4A) and stored on a non - transitory computer - readable medium (e.g., 120). The image (e.g., 420) may then be received at step 302 by the processing circuit 110 of the computing system 101 from the non - transitory computer - readable medium.

[0048] In some situations, an image (e.g., 420) may be stored in a non-transitory computer-readable medium (e.g., 120), or may be pre-generated by the processing circuit 110 itself based on information received from an image capture device (e.g., 141 / 441). For example, the processing circuit 110 may be configured to generate an image (e.g., 420) based on raw camera data received from an image capture device (e.g., 141 / 441), and may be configured to store the generated image in a non-transitory computer-readable medium (e.g., 120). The image may then be received by the processing circuit 110 at step 302 (e.g., by reading the image from the non-transitory computer-readable medium 120).

[0049] In an embodiment, the image (e.g., 420) received at step 302 may be a two-dimensional (2D) array of pixels, or may include such an array, having respective pixel values (also referred to as pixel intensity values) associated with the intensity of a signal sensed by the image capture device 441, such as the intensity of light reflected from the respective surfaces (e.g., the upper surfaces) of the objects 401-404. In some cases, the image (e.g., 420) may be a grayscale image. In such a case, the image (e.g., 420) may include a single 2D array of pixels, where each pixel may have an integer value or a floating-point value within a range, for example, from 0 to 255 or some other range. In some cases, the image (e.g., 420) may be a color image. In such a case, the image (e.g., 420) may include different 2D arrays of pixels, where each pixel of the 2D arrays may represent the intensity of a respective color component (also referred to as a respective color channel). For example, such a color image may include pixels of a first 2D array representing the red channel and indicating the intensity of the red component of the image (e.g., 420), pixels of a second 2D array representing the green channel and indicating the intensity of the green component of the image (e.g., 420), and pixels of a third 2D array representing the blue channel and indicating the intensity of the blue component of the image (e.g., 420).

[0050] In an embodiment, the computing system 101 may be configured to perform a smoothing operation or a blurring operation on an image (e.g., 420). When performed, the smoothing operation may be part of step 302 or may be performed after step 302 to remove, for example, artifacts or noise (e.g., illumination noise) from the image (e.g., 420). Artifacts may be due to, for example, unevenness (e.g., wrinkles) on the surface of an object, effects from the illumination state (e.g., shadows), or some other factor. In some cases, the smoothing operation may involve the application of a structure-preserving filter, such as a Gaussian filter, to the image (e.g., 420).

[0051] In an embodiment, method 300 of FIG. 3 further includes step 306 of generating one or more bitmaps (also referred to as one or more masks) by processing circuit 110 of computing system 101 based on at least one image portion of an image, such as image portion 421 of images 420 of FIGS. 4C and 4D. The image portion (e.g., 421) may be a portion of an image (e.g., 420) representing a particular object (e.g., 401) that is within the field of view (e.g., 443) of an image capture device (e.g., 441), such as an image portion representing a targeted object listed above. Thus, the one or more bitmaps of step 306 may be particularly associated with the targeted object. If the image received at step 302 (e.g., 420) represents multiple objects (e.g., 401-404), step 306 may, in some instances, be based on only or primarily an image portion (e.g., 421) representing the targeted object (e.g., 401). In other words, in such scenarios, at least one image portion on which one or more bitmaps are based may be limited primarily to an image portion representing the targeted object. In another scenario, if the image received at step 302 represents only the targeted object, step 306 may, in some instances, be based on the entire image. In other words, in such scenarios, at least one image portion on which one or more bitmaps are based may include the entire image or substantially the entire image. In such examples, the image portion associated with the targeted object in such scenarios may occupy the entire image or substantially the entire image such that one or more bitmaps of such scenarios may be generated directly based on the entire image or substantially the entire image. In some cases, step 306 may be performed by image classification module 204 of FIG. 2C.

[0052] In an embodiment, one or more visual features for feature detection can be described by one or more bitmaps as being present in at least one image portion (e.g., 421) that represents an object (e.g., 401). The one or more visual features can represent visual details that can be used to compare the appearance of the object to the appearance of a second object (e.g., 404). Some or all of the visual details (when present in the image portion) may incorporate or represent visual markings (when present) that are printed on or otherwise appear on the object (e.g., 401). When creating a template using the image portion (e.g., 421), the one or more visual features (when present) may represent the visual details described by the template or may be used to facilitate comparison of the template to the appearance of a second object (e.g., 404). In such examples, the implementation of object recognition may involve comparing the appearance of the second object (e.g., 404) to the visual details described by the template.

[0053] In an embodiment, visual details or visual features (when present) within the image portion (e.g., 421) can contribute to the visual texture of the image portion (e.g., 421), that is, more specifically, the visual texture of the surface appearance of the object (e.g., 401) represented by the image portion (e.g., 421). Visual texture can refer to the spatial variation of intensity across the image portion (e.g., 421), that is, more specifically, the pixels of the image portion (e.g., 421) where there is variation between pixel intensity values. For example, visual details or one or more visual features (when some are present) can include lines, corners, or patterns represented by regions of pixels with non-uniform pixel intensity values. In some cases, a rapid variation between pixel intensity values can correspond to a high level of visual texture while uniform pixel intensity values can correspond to the absence of visual texture. The presence of visual texture can facilitate a more robust comparison of the appearance of each object, that is, more specifically, a template generated from the appearance of a first object (e.g., 401) to the appearance of a second object (e.g., 404).

[0054] In an embodiment, some or all of the one or more bitmaps may each indicate whether an image portion (e.g., 421) has one or more visual features for feature detection or lacks visual features for feature detection. When the image portion (e.g., 421) has or represents one or more visual features for feature detection, each bitmap of the one or more bitmaps may indicate the number or amount of visual features present in the image portion (e.g., 421) and / or may indicate where within the image portion (e.g., 421) the one or more visual features are located.

[0055] In an embodiment, some or all of the one or more bitmaps may each represent a particular type of visual feature. For example, the types of visual features may include descriptors as a first type of visual feature and edges as a second type of visual feature. When multiple bitmaps are generated, they may include a first bitmap associated with identifying the presence of descriptors (if present) in at least one image portion of the image and a second bitmap associated with identifying the presence of edges (if present) in at least one image portion.

[0056] More specifically, the one or more bitmaps generated in step 306 may, in embodiments, include a descriptor bitmap (also referred to as a descriptor mask) that describes whether one or more descriptors are present in at least one image portion (e.g., 421) of the image (e.g., 420) received in step 302. As discussed in more detail below, the descriptor bitmap may indicate which areas of the image portion (e.g., 421) are devoid of the descriptor and which areas of the image portion (e.g., 421) (if present) are devoid of the descriptor. In some cases, the descriptor bitmap may act as a heat map that indicates the probability of the descriptor being present at various locations in the image portion. A descriptor (also referred to as a feature descriptor) may be a type of visual feature that describes a particular visual detail that appears in the image portion (e.g., 421), such as a corner or pattern in the image portion. In some cases, the visual detail may have a sufficient level of uniqueness in appearance such that it can be distinguished from other visual details or other types of visual details in the received image (e.g., 420). In some cases, the descriptor may act as a fingerprint for that visual detail by encoding the pixels that represent that visual detail into a scalar value or into a vector.

[0057] As discussed above, the descriptor bitmap, if present, may indicate which locations or regions within an image portion (e.g., 421) have visual details that form a descriptor. For example, FIG. 5A depicts an example of a descriptor bitmap 513, generated based on image portion 421. In this example, the descriptor bitmap 513 may be a 2D array of pixels, where the descriptor is a pixel coordinate [a1b1]. T , [a2b2] T , … [a n b n ] T and / or pixel coordinates [a1b1] T , [a2b2] T , … [a n b n ] T Each of the descriptor identification areas 5141, 5142, ... 514 surrounds nmay be shown as being located at. Descriptor identification regions 5141, 5142, … 514 n may be circular regions, or may have some other shape (e.g., a square shape). In some cases, where a pixel value of zero indicates the absence of a descriptor, all pixels within the descriptor identification regions 5141, 5142, … 514 n of the descriptor bitmap 513 may have non-zero values. The pixel coordinates [a1b1] T of the descriptor bitmap 513, [a2b2] T , … [a n b n T (also referred to as pixel positions) correspond to the same pixel coordinates [a1b1] T of the image portion 421, [a2b2] T , … [a n b n T Accordingly, the descriptor bitmap 513 may show that the pixel coordinates [a1b1] T of the image portion 421, [a2b2] T , … [a n b n T have the visual details forming the respective descriptors, and those descriptors generally are located within or around the region of the image portion 421 that occupies the same positions as regions 5141, 5142, … 514 n .

[0058] In an embodiment, the computing system 101 determines one or more positions (e.g., from [a1b1] T to [a n b n ) within an image portion 421 where a descriptor (if any) exists, or one or more regions (e.g., from 5141 to 514 T n ​​​​It may be configured to generate a descriptor bitmap by searching for ). In this embodiment, in the image portion 421, there is sufficient visual detail or sufficient variation in visual detail at one or more positions or regions, and one or more respective descriptors may be formed at such positions or regions. As an example, the computing system 101 of this embodiment may be configured to search for one or more positions by searching at least the image portion 421 for one or more keypoints (also referred to as descriptor keypoints). Each of the one or more keypoints (if some are found) may be a position or region where there is a descriptor. One or more positions (e.g., [a1b1] T from a n b n ) T or one or more regions (e.g., 5141 to 514 n ) may be equal to or based on the one or more keypoints. The search may be performed using feature detection techniques such as the Harris corner detection algorithm, scale-invariant feature transform (SIFT) algorithm, speeded up robust features (SURF) algorithm, feature from accelerated segment test (FAST) detection algorithm, and / or oriented FAST and rotated binary robust independent elementary features (ORB) algorithm. As an example, the computing system 101 may use the SIFT algorithm to search for keypoints in the image portion 421, and each keypoint may be a circular region having center coordinates of the keypoint and a radius represented by a scale parameter value σ (also referred to as the keypoint scale). In this example, the coordinates [a1b1] T , [a2b2] T , … [a n b n )T may be equal to the center coordinates of the keypoint. On the other hand, from the descriptor identification regions 5141 to 514 n may correspond to the circular region identified by the keypoint. More specifically, each of the descriptor identification regions (e.g., region 5141) may be centered on the keypoint center coordinates (e.g., [a1b1] T ) of the corresponding keypoint and may have a size (e.g., radius) equal to or based on the scale parameter value of the corresponding keypoint.

[0059] In an embodiment, the pixels of the descriptor bitmap (e.g., 513) within one or more descriptor identification regions (e.g., from 5141 to 514 n ) may have non-zero pixel values if any such region is found, while some or all of the other pixels of the bitmap may have pixel value zero (or some other defined value). In this example, if all the pixels of a particular descriptor bitmap have pixel value zero, the descriptor bitmap may indicate that no descriptor was found in the corresponding image portion. Alternatively, if some of the pixels of the descriptor bitmap have non-zero values, the descriptor bitmap (e.g., 513) may indicate the number or amount of descriptors in the corresponding image portion (e.g., 421). For example, some of the descriptors or descriptor identification regions within the descriptor bitmap 513 of FIG. 5A may indicate the quantity (e.g., n descriptors) of descriptors in the image portion 421. In this example, the total area of the descriptor identification regions from 5141 to 514 n may indicate the amount of descriptors or descriptor information in the image portion 421. In some cases, if a descriptor identification region (e.g., 5141) is present in the descriptor bitmap, the size of the descriptor identification region may indicate the size of the corresponding descriptor. For example, the radius of the descriptor identification region 5141 may indicate the size of the corresponding descriptor within the image portion 421 located at the pixel coordinates [a1b1] T . In this example, a larger radius may correspond to a descriptor that occupies a larger area.

[0060] In an embodiment, each center of a descriptor identification region (if present) within a descriptor bitmap (e.g., 513) may have a defined non-zero value. For example, the pixel coordinates [a1b1] within the descriptor bitmap 513 of FIG. 5A T from n a n T may each have a defined maximum pixel value. The defined maximum pixel value may be a defined maximum value recognized for the pixels of the descriptor bitmap 513 (or more generally, for the pixels of any bitmap). For example, if each pixel of the bitmap 513 is an integer value represented by 8 bits, the defined maximum pixel value may be 255. In another example, if each pixel is a floating-point value representing a probability value between 0 and 1 (for the probability of a descriptor present in that pixel), the defined maximum pixel value may be 1. In an embodiment, the pixel values of other pixel coordinates within the descriptor identification region may be less than the defined maximum pixel value and / or may be based on the distance from each center coordinate of the descriptor identification region. For example, the pixel value of the pixel coordinates [x y] within the descriptor identification region 5141 T may be equal to or based on the defined maximum pixel value multiplied by a magnification factor less than 1, where the magnification factor is a function (e.g., a Gaussian function) of the distance between the pixel coordinates [x y] T and the center coordinates [a1b1] of the descriptor identification region 5141 T .

[0061] In an embodiment, one or more bitmaps generated in step 306 may include an edge bitmap (also referred to as an edge mask) that describes whether one or more edges are present in at least one image portion (e.g., 421) of an image (e.g., 420) received in step 302. More specifically, the edge bitmap may be for identifying one or more regions of at least one image portion (e.g., 421) that include one or more respective edges detected from at least one image portion (e.g., 421), or for indicating that no edge is detected in at least one image portion. In some cases, the edge bitmap may act as a heatmap indicating the intensity or probability of edges present at various positions of at least one image portion. As an example, FIG. 5B shows edges 4231 to 423 n in an image portion 421, and edge bitmap 523 that identifies regions 5251 to 525 n corresponding to edges 4231 to 423 n in the image portion 421 of FIG. 5B. More specifically, when edges 4231 to 423 n occupy a specific edge position (e.g., pixel coordinates [g m h m ) in the image portion 421 of FIG. 5B, regions 5251 to 525 T (also referred to as edge identification regions) may surround those positions in edge bitmap 523 (e.g., surround pixel coordinates [g n h m ). For example, edge identification regions 5251 to 525 m T may form a band around their edge positions, and the band may have a defined band thickness or width. n In an embodiment, edge identification regions 5251 to 525

[0062] n ​​All pixels within (if present) may have non-zero pixel values, and some or all other pixels of the edge bitmap 523 may have a pixel value of zero. If all pixels of a particular edge bitmap have a pixel value of zero, the edge bitmap may indicate that no edge is detected in the corresponding image portion. If some pixels of a particular edge bitmap have non-zero pixel values, those pixels may indicate one or more positions or regions where one or more edges are located within the image portion 421. In an embodiment, the edge bitmap (e.g., 523) may indicate the quantity or prevalence of edges within the image portion 421. For example, the total number of edge identification regions (e.g., 5251 to 525 n ) may indicate the quantity of edges within the corresponding image portion (e.g., 421), and the area of the edge identification regions (e.g., 5251 to 525 n ) may indicate the prevalence of edges within the image portion (e.g., 421).

[0063] In an embodiment, pixels that are within the edge bitmap (e.g., 523) and at an edge position (e.g., [g m h m ) T may be set to a defined pixel value, such as the defined maximum pixel value discussed above. In such embodiments, other pixels within the edge identification region (e.g., 5251) that surround the edge position (e.g., [g m h m ) T may have a value less than the defined maximum pixel value. For example, pixels within the edge identification region (e.g., 5251) may have pixel values based on the distance from the edge position. By way of example, a pixel [x y] T within the edge identification region 5251 of FIG. 5B may have a pixel value equal to the defined maximum pixel value multiplied by a magnification factor that is less than 1. In some cases, the magnification factor is the distance of the pixel [x y] Tand the distance function to the nearest edge position (e.g., [g m h m ) T ) may be a function (e.g., a Gaussian function).

[0064] In an embodiment, the computing system 101 may be configured to search for edge positions by using an edge detection technique such as a Sobel edge detection algorithm, a Prewitt edge detection algorithm, a Laplacian edge detection algorithm, a Canny edge detection algorithm, or any other edge detection technique. In an embodiment, the edge detection algorithm may identify 2D edges such as straight lines or curves. The detection may be based on, for example, the identification of pixel coordinates where there are abrupt changes in pixel values.

[0065] In an embodiment, one or more bitmaps generated in step 306 may include a standard deviation bitmap (also referred to as a standard deviation mask). The standard deviation bitmap is for describing whether the intensity varies across at least one image portion (e.g., 421), that is, more specifically, for describing how much the intensity varies across at least one image portion. For example, the standard deviation bitmap may form a 2D array of pixels, where each pixel of the standard deviation bitmap may indicate the standard deviation of the pixel values for the corresponding region of pixels within the image portion (e.g., 421). Since the standard deviation is specific to the region, it may be referred to as a local standard deviation. As an example, FIG. 5C shows a standard deviation bitmap 533 generated from the image portion 421. In this example, the pixel value at a specific pixel coordinate (e.g., [u1v1] T or [u2v2] T ) of the standard deviation bitmap 533 is at the same pixel coordinate (e.g., [u1v1] T or [u2v2] TIt may be equal to or based on the local standard deviation (or other measure of variance) of the pixel values in the region of the image portion 421 (e.g., 4321 or 4322) surrounding [[ID=]]. The region of pixels (e.g., 4321 or 4322) for determining the local standard deviation may be a rectangular region having a defined size, such as a square region that is, for example, 3 pixels by 3 pixels. In some implementations, each pixel of the standard deviation bitmap may have a normalized standard deviation value that is equal to the standard deviation of the pixel values of its corresponding region divided by the size of the corresponding region. For example, the pixel value at [u1v1] in the standard deviation bitmap 533 T may be equal to the standard deviation of the pixel values of the region 4321 of the image portion 421 divided by the area of the region 4321 (e.g., 9 square pixels).

[0066] In an embodiment, if a particular pixel of the standard deviation bitmap (e.g., 533) has a pixel value of zero or substantially zero, that pixel may indicate a local standard deviation of zero for the corresponding region of the image portion (e.g., 421). In such embodiments, the corresponding region of the image portion (e.g., 421) may have no or substantially no variation in the pixel values within that region. For example, the pixel at [u2v2] in the standard deviation bitmap 533 T may have a value of zero, which may indicate that the corresponding region 4322 surrounding the same pixel coordinates [u2v2] in the image portion 421 T has substantially uniform pixel values. In an embodiment, if all pixels of the standard deviation bitmap have a pixel value of zero, the standard deviation bitmap may indicate no intensity variation across the image portion on which the standard deviation bitmap is based. In another embodiment, the pixels of the standard deviation bitmap have non-zero values (e.g., the pixel coordinates [u1v1] of the bitmap 533 TIn the case of (), such pixels may indicate that there are intensity variations across at least the corresponding region (e.g., 4322) of the image portion (e.g., 421). In some cases, higher pixel values in the standard deviation bitmap (e.g., 533) may indicate higher local standard deviations, which may indicate a higher level of variation between pixel values within the image portion.

[0067] In an embodiment, step 306 may include generating a plurality of bitmaps, such as a first bitmap that is a descriptor bitmap (e.g., 513) and a second bitmap that is an edge bitmap (e.g., 523). In some cases, the plurality of bitmaps may include at least three bitmaps, such as a descriptor bitmap, an edge bitmap, and a standard deviation bitmap. In this embodiment, it may be possible to combine information from the plurality of bitmaps to produce more complete information regarding how visually distinctive features are present within the image portion. In some cases, the plurality of bitmaps may describe multiple feature types. For example, the first bitmap may indicate whether one or more features of a first feature type, such as a descriptor, are present in at least one image portion (e.g., 421), and the second bitmap may indicate whether one or more features of a second feature type, such as an edge, are present in at least one image portion (e.g., 421).

[0068] In an embodiment, the computing system 101 can be configured to generate one or more bitmaps indicating the effect of the illumination state on the received image (e.g., 420) or a portion of that image (e.g., 421). In some scenarios, the illumination state may result in too much light or other signals being reflected from regions of the surface of an object (e.g., the upper surface of object 401), causing glare in the resulting image portion (e.g., 421) representing the object. For example, the light may be reflected off a region having a shiny material (e.g., a shiny tape). In some scenarios, the illumination state may cause too little light to be reflected from regions of the surface of an object, resulting in shadows in the resulting image portion. For example, the light may be blocked from reaching regions of the surface of the object entirely. One or more of the bitmaps in this example may be referred to as one or more illumination effect bitmaps and can be considered additional bitmaps added to the plurality of bitmaps discussed above. In an embodiment, glare or shadows within a region of an image or image portion may cause the contrast of any visual details within that region to be lost or the visual details to appear too blurred, reducing the reliability of the visual details for use in object recognition.

[0069] In an embodiment, one or more lighting effect bitmaps (also referred to as one or more lighting effect masks) may include a highlight bitmap (also referred to as a highlight mask), and / or a shadow bitmap (also referred to as a shadow mask). The highlight bitmap may indicate one or more regions (if any) of a corresponding image portion (e.g., 421) that exhibit too much glare, or other effects of too much light reflected on a particular portion of the surface of an object. The glare may saturate an area of the image or image portion, thereby losing contrast of visual details (if any) representing that portion of the surface of the object, or causing the visual details to blend into the glare. FIG. 5D depicts an exemplary highlight bitmap 543 generated based on image portion 421. The highlight bitmap 543 may include regions 5471 and 5472 having pixel values such as non-zero pixel values indicating glare. More specifically, regions 5471 and 5472 (which may be referred to as highlight identification regions) may indicate that glare is present within corresponding regions 4271 and 4272 of image portion 421. Regions 4271 and 4272 (which may also be referred to as highlight regions) of image portion 421 may occupy the same positions as highlight identification regions 5471 and 5472 of highlight bitmap 543. In some cases, pixels in the highlight bitmap (e.g., 543) that indicate the presence of glare within a corresponding image portion (e.g., 421), such as pixels within regions 5471 and 5472, may have defined pixel values such as the defined maximum pixel value discussed above. In other cases, pixels within the highlight identification regions of the highlight bitmap (e.g., 543) may have the same pixel values as the corresponding pixels within the highlight regions of the image portion (e.g., 421). In an embodiment, all pixels other than at least one highlight identification region (e.g., 5471 and 5472) may have a pixel value of zero.

[0070] In an embodiment, the computing system 101 may generate a highlight bitmap by detecting an effect due to glare or other excessive brightness within an image portion. Such detection may be based on, for example, detection of pixel values of the image portion 421 that exceed a defined luminance threshold, such as pixel values within regions 4271 and 4272. As an example of a luminance threshold, when the pixel value is an 8-bit integer in the range from 0 to 255, the defined luminance threshold may be, for example, 230 or 240. When the pixel value of a specific pixel coordinate within the image portion 421 exceeds the defined luminance threshold, the computing system 101 may set the pixel value of the same pixel coordinate within the highlight bitmap 543 to a value related to the identification of glare (e.g., 255).

[0071] In an embodiment, the shadow bitmap may indicate an area (if any) of an image portion (e.g., 421) that represents the effect of light being blocked from fully reaching a portion of the surface of an object. Such a darkening effect may cast a shadow on that portion of the surface of the object. In some examples, the shadow may blur or make any visual details in that area of the image portion (e.g., 421) invisible. For example, FIG. 5E shows a shadow region 4281 within the image portion 421. The computing system 101 may detect the shadow region 4281 as an area of the image portion 421 having pixel values that are at least a defined discrimination threshold less than the pixel values of the surrounding area. In some cases, the shadow region 4281 may be detected as an area having pixel values less than a defined dark threshold. For example, when the pixel value is in the range from 0 to 255, the defined dark threshold may be pixel value 10 or 20.

[0072] FIG. 5E further depicts a shadow bitmap 553 generated based on the image portion 421. More specifically, the shadow bitmap 553 may include a shadow identification region 5581 corresponding to the shadow region 4281. More specifically, the shadow identification region 5581 may occupy the same position in the shadow bitmap 553 as the shadow region 4281 in the image portion 421. In some cases, each pixel in the shadow identification region (e.g., 5581) may have a non-zero value, while all pixels of the shadow bitmap 553 outside the shadow identification region may have a pixel value of zero. In some cases, the pixels of the shadow bitmap (e.g., 553) that are within the shadow identification region, if present, may have a defined pixel value, such as a defined maximum pixel value. In some cases, the pixels in the shadow identification region (e.g., 5581) may have the same pixel value as the corresponding pixels in the shadow region (e.g., 4281).

[0073] Referring back to FIG. 3, method 300 may further include step 308 of determining, by the processing circuitry 110 of computing system 101, based on one or more of the bitmaps described above, whether to classify at least one image portion (e.g., 421) as either with texture or without texture (e.g., by image classification module 204). Such classification may refer to whether there is a sufficient amount of visual texture (if present) in the image or image portion, or whether the appearance of the image or image portion is substantially blank or uniform. As described above, at least one image portion may, in some scenarios, be used as a template for object recognition. In such scenarios, step 308 may involve determining whether to classify the template as a textured template or a textureless template. In an embodiment, step 308 may be performed by image classification module 208 of FIG. 2C.

[0074] In an embodiment, step 308 may involve classifying an image portion as having texture if it meets at least one of one or more criteria. In some cases, the at least one criterion may be based on a single bitmap, such as a descriptor bitmap (e.g., 513) or a standard deviation bitmap (e.g., 533). For example, the determination of whether to classify at least one image portion as having texture or not may be based on whether the total number of descriptors indicated by the descriptor bitmap (e.g., 513) exceeds a defined descriptor quantity threshold, or whether the maximum value, minimum value, or representative value of the local standard deviation values in the standard deviation bitmap 533 exceeds a defined standard deviation threshold. As described above, the descriptor bitmap (e.g., 513) may identify one or more regions of at least one image portion (e.g., 421) each including one or more respective descriptors, or the descriptor may indicate that no descriptor is detected within at least one image portion (e.g., 421).

[0075] In an embodiment, at least one criterion for classifying an image portion as having texture may be based on a combination of a descriptor bitmap (e.g., 513) and an edge bitmap (e.g., 523), a combination of a descriptor bitmap (e.g., 513) and a standard deviation bitmap (e.g., 533), a combination of an edge bitmap and a standard deviation bitmap, or all three bitmaps. For example, the determination in step 308 of whether to classify at least one image portion as having texture or not may include generating a fused bitmap (also called a fusion mask) that combines multiple bitmaps, and the classification is based on the fused bitmap. In some cases, the multiple bitmaps may describe multiple respective types of features. By using multiple types of bitmaps to classify corresponding image portions, it may provide the advantage of leveraging information about the presence or absence of multiple types of features, thereby providing a more complete determination about the amount or number of features (if any) present in the image or image portion. For example, an image portion may have a particular visual detail (e.g., a pink area adjacent to a white area) that cannot be identified as a feature by a first bitmap but can be identified as a feature by a second bitmap.

[0076] In an embodiment, generating a fused bitmap may involve summing a plurality of bitmaps, i.e., more specifically, generating a weighted sum of a plurality of bitmaps. For example, the fused bitmap may be equal to, or based on, M1×W1+M2×W2, or M1×W1+M2×W2+M3×W3, where M1 may refer to a first bitmap (e.g., a descriptor bitmap), M2 may refer to a second bitmap (e.g., an edge bitmap), M3 may refer to a third bitmap (e.g., a standard deviation bitmap), and W1, W2, and W3 may be respective weights associated with bitmaps M1, M2, and M3. In this example, bitmaps M1, M2, and M3 may be referred to as feature bitmaps or variation bitmaps because they represent the presence (or absence) of features in an image portion or represent the variation (or absence of variation) of intensity across an image portion. In an embodiment, a sum or other combination of feature bitmaps or variation bitmaps may be referred to as a combined feature bitmap or a combined variation bitmap. Generating a weighted sum of feature bitmaps or variation bitmaps may involve, for example, adding the bitmaps on a pixel-by-pixel basis. For example, the pixel value for the pixel coordinates [x y] T of the fused bitmap may be equal to the sum of W1 multiplied by the pixel value for [x y] T of the first bitmap M1, W2 multiplied by the pixel value for [x y] T of the second bitmap M2, and W3 multiplied by the pixel value for [x y] T of the third bitmap M3. In an embodiment, weights W1, W2, W3 may be predefined. In an embodiment, weights W1, W2, and W3 may be determined by computing system 101 via a machine learning algorithm, as discussed in more detail below.

[0077] In an embodiment, the generation of the fused bitmap may further be based on one or more lighting effect bitmaps, such as a highlight bitmap (e.g., 543) and a shadow bitmap (e.g., 553). For example, the computing system 101 may determine pixel values, also referred to as bitmap pixel values, that describe the visual texture level across at least one image portion (e.g., 421) of the image. The bitmap pixel values may be based on the combined feature bitmap or the combined variation bitmap discussed above, such as pixel value M1×W1+M2×W2, or M1×W1+M2×W2+M3×W3. In this example, the computing system 101 may reduce or otherwise adjust a subset of the determined bitmap pixel values of the combined feature bitmap or the combined variation bitmap, and the adjustment may be based on the highlight map (e.g., 543) and / or the shadow bitmap (e.g., 553). For example, the highlight bitmap or the shadow bitmap may identify one or more regions of at least one image portion (e.g., 421) as being glare, or a shadow, or within a shadow. The computing system 101 may perform an adjustment to reduce the bitmap pixel values in the same one or more regions of the combined feature bitmap or the combined variation bitmap. The reduction may reduce the influence of the pixel values in those regions when classifying the image portion as textured or textureless, as those bitmap pixel values may be affected by lighting that reduces the reliability or quality of the visual information from those one or more regions. In an embodiment, the reduction may be based on multiplying the combined feature bitmap or the combined variation bitmap by the highlight bitmap and / or the shadow bitmap.

[0078] As an example of the above consideration, FIG. 6 shows a fused bitmap 631 generated based on the combination of a feature bitmap and an illumination effect bitmap. More specifically, FIG. 6 depicts a computing system 101 that generates a fused bitmap such that it is equal to (M1×W1+M2×W2+M3×W3)×(M4×W4+M5×W5), where M4 is a highlight bitmap, M5 is a shadow bitmap, and W4 and W5 are respective weights associated with bitmaps M4 and M5. In this example, M1×W1+M2×W2+M3×W3 may form a combined feature bitmap or a combined variation bitmap 621, and the bitmap 621 may be multiplied by a combined illumination effect bitmap 623 equal to (M4×W4+M5×W5).

[0079] As described above, the weights W1 through W5 can be determined by machine learning techniques in the example. For example, the machine learning technique may involve using training data to determine optimal values for the weights W1 through W5. In some cases, the training data may include training images or portions of training images, which may be images or portions of images with a predefined classification as to whether they have texture or not. In such a case, the computing system 101 may be configured to determine optimal values for the weights W1 through W5 that minimize the classification error for the training images. For example, the computing system 101 may be configured to adjust the weights W1 through W5 to their optimal values using a gradient descent process.

[0080] In an embodiment, the computing system 101 may be configured to determine the values of weights W1 to W5 based on predefined information regarding objects that are likely to be within the field of view of the image capture device (e.g., 443). For example, when the computing system 101 receives a sign (e.g., from a warehouse management department) that the image capture device (e.g., 441) has captured or will capture an object that is likely to have many visual markings that will appear as edges, the weight W2 may be assigned a relatively high value to emphasize the edge bitmap M2. When the computing system 101 receives a sign that an object is likely to have visual markings that form a descriptor, the weight W1 may be assigned a relatively high value to emphasize the descriptor bitmap M1. In some cases, the computing system 101 may be configured to determine the values for weights W1 to W5 based on downstream analysis, such as determining which bitmap has more information (e.g., more non-zero values). In such examples, the weight for the bitmap with more information (e.g., M1) may be assigned a relatively high weight. In some cases, the computing system 101 may be configured to assign values to the weights based on defined priorities regarding which type of feature detection to use or emphasize. For example, if the defined priority indicates emphasizing edge-based detection, the computing system may assign a relatively high value to W2. If the defined priority indicates emphasizing descriptor-based detection, the computing system may assign a relatively high value to W1.

[0081] In an embodiment, when the image received in step 302 (e.g., 420) is a color image having a plurality of color components, generation of the fused bitmap (e.g., 631) may involve generating respective intermediate fused bitmaps corresponding to the color components and then combining the intermediate fused bitmaps. More specifically, FIG. 7 shows a color image having a red component, a green component, and a blue component. In such an example, the computing system 101 may be configured to generate at least a first set of bitmaps (from M1_Red to M5_Red) corresponding to a first color component (e.g., red), and a second set of bitmaps (from M1_Green to M5_Green) corresponding to a second color component (e.g., green). In the example of FIG. 7, the computing system 101 may further generate a third set of bitmaps (from M1_Blue to M5_Blue) corresponding to a third color component (e.g., blue). In this embodiment, respective intermediate fused bitmaps, such as Fused_Red, Fused_Green, and Fused_Blue, may be generated from each of the three sets of bitmaps. The three intermediate fused bitmaps may be combined into a single fused bitmap, such as bitmap 631 of FIG. 6.

[0082] As described above, the classification in step 308 may be based on a standard deviation bitmap (e.g., 533) that may represent intensity variations across at least one image portion of the image. In an embodiment, at least one criterion for classifying an image portion as having texture may be based on intensity variations across the fused bitmap (e.g., 631). The variation across the fused bitmap may be quantified, for example, by a standard deviation value of a local region within the fused bitmap. For example, if the maximum value, minimum value, or representative value of such local standard deviation values is equal to or greater than a defined standard deviation threshold, the computing system 101 may classify at least one image portion as having texture.

[0083] In an embodiment, step 308 may involve generating a texture bitmap based on the fused bitmap. In such embodiments, at least one criterion for classifying an image portion as having texture may be based on the texture bitmap. FIG. 6 depicts a fused bitmap 631 that is being converted into a texture bitmap 641. In an embodiment, the texture bitmap may be for identifying which one or more regions of a corresponding image portion (e.g., 421) have a sufficient level of visual texture or for indicating that there are no regions in the image portion (e.g., 421) having a sufficient level of visual texture. More specifically, the texture bitmap may have a texture identification region and / or a no-texture identification region. A texture identification region, such as region 643 in texture bitmap 641, may have pixel values indicating that the corresponding region of the image portion, which may be referred to as a textured region, has at least a defined texture level. A no-texture identification region, such as region 645 in texture bitmap 641, may have pixel values indicating that the corresponding region of the image portion, which may be referred to as a no-texture region, does not have a defined texture level. The texture region in the image portion (e.g., 421) may occupy the same position (e.g., the same coordinates) in the texture bitmap 641 as that occupied by the texture identification region 643. Similarly, the no-texture region in the image portion may occupy the same position in the texture bitmap 641 as that occupied by the no-texture identification region 645. Thus, the texture bitmap 641 may be for identifying how much of the image portion (if any) has a sufficient level of visual texture and how much of the image portion (if any) lacks a sufficient level of visual texture.

[0084] In an embodiment, the computing system 101 may be configured to generate a texture bitmap (e.g., 641) by comparing pixels of a fused bitmap (e.g., 631) to a defined texture level threshold, such as a defined pixel value threshold. In such an example, for each pixel coordinate of the fused bitmap (e.g., 631), the computing system 101 may determine whether the pixel value of the fused bitmap at that pixel coordinate is equal to or exceeds the defined pixel value threshold. If the pixel value of the fused bitmap at that pixel coordinate is equal to or exceeds the defined pixel value threshold, the computing system 101 may assign, for example, a non-zero value to the same pixel coordinate in the texture bitmap (e.g., 641). As an example, the pixel coordinates assigned a non-zero value may be pixel coordinates within the texture identification region 643. The above considerations involve the assignment of non-zero values, but any value associated with indicating a sufficient level of texture may be assigned. If the pixel value of the fused bitmap (e.g., 631) at that pixel coordinate is less than the defined pixel value threshold, the computing system 101 may assign, for example, a zero value to the same pixel coordinate in the texture bitmap. As an example, the pixel coordinates assigned a zero value may be pixel coordinates within the no-texture identification region 645. The above considerations involve the assignment of zero values, but any value associated with indicating an insufficient level of texture may be assigned.

[0085] In an embodiment, the texture bitmap may be a binary mask in which all pixels in the texture bitmap can have only one of two pixel values, such as either 0 or 1. For example, all pixels in the texture identification region 643 of the texture bitmap 641 may have a pixel value of 1, while all pixels in the textureless identification region 645 may have a value of 0. In this example, pixels having a pixel value of 1 in the texture bitmap may indicate that the corresponding region of the image portion (e.g., 421) is a textured region, while pixels having a pixel value of 0 in the texture bitmap 641 may indicate that the corresponding region of the image portion (e.g., 421) is a textureless region.

[0086] In an embodiment, at least one criterion for classifying an image portion (e.g., 421) as having texture may be based on the size (e.g., total area) of one or more texture identification regions (if any) in the texture bitmap (e.g., 641), or on the size of one or more textureless identification regions (if any) in the texture bitmap (e.g., 641). The criterion may also be based on the size of one or more textured regions (if any) of the image portion (e.g., 421), or on the size of one or more textureless regions (if any) of the image portion. The size of one or more texture identification regions (if any) may be equal to or substantially equal to the size of one or more textured regions (if any), while the size of one or more textureless identification regions (if any) may be equal to or substantially equal to the size of one or more textureless regions (if any).

[0087] As an example of the above criteria, the computing system 101 may determine the total area with texture indicated by the texture bitmap, and based on the total area with texture, classify an image portion (e.g., 421) as having texture or not having texture. The total area with texture may indicate the total area of all texture identification regions (e.g., 643) in the texture bitmap (e.g., 641), or the total area of all corresponding regions with texture in the image portion (e.g., 421). If there are no texture identification regions in the texture bitmap (e.g., 641), or no regions with texture in the image portion (e.g., 421), the total area with texture may be zero. In some cases, the computing system 101 may classify an image portion (e.g., 421) as having texture if the total area with texture is equal to or greater than a defined area threshold, and classify the image portion (e.g., 421) as not having texture if the total area with texture is less than the defined area threshold.

[0088] In an embodiment, at least one criterion for classifying an image portion as having texture or not having texture may be a ratio P, which, if present, is the ratio of the image portion (e.g., 421) occupied by one or more regions with texture, or, if present, the ratio of the texture bitmap (e.g., 641) occupied by one or more texture identification regions (e.g., 643). texture If there are no regions with texture in the image portion, or no texture identification regions in the corresponding texture bitmap, the ratio P texture may be zero. In an embodiment, at least one criterion may be a ratio P, which, if present, is the ratio of the image portion (e.g., 421) occupied by one or more regions without texture, or, if present, the ratio of the texture bitmap (e.g., 641) occupied by one or more regions without texture identification (e.g., 643). textureless It may also be based on this.

[0089] In an embodiment, at least one criterion for classifying an image portion as having texture or not having texture is the ratio Ptexture (In this example, it may be the first ratio) and the ratio P textureless (In this example, it may be the second ratio). For example, such an embodiment is based on the ratio P texture / P textureless exceeding a comparison threshold T1 (e.g., 5) between the defined texture and no texture, may involve classifying at least one image portion (e.g., 421) as having texture.

[0090] In an embodiment, at least one criterion for classifying an image portion (e.g., 421) as having texture or no texture is the ratio of P texture in the image portion (e.g., 421) or in the image (e.g., 420) received in step 302 image to the total number of pixels Num textureless and / or the ratio of P image to Num texture / Num image is greater than a comparison threshold T2 (e.g., 0.9) between the defined texture and the image size, and / or the ratio of P textureless / Num image is less than a comparison threshold T3 (e.g., 0.1) between no defined texture and the image size, at least the image portion (e.g., 421) may be classified as having texture.

[0091] In an embodiment, the computing system 101 may combine some or all of the above criteria involved in classifying an image portion as having texture or no texture. In some cases, the computing system 101 may be configured to perform step 308 by classifying the image portion (e.g., 421) as having texture if any one of the above criteria is met, and classifying the image portion as having no texture if none of the above criteria are met.

[0092] For example, as part of evaluating a first criterion, the computing system 101 may determine whether the number of descriptors in a descriptor bitmap (e.g., 513) is greater than a defined descriptor quantity threshold. If this first criterion is met, the computing system 101 may classify the image portion (e.g., 421) as having texture. If the first criterion is not met, the computing system 101 may evaluate a second criterion by determining whether P texture / P textureless >T1. If the second criterion is met, the computing system 101 may classify the image portion (e.g., 421) as having texture. If the second criterion is not met, the computing system 101 may evaluate a third criterion by determining whether P textureless / Num image >T2 and / or P textureless / Num image <T3. If the third criterion is met, the computing system 101 may classify the image portion (e.g., 421) as having texture. If the third criterion is not met, the computing system 101 may evaluate a fourth criterion by determining whether the maximum, minimum, or average of the standard deviation values indicated by a standard deviation bitmap (e.g., 533) or a fusion bitmap (e.g., 631) is greater than a defined standard deviation threshold. If the fourth criterion is met, the computing system may classify the image portion (e.g., 421) as having texture. If none of the above criteria are met, the computing system 101 may classify the image portion (e.g., 421) as having no texture.

[0093] In an embodiment, steps 306 and 308 may be repeated for one or more other image portions of the image received at step 302. For example, the received image (e.g., 420) may represent a plurality of objects such as objects 401-404 in FIG. 4A. In some situations, more than one template may be generated based on the plurality of objects. As an example, a first template may be generated based on image portion 421, which describes the appearance of object 401, as discussed above. In this embodiment, a second template may be generated based on a second image portion 422, while a third template may be generated based on a third image portion 423, and image portions 422 and 423 are depicted in FIGS. 8A-8C. Image portion 422 may represent object 402, while image portion 423 may represent object 403. By the computing system 101 of this example, image portions 422 and 423 are extracted from image 420, and steps 306 and 308 may be performed on those image portions to generate the second and third templates respectively based on those image portions 422, 423. In one example, image portion 422 may be classified as a textureless template. In some implementations, image portion 423 may also be classified as a textureless template. Although image portion 423 may display one or more edges of a piece of tape, the feature bitmap, variance bitmap, and fusion bitmap generated from only one or more edges may be insufficient in this example to produce a textured classification.

[0094] Returning to FIG. 3, method 300 may include step 310, where the processing circuit 110 of computing system 101 executes an operation plan for robot interaction with one or more objects (e.g., 401-404 in FIG. 4A) based on whether at least one image portion (e.g., 421) is classified as having texture or being textureless. In an embodiment, step 308 may be performed by the image classification module 204 and / or the operation plan module 208 of FIG. 2C.

[0095] In an embodiment, step 310 may involve performing object recognition on one or more of the objects, such as objects 401-404, represented by the image 420 and within the device field of view (e.g., 443) of the image capture device (e.g., 441). For example, as discussed above, the image portion 421 representing object 401 may be used as a template or to generate a template, and the object recognition may involve determining whether the remaining objects 402-404 within the device field of view 443 match the template. As an example, the computing system 101 may be configured to determine whether a portion of the image 420 representing object 402, 403, or 404 matches the template, where the template is generated based on the appearance of object 401. In some cases, the object recognition may be based on whether the template is classified as a textured template or a non-textured template. For example, the classification of the template may affect where the template is stored and / or for how long the template is stored. The implementation of object recognition based on non-textured templates or textured templates is discussed in more detail in U.S. Patent Application No. ______ (Attorney Docket No. MJ0054-US / 0077-0012US1), filed on the same date as this specification and entitled "METHOD AND COMPUTING SYSTEM FOR OBJECT RECOGNITION OR OBJECT REGISTRATION BASED ON IMAGE CLASSIFICATION", the entire content of which is incorporated herein by reference. As described above, object recognition may generate information about, for example, the size of the object, which may be used to plan an interaction between the object (e.g., 404) and the robot. In an embodiment, step 310 may be omitted. For example, such embodiments may include steps 302, 306, 308 and may include a method that stops upon completion of step 308.

[0096] In an embodiment, the computing system 101 may be configured to determine the confidence of object recognition, and the determination may be based on whether the template is with or without texture. For example, if the appearance of an object (e.g., 403) matches only the textureless template, such a match may be assigned a relatively low confidence. If the appearance of an object (e.g., 404) matches the textured template, such a match may be assigned a relatively high confidence. In some cases, the computing system 101 may be configured to perform additional object recognition operations, such as operations based on another technique or additional information, to attempt to improve the robustness of object recognition. In some cases, the computing system 101 may perform an action plan based on the confidence. For example, if the confidence is relatively low, the computing system 101 may be configured to limit the speed of the robot when the robot attempts to pick up an object or interact with the object in another way so that the interaction of the robot (e.g., 461) can proceed with a higher level of attention.

[0097] Additional considerations regarding various embodiments

[0098] Embodiment 1 relates to a method of image classification. The method can be performed, for example, by a computing system that executes instructions on a non-transitory computer-readable medium. The method of this embodiment includes receiving an image by the computing system, where the computing system is configured to communicate with an image capture device, and the image is generated by the image capture device and is for representing one or more objects within the field of view of the image capture device. The method further includes generating, by the computing system, one or more bitmaps based on at least one image portion of the image, where the one or more bitmaps and the at least one image portion are associated with a first object among the one or more objects, and the one or more bitmaps describe whether one or more visual features for feature detection are present in the at least one image portion or describe whether there are intensity variations across the at least one image portion. Additionally, the method includes determining, based on the one or more bitmaps, whether to classify the at least one image portion as either with texture or without texture, and executing an operation plan for robotic interaction with the one or more objects based on whether the at least one image portion is classified as with texture or without texture.

[0099] Embodiment 2 includes the method of Embodiment 1. In this embodiment, the one or more bitmaps include a descriptor bitmap for identifying one or more regions of at least one of the at least one image portion, where the one or more descriptors indicate whether the one or more descriptors are present in the at least one image portion or include the one or more respective descriptors detected from the at least one image portion. Determining whether to classify the at least one image portion as either with texture or without texture is based on whether the total number of descriptors identified by the descriptor bitmap exceeds a defined descriptor quantity threshold.

[0100] Embodiment 3 includes the method of Embodiment 1 or 2. In this embodiment, one or more bitmaps include a plurality of bitmaps having a first bitmap and a second bitmap. The first bitmap is generated based on at least one image portion and describes whether one or more visual features of a first feature type are present in the at least one image portion. Further, in this embodiment, the second bitmap is generated based on at least one image portion and describes whether one or more visual features of a second feature type are present in the at least one image portion, and determining whether to classify the at least one image portion as either with texture or without texture includes generating a fused bitmap by combining the plurality of bitmaps, and the at least one image portion is classified as either with texture or without texture based on the fused bitmap.

[0101] Embodiment 4 includes the method of Embodiment 3. In this embodiment, the first bitmap is a descriptor bitmap for identifying one or more regions of at least one of the at least one image portion, including one or more respective descriptors detected from the at least one image portion, or for indicating that the descriptor is not identified within the at least one image portion, and the second bitmap is an edge bitmap for identifying one or more regions of at least one of the at least one image portion, including one or more respective edges detected from the at least one image portion, or for indicating that the edge is not detected within the at least one image portion.

[0102] Embodiment 5 includes the method of Embodiment 4. In this embodiment, the plurality of bitmaps includes a third bitmap that, for each pixel of the at least one image portion, is a standard deviation bitmap indicating the standard deviation between pixel intensity values around the pixel.

[0103] Embodiment 6 includes any one of the methods of Embodiments 3 to 5. In this embodiment, determining whether to classify at least one image portion as having texture or no texture includes, by a computing system, converting a fused bitmap into a texture bitmap. Further, in this embodiment, the texture bitmap is for identifying one or more textured regions of at least one image portion or for indicating that there are no textured regions in at least one image portion, and the texture bitmap is for further identifying one or more textureless regions of at least one image portion or for indicating that there are no textureless regions in at least one image portion, the one or more textured regions being one or more regions of at least one image portion having at least a defined texture level, the one or more textureless regions being one or more regions of at least one image portion having a lower texture level than the defined texture level, and determining whether to classify at least one image portion as having texture or no texture is based on the texture bitmap.

[0104] Embodiment 7 includes the method of Embodiment 6. In this embodiment, determining whether to classify at least one image portion as having texture or no texture is based on at least one of the total textured areas indicated by the texture bitmap, the total textured area being the total area of one or more textured regions or zero if the texture bitmap indicates that there are no textured regions in at least one image portion.

[0105] Embodiment 8 includes any one of the methods of Embodiments 3 to 7. In this embodiment, determining whether to classify at least one image portion as having texture or no texture is based on the presence or absence of fluctuations in pixel intensity values across the fused bitmap or on the amount of fluctuations in pixel intensity values across the fused bitmap.

[0106] Embodiment 9 includes any one of the methods of Embodiments 2 to 8. In this embodiment, determining whether to classify at least one image portion as having texture or no texture includes: a) classifying at least one image portion as having texture when the number of descriptors identified by the descriptor bitmap is greater than a defined descriptor quantity threshold; b) classifying at least one image portion as having texture when the ratio of the first ratio to the second ratio exceeds a defined texture-to-no-texture comparison threshold, where the first ratio is the ratio of at least one image portion occupied by one or more textured regions or zero if there are no textured regions in at least one image portion, and the second ratio is the ratio of at least one image portion occupied by one or more textureless regions; c) classifying at least one image portion as having texture when the ratio of the first ratio to the size of at least one image portion is greater than a defined texture-to-image size comparison threshold, or when the ratio of the second ratio to the size of at least one image portion is less than a defined textureless-to-image size comparison threshold; or d) classifying at least one image portion as having texture when the maximum or minimum value of the standard deviation of the local region of each pixel of the fusion bitmap is greater than a defined standard deviation threshold, including at least one of the above.

[0107] Embodiment 10 includes any one of the methods of Embodiments 1 to 9. In this embodiment, the method further includes generating an additional bitmap that describes the influence of the illumination state in which the image was generated on at least one image portion.

[0108] Embodiment 11 includes the method of Embodiment 10. In this embodiment, the additional bitmap is a highlight bitmap that identifies, as a result of the illumination state, one or more regions exceeding a defined luminance threshold within at least one image portion, or a shadow bitmap that identifies one or more regions within a shadow in at least one image portion, and includes at least one of them.

[0109] Embodiment 12 includes the method of any one of Embodiments 3 to 11. In this embodiment, generating the fused bitmap includes determining bitmap pixel values that describe the texture level across at least one image portion based at least on the first bitmap and the second bitmap, and reducing a subset of the determined bitmap pixel values based on the highlight bitmap or the shadow bitmap, wherein the subset of the bitmap pixel values to be reduced corresponds to one or more regions of at least one image portion that are identified by the highlight bitmap when exceeding a defined luminance threshold, or are identified by the shadow bitmap as being in a shadow.

[0110] Embodiment 13 includes the method of any one of Embodiments 3 to 12. In this embodiment, generating the fused bitmap is based at least on the weighted sum of the first bitmap and the second bitmap, and the weighted sum of the highlight bitmap and the shadow bitmap.

[0111] Embodiment 14 includes the method of any one of Embodiments 3 to 13. In this embodiment, the image received by the computing system is a color image including a plurality of color components, the first bitmap and the second bitmap belong to a first set of bitmaps associated with a first color component among the plurality of color components, the method includes generating a second set of bitmaps associated with a second color component among the plurality of color components, and the fused bitmap is generated based at least on the first set of bitmaps and the second set of bitmaps.

[0112] Embodiment 15 includes the method of Embodiment 14. In this embodiment, the method further includes generating a first intermediate fusion bitmap by combining a first set of bitmaps, wherein the first intermediate fusion bitmap is associated with a first color component, and generating a second intermediate fusion bitmap by combining a second set of bitmaps, wherein the second intermediate fusion bitmap is associated with a second color component, and the fusion bitmap is generated by combining at least the first intermediate fusion bitmap and the second intermediate fusion bitmap.

[0113] Embodiment 16 includes the method of any one of Embodiments 1 to 15. In this embodiment, the method further includes applying a smoothing operation to the image to produce an updated image before one or more bitmaps are generated, and at least one image from which one or more bitmaps are generated is extracted from the updated image.

[0114] It will be apparent to those skilled in the relevant art that other suitable modifications and adaptations to the methods and uses described herein can be made without departing from the scope of any of the embodiments. The embodiments described above are illustrative examples and should not be construed as limiting the present invention to these specific embodiments. It should be understood that the various embodiments disclosed herein may be combined in combinations different from those specifically presented in the description and the accompanying drawings. By way of example, it should also be understood that any particular act or event of any of the processes or methods described herein may be performed in a different order, may be added, incorporated, or completely omitted (for example, all acts or events described may not be necessary to carry out the method or process). In addition, although certain features of the embodiments herein are described as being performed by a single component, module, or unit for clarity, it should be understood that the features and functions described herein may be performed by any combination of components, modules, or units. Accordingly, various changes and modifications can be made by those skilled in the art without departing from the spirit or scope of the invention as defined in the appended claims.

Claims

1. Receiving, by a computing system, an image representing one or more objects; Generating, by the computing system, one or more masks, wherein each of the one or more masks describes whether one or more visual features are present in at least one image portion of the image or whether there are intensity variations across the at least one image portion; Generating, from the one or more masks, a texture mask, wherein the texture mask is for identifying one or more textured regions and / or one or more textureless regions of the at least one image portion; Determining, by the computing system, a texture classification of the at least one image portion based on the generated texture mask; Executing, based on the texture classification of the at least one image portion, an operation plan for robotic interaction with the one or more objects; A method for image classification, comprising the steps above.

2. The one or more masks include a descriptor mask, and the descriptor mask is for indicating whether one or more descriptors are present in the at least one image portion or for identifying one or more regions of the at least one image portion that include respective ones of the one or more descriptors detected from the at least one image portion. The step of determining the texture classification of the at least one image portion is based on whether the total number of descriptors identified by the descriptor mask exceeds a defined descriptor quantity threshold, according to the method of Claim 1.

3. The one or more masks include a plurality of masks having a first mask and a second mask. The first mask is generated based on the at least one image portion and describes whether one or more visual features of a first feature type are present in the at least one image portion. The second mask is generated based on the at least one image portion and describes whether one or more visual features of a second feature type are present in the at least one image portion. The step of determining the texture classification of the at least one image portion includes generating a fused mask by combining the plurality of masks. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ The method according to claim 1, wherein at least one of the image portions is classified with or without texture based on the fusion mask.

4. The first mask is a descriptor mask for identifying one or more regions of at least one of the image portions, including one or more respective descriptors detected from the at least one image portion, or a descriptor mask for indicating that no descriptor is detected in the at least one image portion. The second mask is an edge mask for identifying one or more regions of at least one of the image portions, including one or more respective edges detected from the at least one image portion, or an edge mask for indicating that no edge is detected in the at least one image portion, according to the method of claim 3.

5. The plurality of masks has a third mask, wherein the third mask is a standard deviation mask for indicating, for each pixel of the at least one image portion, the standard deviation between the pixel intensity values around the pixel, according to the method of claim 4.

6. Determining the texture classification of the at least one image portion includes converting, by the computing system, the fusion mask into a texture mask, wherein the texture mask is for identifying one or more textured regions of the at least one image portion or for indicating that there are no textured regions in the at least one image portion, wherein the texture mask is further for identifying one or more textureless regions of the at least one image portion or for indicating that there are no textureless regions in the at least one image portion, wherein the one or more textured regions are at least one or more regions of the at least one image portion having a defined texture level, and the one or more textureless regions are at least one or more regions of the at least one image portion having a lower texture level than the defined texture level, Determining the texture classification of the at least one image portion is based on the texture mask, according to the method of claim 4.

7. Determining the texture classification of the at least one image portion is based on at least one of the total areas with texture indicated by the texture mask, The total area with texture is the total area of the one or more areas with texture, or zero when the texture mask indicates that there are no areas with texture in the at least one image portion, the method of claim 6. **Claim 8** Determining the texture classification of the at least one image portion is based on the presence or absence of fluctuations in pixel intensity values passing through the fusion mask, or based on the amount of fluctuations in pixel intensity values passing through the fusion mask, the method of claim 6. **Claim 9** Determining the texture classification of the at least one image portion is a) classifying the at least one image portion as having texture when the number of descriptors identified by the descriptor mask is greater than a defined quantity threshold of descriptors, b) classifying the at least one image portion as having texture when the ratio of a first ratio to a second ratio exceeds a defined comparison threshold between texture and no texture, where the first ratio is the ratio of the at least one image portion occupied by the one or more areas with texture, or zero if there are no areas with texture in the at least one image portion, and the second ratio is the ratio of the at least one image portion occupied by the one or more areas without texture, c) classifying the at least one image portion as having texture when the ratio of the first ratio to the size of the at least one image portion is greater than a defined comparison threshold between texture and image size, or when the ratio of the second ratio to the size of the at least one image portion is less than a defined comparison threshold between no texture and image size, or d) classifying the at least one image portion as having texture when the maximum or minimum value of the standard deviation with respect to the local region of each pixel of the fusion mask is greater than a defined standard deviation threshold, including at least one of, the method of claim 6. **Claim 10** further comprising generating an additional mask, The method according to claim 3, wherein the additional mask describes the influence on the at least one image portion from the illumination state in which the image was generated. **Claim 11** The additional mask is a highlight mask that identifies, within the at least one image portion, one or more regions that exceed a defined luminance threshold as a result of the illumination state, or includes at least one of a shadow mask that identifies, within the at least one image portion, one or more regions that are in shadow, the method according to claim 10. **Claim 12** Generating the fusion mask includes determining mask pixel values that describe a texture level across the at least one image portion based at least on the first mask and the second mask, and reducing a subset of the determined mask pixel values based on the highlight mask or the shadow mask, wherein the subset of mask pixel values to be reduced is identified by the highlight mask when it exceeds the defined luminance threshold, or is identified by the shadow mask as being in shadow, corresponding to one or more regions of the at least one image portion, the method according to claim 11. **Claim 13** The method according to claim 11, wherein generating the fusion mask is based at least on a weighted sum of the first mask and the second mask, and a weighted sum of the highlight mask and the shadow mask. **Claim 14** The image received by the computing system is a color image including a plurality of color components, the first mask and the second mask belong to a first set of masks associated with a first color component of the plurality of color components, the method includes generating a second set of masks associated with a second color component of the plurality of color components, the fusion mask is generated based at least on the first set of masks and the second set of masks, the method according to claim 3. **Claim 15** further includes generating a first intermediate fusion mask associated with the first color component by combining the first set of masks, and generating a second intermediate fusion mask associated with the second color component by combining the second set of masks. The method according to claim 14, wherein the fusion mask is generated by combining at least the first intermediate fusion mask and the second intermediate fusion mask.

16. Further comprising applying a smoothing operation to the image to produce an updated image before the one or more masks are generated, The method according to claim 1, wherein the at least one image portion for which the one or more masks are generated is extracted from the updated image.

17. A non-transitory computer-readable medium, At least one processing circuit, When the non-transitory computer-readable medium stores an image representing one or more objects, the at least one processing circuit Receiving the image, Generating one or more masks, each of the one or more masks indicating whether one or more visual features are present in at least one image portion of the image or describing whether there are intensity variations across the at least one image portion, Generating a texture mask from the one or more masks, the texture mask being for identifying one or more textured regions and / or one or more textureless regions of the at least one image portion, Determining a texture classification of the at least one image portion based on the generated texture mask, Executing an action plan for robot interaction with the one or more objects based on the texture classification of the at least one image portion, A computing system for image classification, configured to perform the above.

18. The one or more masks include a descriptor mask, The descriptor mask Is for indicating whether one or more descriptors are present in the at least one image portion or Is for identifying one or more regions of the at least one image portion, each including one or more descriptors detected from the at least one image portion, And The at least one processing circuit The computing system according to claim 17, configured to determine the texture classification of the at least one image portion based on whether the total number of descriptors identified by the descriptor mask exceeds a defined descriptor quantity threshold.

19. The one or more masks include a plurality of masks having a first mask and a second mask, the first mask is generated based on the at least one image portion and describes whether one or more visual features of a first feature type are present in the at least one image portion, the second mask is generated based on the at least one image portion and describes whether one or more visual features of a second feature type are present in the at least one image portion, the at least one processing circuit is configured to determine the texture classification of the at least one image portion by generating a fused mask that combines the plurality of masks, the at least one image portion is classified as having texture or not having texture based on the fused mask, the computing system according to claim 17.

Citation Information

Patent Citations

  • Takeout device

    JP1988017735A

  • Image inspection and recognition method, and reference data generating method used for same and device for them

    JP1996069533A

  • Information processor, method, and robot system

    JP2019063984A

  • Information processor, information processing method, program, and system

    JP2019125345A

  • Controller, robot, robot system, and method for recognizing object

    JP2019158427A