METHOD, SYSTEM AND DEVICE FOR DETERMINING A BEARING STRUCTURE DEPTH

The method and system improve the accuracy of determining support structure depth by using a mobile automation device with sensors and confidence masks to filter and weight depth measurements, addressing the challenges of complex inventory environments.

DE112019004975B4Active Publication Date: 2026-03-19ZEBRA TECHNOLOGIES CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2019-09-05
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing inventory management systems face challenges in accurately determining the depth of support structures, such as shelves, due to varying object characteristics, frequent changes in placement and quantity, and varying imaging conditions, which affect the accuracy of capturing information about objects in complex and volatile environments.

Method used

A method and system that utilize a mobile automation device equipped with sensors to capture point clouds and images, apply a confidence mask to identify unobscured depth measurements, and determine the support structure depth based on these measurements and confidence levels, using a server to process the data and generate a shelf depth.

Benefits of technology

Enhances the accuracy of determining support structure depth by filtering out obscured points and weighting measurements by confidence levels, resulting in precise shelf depth determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for determining the support structure depth of a support structure with a front and a back separated by the support structure depth comprises: obtaining a point cloud of the support structure and a mask indicating, for a plurality of sections of an image of the support structure acquired from a capture position, respective confidence levels that the sections represent the back of the support structure; selecting an initial set of points from the point cloud that are within a field of view radiating from the capture position; selecting an unobscured subset of depth measurements from the initial set of points, wherein the depth measurements in the unobscured subset correspond to respective image coordinates; retrieving a confidence level for each of the depth measurements in the unobscured subset from the mask;and determining the carrier structure depth based on the depth measurements in the unhidden subset and the retrieved confidence levels.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Environments where inventories of objects are managed, such as products for purchase in a retail setting, can be complex and volatile. For example, a given environment may contain a variety of objects with differing characteristics (size, shape, price, and the like). Furthermore, the placement and quantity of objects within the environment can change frequently. In addition, imaging conditions, such as lighting, can vary both over time and at different locations within the environment. These factors can reduce the accuracy with which information about the objects in the environment can be captured.

[0002] US 9,996,818 B1 describes a system for counting stacked items using image analysis. An image of a storage location containing stacked items is captured and processed to determine the number of items stacked at that location. In some cases, the object closest to the camera capturing the image may be the only object visible in the image. Using image analysis, the object's distance from the camera and the storage location's shelf can be determined. With this information and known measurements of the object, the number of items stacked at that location can be calculated.

[0003] US 2019 / 0073550A1 describes a sensor calibration target configured for sensor calibration relative to a common reference frame. The sensor calibration target includes a first surface at a first predefined depth, bearing a first set of markers at corresponding first heights and with corresponding first predefined displacements, each of the first markers encoding a corresponding first height. The sensor calibration target further includes a second surface at a second predefined depth, bearing a second set of markers at corresponding second heights and with corresponding second predefined displacements, each of the second markers encoding a corresponding second height.

[0004] The document ZENG, Andy, et al. Multi-view self-supervised deep learning for 6d pose estimation in the amazon picking challenge. In: 2017 IEEE international conference on robotics and automation (ICRA). IEEE, 2017. pp. 1386-1383, describes segmenting and labeling multiple views of a scene using a neural network and fitting scanned 3D object models to the resulting segmentation to obtain a 6D object pose. BRIEF DESCRIPTION OF THE DIFFERENT VIEWS OF THE DRAWINGS

[0005] The accompanying figures, in which identical reference numerals denote identical or functionally similar elements in the individual views, are incorporated into the disclosure together with the following detailed description and form an integral part of the disclosure and serve to further illustrate embodiments of concepts comprising the claimed invention described herein and to explain various principles and advantages of these embodiments. Fig. Figure 1 is a schematic representation of a mobile automation system. Fig. Figure 2A shows a mobile automation device in the system of Fig. 1. Fig. 2B is a block diagram of certain internal hardware components of the mobile automation device in the system of Fig. 1. Fig. Figure 3 is a flowchart of a procedure for determining a support structure depth. Fig. 4A is a diagram of a point cloud and a shelf plane, which is shown in block 305 of the procedure by Fig. 3 will be received. Fig. 4B is a diagram of example images produced by the system's device. Fig. 1 recorded and at block 310 of the procedure from Fig. 3 will be received. Fig. 5A is a diagram that is one of the images of Fig. 4B presents in more detail. Fig. 5B is a diagram showing a sample back side of a shelf mask, featuring the image of Fig. 5A corresponds. Fig. 6A is a flowchart of a procedure for carrying out block 315 of the procedure of Fig. 3. Fig. 6B is a diagram illustrating the execution of the procedure of Fig. 6A in conjunction with the point cloud of Fig. 4A is shown. Fig. 7A is a flowchart of a procedure for carrying out block 320 of the procedure of Fig. 3. Fig. 7B is a diagram illustrating the execution of the procedure of Fig. 7A in conjunction with the image of Fig. 5A is shown. Fig. 8A and Fig. 8B are diagrams that illustrate an exemplary implementation of Block 325 of the procedure by Fig. Show 3. Fig. 8C is a diagram showing a further embodiment of block 325 of the method of Fig. 3 shows. Fig. Figure 9 is a diagram showing a support structure depth achieved by carrying out the procedure of Fig. 3 is determined.

[0006] Experts will recognize that elements in the figures are shown for the sake of simplicity and clarity and are not necessarily drawn to scale. For example, the dimensions of some of the elements in the figures may be exaggerated relative to other elements to improve the understanding of embodiments of the present invention.

[0007] Where appropriate, the apparatus and process components have been represented by conventional symbols in the drawings, which show only those specific details relevant to understanding the embodiments of the present invention, so as not to obscure the disclosure with details that are readily apparent to those skilled in the field who refer to the present description. DETAILED DESCRIPTION

[0008] The examples disclosed herein relate to a method for determining the support structure depth of a support structure having a front and a back separated by the support structure depth, wherein the method comprises: obtaining (i) a point cloud of the support structure and (ii) a mask indicating, for a plurality of sections of an image of the support structure acquired from a capture position, respective confidence levels that the sections represent the back of the support structure; selecting an initial set of points from the point cloud that are within a field of view radiating from the capture position; selecting an unobscured subset of depth measurements from the initial set of points, wherein the depth measurements in the unobscured subset correspond to respective image coordinates;Retrieving a confidence level for each of the depth measurements in the unhidden subset from the mask; and determining the carrier structure depth based on the depth measurements in the unhidden subset and the retrieved confidence levels.

[0009] Further examples disclosed herein relate to a computer device for determining a support structure depth of a support structure having a front and a back separated by the support structure depth, the computer device comprising: a memory that stores (i) a point cloud of the support structure and (ii) a mask that indicates, for a plurality of sections of an image of the support structure acquired from a capture position, respective confidence levels that the sections represent the back of the support structure; an imaging controller connected to the memory and configured to: select from the point cloud an initial set of points located within a field of view extending from the capture position;to select an unhidden subset of depth measurements from the initial set of points, where the depth measurements in the unhidden subset correspond to the respective image coordinates; to retrieve a confidence level for each of the depth measurements in the unhidden subset from the mask; and to determine the support structure depth based on the depth measurements in the unhidden subset and the retrieved confidence levels.

[0010] Further examples disclosed herein are directed to a computer-readable medium that stores computer-readable instructions that can be executed by a server's processor, wherein the execution of the computer-readable instructions causes the server to: (i) obtain a point cloud of the support structure and (ii) a mask indicating, for a plurality of sections of an image of the support structure acquired from a capture position, respective confidence levels that the sections represent the back side of the support structure; select an initial set of points from the point cloud that are within a field of view radiating from the capture position; select an unobscured subset of depth measurements from the initial set of points, wherein the depth measurements in the unobscured subset correspond to respective image coordinates;Retrieving a confidence level for each of the depth measurements in the unhidden subset from the mask; and determining the carrier structure depth based on the depth measurements in the unhidden subset and the retrieved confidence levels.

[0011] Fig. Figure 1 shows a mobile automation system 100 according to the teachings of this disclosure. The system 100 is shown to be used in a retail environment, but in other embodiments it can be used in a variety of other environments, including warehouses, hospitals, and the like. The system 100 comprises a server 101 that communicates with at least one mobile automation device 103 (here also simply referred to as device 103) and at least one client computer device 105 via communication links 107, which in this example include wireless connections. In this example, the connections 107 are provided by a wireless local area network (WLAN) provided within the retail environment by one or more access points (not shown).In other examples, the server 101, the client device 105, or both are located outside the retail environment, and the connections 107 therefore include wide-area networks such as the internet, mobile networks, and the like. In this example, the system 100 also includes a dock 108 for the device 103. The dock 108 communicates with the server 101 via a connection 109, which in this example is a wired connection. In other examples, however, connection 109 is a wireless connection.

[0012] The client computer device 105 is in Fig. 1 is represented as a mobile computing device, such as a tablet, smartphone, or the like. In other examples, the client device 105 is implemented as a different type of computing device, such as a desktop computer, a laptop computer, another server, a kiosk, a monitor, and the like. The system 100 can comprise a variety of client devices 105 that communicate with the server 101 via appropriate connections 107.

[0013] In the example shown, System 100 is used in a retail environment that includes a variety of support structures such as shelf modules 110-1, 110-2, 110-3, etc. (collectively referred to as shelves 110 and generally as a shelf 110 – this nomenclature is also used for other elements discussed here). Other examples may include additional types of support structures, such as pegboards. Each shelf module 110 supports a variety of products 112. Each shelf module 110 includes a shelf back panel 116-1, 116-2, 116-3 and a support surface (e.g., support surface 117-3 as shown in [reference missing]). Fig. 1 shown), which extends from the back of the shelf 116 to a shelf edge 118-1, 118-2, 118-3.

[0014] The shelf modules 110 are typically arranged in a multitude of aisles, each containing a multitude of modules 110 aligned end to end. In such arrangements, the shelf edges 118 face into the aisles through which customers in the retail environment, as well as the device 103, can move. As from Fig. As can be seen from Figure 1, the term "shelf edge" 118 used here, which can also be described as the edge of a support surface (e.g., the support surfaces 117), refers to a surface that is bounded by adjacent surfaces with different angles of inclination. In the Fig. In the example shown, the shelf edge 118-3 is at an angle of approximately ninety degrees relative to each of the support surfaces 117-3 and the underside (not shown) of the support surface 117-3. In other examples, the angles between the shelf edge 118-3 and the adjacent surfaces, such as the support surface 117-3, are more or less than ninety degrees. The shelf edges 118 define a front of the shelves 110, which is separated from the shelf backs 116 by a shelf depth. A common reference frame 102 is in Fig. Figure 1 is shown. In the present example, the shelf depth is defined in the Y dimension of the reference frame 102, while the shelf backs 116 and the shelf edges 118 are shown parallel to the XZ plane.

[0015] The device 103 is used in the retail environment and communicates with the server 101 (e.g. via connection 107) to operate autonomously or semi-autonomously along a length 119 (in Fig. The device 103 (shown as parallel to the X-axis of the reference frame 102) navigates at least part of the shelves 110. The device 103, autonomous or in conjunction with the server 101, is configured to continuously determine its position within the environment, for example, with respect to a map of the environment. The device 103 can also be configured to update the map (e.g., via a simultaneous mapping and localization process, or SLAM).

[0016] The device 103 is equipped with a variety of navigation and data acquisition sensors 104, such as image sensors (e.g., one or more digital cameras) and depth sensors (e.g., one or more LiDAR (Light Detection and Ranging) sensors, one or more depth cameras using structured light patterns, such as infrared light, or the like). The device 103 can be configured to use the sensors 104 both to navigate between the shelves 110 (e.g., according to the paths mentioned above) and to acquire shelf data, such as point clouds and image data, during such navigation.

[0017] Server 101 contains a dedicated image processing controller, such as a processor 120, specifically designed to control and / or support the mobile automation device 103 in navigating its environment and acquiring data. The processor 120 can also be configured to retrieve the acquired data via a communication interface 124, store it in a memory 132, and subsequently process it (e.g., to identify objects such as shelf products in the acquired data and determine corresponding status information). Server 101 can also be configured to send status messages (e.g., messages indicating that products are out of stock, low, or misplaced) to the client device 105 in response to the acquisition of product status data. The client device 105 contains one or more controllers (e.g.,central processing units (CPUs) and / or field-programmable gate arrays (FPGAs) and the like, which are configured to process (e.g. display) the messages received from Server 101.

[0018] The processor 120 is connected to a non-volatile, computer-readable storage medium, such as the memory 122 mentioned above, on which computer-readable instructions for performing various functions are stored, including controlling the device 103 for acquiring shelf data, post-processing the shelf data, and generating and providing certain navigation data to the device 103, such as target locations for shelf data acquisition. The memory 122 comprises a combination of volatile (e.g., random access memory or RAM) and non-volatile memory (e.g., read-only memory or ROM, electrically erasable programmable read-only memory or EEPROM, flash memory). The processor 120 and the memory 122 each comprise one or more integrated circuits. In some embodiments, the processor 120 is implemented as one or more central processing units (CPUs) and / or graphics processing units (GPUs).

[0019] Server 101 also includes the aforementioned communication interface 124, which is connected to processor 120. Communication interface 124 comprises suitable hardware (e.g., transmitters, receivers, network interface controllers, and the like) that enables Server 101 to communicate with other computing devices—specifically, Device 103, Client Device 105, and Dock 108—via connections 107 and 109. Connections 107 and 109 can be direct connections or connections that traverse one or more networks, including both local and wide area networks. The specific components of communication interface 124 are selected according to the type of network or other connections over which Server 101 will communicate.In the present example, as already mentioned, a wireless local area network is implemented in the retail environment through the use of one or more wireless access points. The connections 107 therefore include either wireless connections between the device 103 and the mobile device 105 and the access points mentioned above, as well as a wired connection (e.g., an Ethernet-based connection) between the server 101 and the access point.

[0020] Memory 122 stores a variety of applications, each containing a variety of computer-readable instructions that can be executed by Processor 120. The execution of these instructions by Processor 120 configures Server 101 to perform various actions discussed herein. The instructions stored in Memory 122 include a control application 128, which may also be implemented as a sequence of logically distinct applications. In general, Processor 120 is configured to execute Application 128, or parts thereof, and in conjunction with other components of Server 101, to implement various functions related to controlling Device 103 to navigate between Shelves 110 and acquire data, as well as to retrieve the acquired data and perform various post-processing operations on it.In the present example, which will be discussed in more detail below, the server 101 is configured by the execution of the application 128 to determine a shelf depth for one or more shelves 110 based on the acquired data (e.g. from the device 103), including the point cloud and image data representing the shelves 110.

[0021] The processor 120, as configured via the execution of the control application 128, is also referred to here as the controller 120. As will now become apparent, some or all of the functionality implemented by the controller 120 described below can also be performed by pre-configured special hardware controllers (e.g., one or more logic circuit arrangements specifically configured to optimize the speed of image processing, e.g., via FPGAs and / or application-specific integrated circuits (ASICs) configured for this purpose) instead of by the processor 120 executing the control application 128.

[0022] In the Fig. 2A and Fig. Figure 2B shows the mobile automation device 103 in more detail. The device 103 comprises a frame 201 containing a drive mechanism 203 (e.g., one or more electric motors that drive wheels, rails, or the like). The device 103 further comprises a sensor mast 205, which is supported on the frame 201 and, in this example, extends upwards (e.g., substantially vertically) from the frame 201. The mast 205 carries the aforementioned sensors 104. The sensors 104 include, in particular, at least one image sensor 207, such as a digital camera, and at least one depth sensor 209, such as a 3D digital camera. The device 103 also includes additional depth sensors, such as LiDAR sensors 211. In other examples, the device 103 includes additional sensors, such as... B. one or more RFID readers, temperature sensors and the like.

[0023] In the present example, the mast 205 carries seven digital cameras 207-1 to 207-7 and two LiDAR sensors 211-1 and 211-2. The mast 205 also carries a plurality of lighting assemblies 213, which are configured to illuminate the fields of view of the respective cameras 207. That is, lighting assembly 213-1 illuminates the field of view of camera 207-1, and so on. The sensors 207 and 211 are oriented on the mast 205 such that the fields of view of each sensor face a shelf 110, along the length 119 of which the device 103 moves. The device 103 is configured to track a location of the device 103 (e.g., a location of the center point of the rack 201) in the common reference frame 102 previously established in the retail facility, so that the data acquired by the mobile automation device 103 can be registered in the common reference frame.

[0024] The mobile automation device 103 includes a special controller, such as a processor 220, as shown in Fig. Figure 2B shows a processor 220 connected to a non-transient, computer-readable storage medium, such as a memory 222. The memory 222 comprises a combination of volatile (e.g., Random Access Memory or RAM) and non-volatile memory (e.g., Read Only Memory or ROM, Electrically Erasable Programmable Read Only Memory or EEPROM, Flash Memory). The processor 220 and the memory 222 each comprise one or more integrated circuits. The memory 222 stores computer-readable instructions for execution by the processor 220. In particular, the memory 222 stores a control application 228 which, when executed by the processor 220, configures the processor 220 to perform various functions related to the navigation of the device 103 and the acquisition of data for subsequent processing, e.g., by the server 101.In some embodiments, such post-processing can be carried out by the device 103 itself via the execution of application 228. In other examples, application 228 can also be implemented as a series of different applications.

[0025] When configured by the execution of application 228, the processor 220 can also be referred to as the imaging controller 220. Those skilled in the art will recognize that the functionality implemented by the processor 220 through the execution of application 228 can also be implemented by one or more specially designed hardware and firmware components, including logic circuit configurations optimized for processing image and / or depth sensor data, such as specially configured FPGAs, ASICs, and the like in other embodiments.

[0026] The memory 222 can also store an archive 232, which contains, for example, one or more maps representing the environment in which the device 103 operates, for use during the execution of the application 228. The device 103 can communicate with the server 101 via a communication interface 224 through the Fig. The connection 107 shown in the diagram allows communication, for example, to receive instructions for navigation to specific locations and to initiate data collection processes. The communication interface 224 also enables the device 103 to communicate with the server 101 via the dock 108 and the connection 109.

[0027] As will become clear in the following discussion, in other examples some or all of the processing operations performed by server 101 can be performed by device 103, and some or all of the processing operations performed by device 103 can be performed by server 101. That is, although in the example shown application 128 is located in server 101, in other embodiments some or all of the actions described below for determining the shelf depth of shelves 110 from acquired data can be performed by processor 220 of device 103, either in conjunction with or independently of processor 120 of server 101.As experts will recognize, the distribution of such calculations between the server 101 and the mobile automation device 103 may depend on the respective processing speeds of the processors 120 and 220, the quality and bandwidth of the connection 107, and the criticality level of the underlying instruction(s).

[0028] The functionality of application 128 will now be described in more detail. In particular, the determination of the carrier structure depth mentioned above, as performed by server 101, will be described. Fig. Figure 3 describes a method 300 for determining the carrier structure depth. Method 300 is described in connection with its execution by Server 101, with reference to the information in Fig. The components shown in the diagram are described.

[0029] In block 305, server 101 is configured to receive a point cloud of the support structure as well as a plane definition that corresponds to the front face of the support structure. In the present example, where the support structures are shelves like the ones in Fig. Since the shelves 110 shown in Block 305 are represented by the point cloud obtained in Block 305, this represents at least part of a shelf module 110 (and can represent a multitude of shelf modules 110), and the plane definition corresponds to a shelf plane that corresponds to the front face of the shelf modules 110. In other words, the plane definition defines a plane that contains the shelf edges 118.

[0030] The point cloud and plane definition obtained in block 305 can be retrieved from archive 132. For example, server 101 may have previously received acquired data from device 103, including a multitude of lidar scans of the shelf modules 110, and generated a point cloud from these lidar scans. Each point in the point cloud represents a point on a surface of the shelves 110, products 112, and the like (e.g., a point encountered by the scan line of a lidar sensor 211) and is defined by a set of coordinates (X, Y, and Z) in reference frame 102. The plane definition may also have been previously generated by server 101 and stored in archive 132, e.g., from the point cloud mentioned above.Server 101 can, for example, be configured to process the point cloud, the lidar raw data, the image data captured by the cameras 207, or a combination thereof, to identify the shelf edges 118 according to the predefined properties of the shelf edges 118. Examples of such properties are that the shelf edges 118 are likely to be essentially planar and also likely to be closer to the device 103 than other objects (such as the shelf backs 116 and products 112) when the device 103 traverses the length 119 of a shelf module 110. The plane definition can be obtained in a variety of suitable formats, such as a suitable set of parameters that define the plane. An example of such parameters is a normal vector (i.e.,a vector defined according to the reference frame 102 and perpendicular to the plane) and a depth (which specifies the distance along the normal vector from the origin of the reference frame 102 to the plane).

[0031] Referring to Fig. Figure 4A shows a point cloud 400, which depicts the shelf module 110-3. The shelf back 116-3, as well as the shelf 117-3 and the shelf edge 118-3, are therefore shown in the point cloud 400. Also in Fig. Figure 4A shows a plane definition 404 that corresponds to the front of the shelf module 110-3 (i.e., the plane definition 404 contains the shelf edges 118-3). The point cloud 400 and the plane definition 404 do not need to be in the Fig. The graphical form shown in Figure 4A is available. As experts will recognize, the point cloud can be obtained as a list of coordinates and the plane definition 404 as the parameters mentioned above.

[0032] Back to Fig. In block 310, server 101 is configured to receive at least one image of the support structure as captured (e.g., by device 103) from a specific capture position. The capture position is the position and orientation of the capture device, such as a camera 207, within the reference frame. Device 103, as mentioned above, is configured to traverse one or more shelf modules 110 and capture images of them. As can now be seen, each image capture occurs at a specific position and orientation of device 103. Furthermore, device 103 includes a plurality of cameras 207, as shown in Fig. Figure 2A shows each camera 207 having a predefined physical position and orientation on the device 103. Thus, for each position (i.e., position and orientation) of the device 103, a plurality of images can be captured, one for each camera 207. Each image therefore corresponds to a specific capture position, i.e., the physical position of the camera 207 according to the reference frame 102.

[0033] Fig. Figure 4B illustrates the acquisition of two example images 408-1 and 408-2 by the device 103 while the device 103 moves through the shelf module 110-3 in a direction of movement 406. Specifically, in a first device position 412-1, the device 103 controls the camera 207-1 to acquire the first image 408-1. The position and orientation of the camera 207-1 at the time of acquiring the first image 408-1 thus corresponds to a first acquisition position. Later, during the movement past the shelf module 110-3, in a second device position 412-2, the device 103 controls the camera 207-1 to acquire the second image 408-2. As can now be seen, the second image 408-2 corresponds to a second capture position, which is defined by the device position 412-2 and the physical orientation of the camera 207-1 relative to the device 103.Each of the other cameras 207 can also be controlled to capture images in any device position 412. The images captured by these other cameras 207 correspond to even more capture positions.

[0034] Back to Fig. In block 310, server 101 is also configured to receive a mask, also known as a BoS (Back of Shelf) mask or BoS map, for example, by retrieving it from archive 132. The mask corresponds to the at least one image mentioned above. That is, for every image retrieved in block 310, a corresponding mask can also be retrieved. The mask is derived from the corresponding image and indicates, for each of a multitude of sections of the image, a confidence level that the section represents the back of the shelf 116. The sections can be single pixels if the mask has the same resolution as the image. In other examples, the mask has a lower resolution than the image, and each confidence level in the mask therefore corresponds to a section of the image containing multiple pixels.

[0035] Fig. 5A illustrates this in Fig. 4A shown in Figure 408-1. As in Fig. As shown in Figure 5A, section 500 of Figure 408-1 represents the back of the shelf 116-3. Sections 504-1, 504-2, and 504-3 show the products 112, and section 508 shows the shelf edge 118-3. Fig. Figure 5B shows a mask 512 derived from image 408-1. Various mechanisms can be used to generate the mask. For example, image 408-1 can be decomposed into fields of a predefined size (e.g., 5 x 5 pixels), and each field can be classified by a suitable classification operation to generate a confidence level indicating the degree to which the field matches a reference image of the shelf back 116-3. The mask 512 can then be constructed by combining the confidence levels assigned to each field.

[0036] In Fig. Figure 5B shows the confidence levels of mask 512 in grayscale. Darker areas of mask 512 indicate lower confidence that the corresponding section of image 408-1 represents shelf back 116-3 (or, in other words, higher confidence that the corresponding section of image 408-1 does not represent shelf back 116-3), and lighter areas of mask 512 indicate higher confidence that the corresponding section of image 408-1 represents shelf back 116-3. For example, an area 516 indicates a confidence level of zero that section 508 of image 408-1 represents shelf back 116-3. Another area 520 of mask 512 indicates a maximum confidence level (e.g., 100%) that the corresponding section of image 408-1 represents the back of shelf 116-3. Other areas of mask 512 indicate medium confidence levels.For example, area 524 indicates a confidence level of approximately 50%, because the pattern on product 112, shown in section 504-3 of image 408-1, resembles the shelf back 116-3. Area 528 of mask 512, on the other hand, indicates a confidence level of approximately 30%.

[0037] About the in Fig. Beyond the grayscale image shown in Figure 5B, various other mechanisms for storing the confidence levels of mask 512 are conceivable. For example, the confidence levels can be stored in a list, with associated sets of image coordinates indicating which section of image 408-1 corresponds to the respective confidence level.

[0038] After the point cloud, the layer definition, the image(s), and the mask(s) in blocks 305 and 310 have been obtained, server 101 is configured to identify a subset of the points in the point cloud for which corresponding confidence levels exist in mask 512. That is, server 101 identifies points in the point cloud that were visible to camera 207 at the time the image was captured. Server 101 is then configured to use the depths of such points relative to the shelf level, in conjunction with the corresponding confidence levels from mask 512, to determine the depth of the shelf back 116 relative to the shelf level. The above functionality is explained in more detail below.

[0039] Back to Fig. Block 315 of server 101 is configured to select an initial set of points from the point cloud that fall within a field of view defined by the aforementioned acquisition position. Fig. Figure 4B shows the field of view of the camera 207-2 in each detection position 412, indicated by dashed lines. The detection position is defined according to the reference frame 102, and the position and extent of the field of view within the reference frame 102 can also be defined according to the predefined operating parameters (e.g., focal length) of the camera 207.

[0040] Server 101 can be configured to evaluate each point in the point cloud in Block 315 to determine whether the point falls within the field of view corresponding to the image obtained in Block 310. For example, Server 101 can be configured to define the field of view as a volume within the reference frame 102 and determine whether each point in the point cloud falls within the defined volume. Points that fall within the defined volume are selected for the initial set. However, in some examples, Server 101 is configured to perform a tree-based search to generate the initial set of points, as shown below in conjunction with Fig. 6 explained.

[0041] In Fig. Section 6A describes a procedure 600 for selecting the initial set of points in block 315. In block 605, server 101 is configured to generate a tree data structure, such as a kd (k-dimensional) tree, an octree, or the like. In this example, a kd tree is generated in block 605. The tree data structure contains, for each point in the point cloud, first- and second-dimensional coordinates orthogonal to the point's depth. That is, each point is represented in the tree by its X and Z coordinates according to reference frame 102, omitting the Y coordinate for selecting the initial set (the Y coordinates are used later in procedure 300, as explained below).

[0042] As experts understand, the kd-tree can be constructed by determining the median of one of the two dimensions mentioned above (e.g., the X dimension). All points whose X-coordinate is below the median are assigned to a first branch of the tree, while the remaining points are assigned to a second branch. For each branch, the median of the other coordinate (in this example, Z) is determined, and the points assigned to that branch are further subdivided depending on whether their Z-coordinates are above or below the Z-median. This process is repeated, with the points between the pairs of branches being further subdivided based on alternating dimensional medians (i.e., a subdivision based on the X dimension, followed by a subdivision based on the Z dimension, followed by another subdivision based on the X dimension, and so on), until each node of the tree contains a single point.

[0043] In block 610, server 101 is configured to determine the coordinates of a field-of-view center in the two dimensions shown in the tree. As mentioned above, the volume defined by the field of view is determined from the operating parameters of camera 207 and the detection position. Fig. Figure 6B shows a field of view 602 of camera 207-2 in the device position 412-1. The center of the field of view 602 is defined in three dimensions by line 604. To determine two-dimensional coordinates of the center of the field of view in two dimensions (i.e., in the X and Z dimensions), server 101 is configured to select a predefined depth and determine the coordinates where the center line 604 intersects the predefined depth. The predefined depth can be stored in memory 122 as a depth that is added to the shelf-level depth obtained in block 305. In other examples, the point cloud, the shelf level, the capture positions, and the like can be transformed into a reference frame whose origin is on the shelf level itself (e.g., whose XZ plane is on the shelf level) to simplify the calculations described here. As in Fig. As shown in Figure 6B, the center line 604 intersects the predefined depth at a FOV (Field of View) center 608. The predefined depth is preferably chosen to exceed the depth of the shelf back 116 (although the depth of the shelf back 116 may not be known exactly).

[0044] In block 615, server 101 is configured to select points for the set by retrieving points from the tree that lie within a predefined radius of center 608. Fig. Figure 6B shows a predefined radius 612 extending from the center 608. As experts will now see, there are various mechanisms for performing radius-based searches in trees such as kd-trees. In the present example, as in Fig. As shown in 6B, the points retrieved in block 615 include the example points 616, while other points 618 are not retrieved because they are further away from the center 608 than the radius 612.

[0045] In block 620, server 101 can be configured to check whether the three-dimensional position of each point retrieved in block 615 lies within the FOV 602, since the predefined radius 612 may extend beyond the actual limits of the FOV 602. In other examples, block 620 can be omitted. If performed, the check in block 620 can use a transformation matrix, also known as a camera calibration matrix, configured to transform three-dimensional coordinates from the point cloud into two-dimensional coordinates in an image reference frame (e.g., pixel coordinates in image 408-1). Therefore, for each point retrieved in block 615, the check in block 620 can include generating the corresponding image coordinates and determining whether the image coordinates lie within the limits of image 408-1.

[0046] Back to Fig. 3. After the initial set of points 616 within the FOV 602 has been selected, server 101 in block 320 is configured to select an unobstructed set of depth measurements from the points in the initial set. The initial set of points selected in block 315, although within the FOV 602, may still contain points that were not imaged by camera 207 because they are obscured from camera 207's view by other objects. For example, with further reference to Fig. 6B, although it lies within the FOV 602, corresponds to a section of shelf 117, which camera 207-2 in the Fig. The acquisition position shown in 6B cannot be depicted because a product 112 is located between camera 207-2 and point 616a. Point 616a may appear in the point cloud because a lidar scanner is positioned differently than camera 207-2, because a lidar scanner acquired point 616a from a subsequent device position 412, or similar reasons. In other words, point 616a is a hidden point for which image 408-1 and mask 512 do not have corresponding data. In block 320, server 101 is configured to remove such hidden points from further consideration and retain data for unhidden points, such as point 616b.

[0047] In general, the selection in Block 320 assumes that for every occluded point in the point cloud, there is also an unoccluded point in the point cloud that corresponds to the object responsible for the occlusion. Block 320 further assumes that the aforementioned unoccluded point is visible to camera 207 and is therefore shown in Figure 408-1. Fig. 7A is an example procedure 700 for selecting the unhidden subset of depth measurements.

[0048] In block 705, server 101 is configured to determine the image coordinates for each point in the source set selected in block 315. As mentioned above, the image coordinates can be obtained using the camera calibration matrix in a process also known as forward projection (i.e., projecting a point in three dimensions "forward" into a captured image, as opposed to back projection, where a point in an image is projected "backward" into the point cloud). Fig. Figure 7B illustrates the results of Block 705 for the previously discussed points 616a and 616b. Points 616a and 616b correspond to image coordinates defined according to an image reference system 702 (which in this example lies parallel to the XZ plane of reference system 102). As in Fig. As can be seen in Figure 7B, the depth measurements in reference frame 102, which are assigned to points 616, are also retained when carrying out procedure 700, although they are not directly represented in the image coordinates (which are two-dimensional). The depth measurements can be stored in a list 704 or another suitable format in conjunction with the image coordinates. Further example points 706, 708, 709, 712, and 714 are also shown. As shown in the list 704 of depth measurements, point 708 is on the surface of a product, while point 709 is behind the product, e.g., on the back of the shelf 116 (at a depth of 528 mm, compared to a depth of 235 mm for point 708). Points 712 and 714 are also located on the back of shelf 116 and have corresponding depths of 530 mm and 522 mm respectively.

[0049] In block 710, server 101 is configured to create a tree data structure, for example, another kd tree containing the image coordinates determined in block 705. In block 715, server 101 is configured to select neighboring groups of points. Specifically, server 101 is configured to retrieve the nearest neighbors of a selected point in the tree (for example, a predefined number of neighbors, neighbors within a predefined radius, or a combination of the above). Server 101 is further configured to select the neighbor with the shallowest depth from the nearest neighbors retrieved in block 715. Thus, referring again to Fig. 7B and starting with point 616a, the nearest neighbor is point 616b, and the shallowest depth between points 616a and 616b is the depth assigned to point 616b. The depth measurement (and the corresponding image coordinates) of point 616b is therefore retained for inclusion in the unobscured subset of depth measurements, while point 616a is discarded. The execution of block 715 is repeated for each remaining point in the initial set until a determination in block 720 indicates that no more points are to be processed. For the in Fig. Example points shown in 7B,

[0050] Once all points from the initial set have been processed and the subset of unobscured depth measurements has been selected, Server 101 returns to Block 325 of Procedure 300. In Block 325, Server 101 can optionally be configured to select a final subset of depth measurements from the unobscured subset of depth measurements. For example, if one selects the Fig. Taking the points shown in 7B, the resulting unobscured subset of depth measurements is in Fig. Figure 8A shows that points 616a and 709 have been discarded from the unhidden subset 800. In block 325, server 101 can be configured to perform one or more additional filter operations to exclude further points from the unhidden subset.

[0051] A first example of a filtering operation applied in block 325 is to discard all points from mask 512 with a BoS confidence level below a predefined threshold. In this example, the predefined threshold is 55% (however, it is understood that various other thresholds could be applied instead). Fig. Figure 8B shows mask 512 with confidence levels 816b, 806, 808, 812, and 814, corresponding to the image coordinates of points 616b, 706, 708, 712, and 714, respectively. In this example, confidence levels 816b, 806, 808, 812, and 814 are assumed to be 30%, 0%, 50%, 100%, and 90%, respectively. Points with confidence levels below 55% (i.e., points 616b, 706, and 708) are therefore discarded, and the final subset of depth measurements includes the depth measurements for points 712 and 714, along with their associated image coordinates.

[0052] Other examples of the filtering performed in Block 325 include discarding points with depth measurements that exceed a predefined maximum depth threshold. Fig. Figure 8C shows another example image 818, which was captured from a different fixture position (and thus in a different acquisition position) than the position in which image 418-1 was captured. In image 818, an edge 819 of the shelf module 110-3 is visible, and certain points in both image 818 and the point cloud therefore correspond to areas of the facility beyond the shelf module 110-3. For example, point 820 may have an associated depth measurement of 2500 mm. The maximum threshold mentioned above can be chosen as the maximum known shelf depth in the entire facility (e.g., 700 mm). Point 820 can therefore be discarded in block 325.

[0053] Back to Fig. In block 330, the depth measurements of the final subset are weighted according to the corresponding confidence levels from mask 512. In this example, the depth measurements for points 712 and 714 are weighted according to their respective confidence levels (100% and 90%, respectively). For example, the depths can be multiplied by their respective weights (e.g., 530 x 1 and 522 x 0.9). In block 335, the shelf height is determined from the weighted depths. This means that the shelf depth determined in block 335 is a weighted average of the depth measurements in the final subset from block 325. In the present example, the weighted average of the depth measurements for points 712 and 714 is determined by summing the weighted depths and dividing the result by the sum of the weights (i.e., 1.9 or 190%), which yields a result of 526.2 mm. Fig.Figure 9 shows the determined shelf depth as a dashed line 900, which starts from shelf level 404 (and is perpendicular to shelf level 404).

[0054] In block 340, server 101 is configured to determine whether there are any remaining capture positions to process (i.e., whether there are any further fixture positions for the current camera or whether any other cameras remain in the current fixture position). If the determination in block 340 is positive, the execution of procedure 300 is repeated for all subsequent images and corresponding masks. If the determination in block 340 is negative, the execution of procedure 300 is terminated. In some examples, block 335 is executed only after a negative determination in block 340 and uses the multiple weighted final sets of depth measurements from each execution of block 330 to determine a single shelf depth for shelf module 110.The shelf depth determined by carrying out procedure 300 can, for example, be returned to another application of server 101 (or to another computer device) to identify gaps in shelves 110 or other object condition data.

[0055] Specific embodiments have been described in the foregoing description. However, a person skilled in the art will recognize that various modifications and alterations can be made without altering the scope of protection of the invention as defined in the claims below. Accordingly, the description and figures are to be regarded in an illustrative rather than a limiting sense, and all such modifications are to be included within the scope of the present teachings.

[0056] The benefits, advantages, solutions to problems, and all elements that may lead to the occurrence or enhancement of a benefit, advantage, or solution are not to be understood as critical, necessary, or essential features or elements in some or all of the claims. The invention is defined solely by the attached claims, including any amendments made during the pendency of this application and all equivalents of the granted claims.

[0057] Furthermore, in this document, relational terms such as first and second, upper and lower, and the like may be used merely to distinguish one entity or action from another, without necessarily requiring or implying any actual relationship or order of such an entity or action between such entities or actions. The expressions "includes," "comprising," "has," "have," "exhibits," "exhibiting," "contains," "containing," or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, procedure, product, or device that includes, has, exhibits, or contains a list of elements may not only have those elements but may also have other elements not expressly listed or inherent in such process, procedure, product, or device. An element that "includes," "has," "exhibits," or "contains"The use of the term "a" does not, without further limitations, preclude the existence of additional identical elements in the process, method, product, or apparatus that includes, has, features, or contains the element. The terms "a" and "a" are defined as one or more unless expressly stated otherwise herein. The terms "essentially," "generally," "approximately," "about," or any other version thereof are defined in such a way as to be understood by a person skilled in the art in this field, and in one non-restrictive embodiment, the expression is defined as within 10%, in another embodiment as within 5%, in yet another embodiment as within 1%, and in yet another embodiment as within 0.5%. The term "coupled," as used herein, is defined as connected, but not necessarily directly and not necessarily mechanically.A device or structure that is “designed” in a certain way is at least also designed in that way, but may also be designed in ways that are not listed.

[0058] It is understood that some embodiments may include one or more generic or specialized processors (or “processing devices”) such as microprocessors, digital signal processors, custom processors, and field-programmable gate arrays (FPGAs), and uniquely stored program instructions (including both software and firmware) that control the one or more processors to implement, in conjunction with certain non-processor circuitry, some, most, or all of the functions of the method and / or device described herein. Alternatively, some or all of the functions may be implemented by a state machine that does not have any stored program instructions, or in one or more application-specific integrated circuits (ASICs) in which each function, or some combinations of certain functions, are implemented as user-defined logic.Of course, a combination of the two approaches can be used.

[0059] Furthermore, an embodiment may be implemented as a computer-readable storage medium on which computer-readable code is stored for programming a computer (which, for example, includes a processor) to execute a method as described and claimed herein. Examples of such computer-readable storage media include, but are not limited to, a hard disk, a CD-ROM, an optical storage device, a magnetic storage device, a ROM (read-only memory), a PROM (programmable read-only memory), an EPROM (erasable programmable read-only memory), an EEPROM (electrically erasable programmable read-only memory).Furthermore, it is assumed that an average professional, regardless of possible significant effort and many design choices motivated, for example, by available time, current technology, and economic considerations, will be readily able to generate such software instructions, programs, and ICs with minimal experimentation if guided by the concepts and principles disclosed herein.

[0060] The summary of the disclosure is provided to enable the reader to quickly ascertain the essence of the technical disclosure. It is provided with the understanding that it is not intended to be used for interpreting or limiting the scope or meaning of the claims. Furthermore, it can be inferred from the preceding detailed description that various features in different embodiments have been summarized for the purpose of streamlining the disclosure. This type of disclosure is not to be interpreted as reflecting the intention that the claimed embodiments require more features than are expressly stated in each claim. Rather, as the following claims demonstrate, the inventive step lies in fewer than all the features of a single disclosed embodiment.The following claims are hereby incorporated into the detailed description, each claim being a separately claimed subject matter.

Claims

[1] Method for determining a support structure depth of a support structure having a front and a back side separated by the support structure depth, the method comprising: Obtain (i) a point cloud (400) of the support structure, and (ii) a mask (512) indicating, for a plurality of sections of an image of the support structure acquired from a capture position, respective confidence levels that the sections represent the back side of the support structure; Selecting an initial set of points from the point cloud (400) that are within a field of view originating from the acquisition position; Selecting a non-hidden subset of depth measurements from the initial set of points, where the depth measurements in the non-hidden subset correspond to the respective image coordinates; Retrieving a confidence level from mask (512) for each of the depth measurements in the unmasked subset; and Determining the carrier structure depth based on the depth measurements in the unhidden subset and the retrieved confidence levels. [2] The method of claim 1, further comprising: Receiving another mask (512) that corresponds to another image of the support structure, which was captured from another acquisition position; Select another starting set of points; Select another unobscured set of depth measurements; Retrieving another confidence level for each of the depth measurements in the further unhidden subset. [3] The method of claim 2, further comprising: Determining the carrier structure depth based on the depth measurements in the unhidden subset, the depth measurements in the further unhidden subset, the confidence levels, and the further confidence levels. [4] Method according to claim 1, wherein a detection position defines a camera position and orientation within a common reference frame. [5] Method according to claim 1, wherein the selection of the initial set of points comprises: Generating a tree data structure that includes a first and a second dimension orthogonal to the support structure depth for each point in the point cloud (400); Determining a visual field center in the first and second dimensions; and Retrieving the initial set of points within a predefined radius of the viewport center from the tree data structure. [6] Method according to claim 1, wherein selecting the unobscured subset of depth measurements comprises: Determining image coordinates that correspond to each of the initial set of points; Identifying neighboring groups of the image coordinates; and Select the image coordinate that corresponds to the smallest depth measurement for each neighboring group. [7] Method according to claim 1, wherein determining the support structure depth comprises: Obtaining a plane definition (404) that corresponds to the front face of the support structure; Transforming each depth measurement of the unhidden subset of depth measurements into a depth relative to the level definition (404); Weights of each transformed depth measurement according to the retrieved confidence levels; and Determining an average of the weighted depth measurements. [8] Method according to claim 1, further comprising: Before determining the support structure depth, discard depth measurements for which the retrieved confidence levels are below a minimum confidence threshold. [9] Method according to claim 1, further comprising: Before determining the support structure depth, discard depth measurements that exceed a maximum depth threshold. [10] Computer device for determining a support structure depth of a support structure having a front and a back side separated by the support structure depth, the computer device comprising: a memory that stores (i) a point cloud (400) of the support structure and (ii) a mask (512) that indicates, for a plurality of sections of an image of the support structure acquired from a capture position, respective confidence levels that the sections represent the back side of the support structure; an imaging controller that is connected to the memory and configured to: selects an initial set of points from the point cloud (400) that are within a field of view originating from the acquisition position; selects a non-hidden subset of depth measurements from the initial set of points, where the depth measurements in the non-hidden subset correspond to the respective image coordinates; retrieves a confidence level from mask (512) for each of the depth measurements in the unhidden subset; and The support structure depth is determined based on the depth measurements in the unhidden subset and the retrieved confidence levels. [11] Computer device according to claim 10, wherein the imaging control is further configured such that it: a further mask (512) is obtained, which corresponds to a further image of the support structure that was captured from a further capture position; selects another initial set of points; selects another unobscured set of depth measurements; retrieves another confidence level for each of the depth measurements in the further unhidden subset. [12] Computer device according to claim 11, wherein the imaging control is further configured such that it: The support structure depth is determined based on the depth measurements in the unhidden subset, the depth measurements in the further unhidden subset, the confidence levels, and the further confidence levels. [13] Computer device according to claim 10, wherein a detection position defines a camera position and orientation within a common reference frame. [14] Computer device according to claim 10, wherein the imaging control is further configured to select the initial set of points in order to: to create a tree data structure that includes a first and a second dimension orthogonal to the support structure depth for each point in the point cloud (400); to determine a visual field center in the first and second dimensions; and to retrieve the initial set of points within a predefined radius of the center of the field of view from the tree data structure. [15] Computer device according to claim 10, wherein the imaging control is further configured to select the unobscured subset of depth measurements in order to: to determine image coordinates that correspond to each of the initial set of points; to identify neighboring groups of the image coordinates; and For each neighboring group, select the image coordinate that corresponds to the smallest depth measurement. [16] Computer device according to claim 10, wherein the imaging control is further configured to determine the support structure depth in order to: to obtain a plane definition (404) that corresponds to the front face of the support structure; to transform each depth measurement of the unhidden subset of depth measurements into a depth relative to the level definition (404); to weight each transformed depth measurement according to the retrieved confidence levels; and to determine an average of the weighted depth measurements. [17] Computer device according to claim 10, wherein the imaging control is further configured such that it: Before determining the support structure depth, it rejects depth measurements for which the retrieved confidence levels are below a minimum confidence threshold. [18] Computer device according to claim 10, wherein the imaging control is further configured such that it: Before determining the support structure depth, it rejects depth measurements that exceed a maximum depth threshold. [19] Computer-readable medium on which computer-executable instructions are stored, the instructions comprising: Obtain (i) a point cloud (400) of the support structure, and (ii) a mask (512) indicating, for a plurality of sections of an image of the support structure acquired from a capture position, respective confidence levels that the sections represent the back side of the support structure; Selecting an initial set of points from the point cloud (400) that are within a field of view originating from the acquisition position; Selecting a non-hidden subset of depth measurements from the initial set of points, where the depth measurements in the non-hidden subset correspond to the respective image coordinates; Retrieving a confidence level from mask (512) for each of the depth measurements in the unmasked subset; and Determining the carrier structure depth based on the depth measurements in the unhidden subset and the retrieved confidence levels. [20] Computer-readable medium according to claim 19, wherein the instructions further comprise: Determining the carrier structure depth based on the depth measurements in the unhidden subset, the depth measurements in the further unhidden subset, the confidence levels, and the further confidence levels.

Citation Information

Patent Citations

  • Imaging-based sensor calibration

    US20190073550A1

  • Counting inventory items using image analysis and depth information

    US9996818B1