Detecting identification markers on picking stations

The use of neural networks and ultra-wide-angle cameras in the grid-based storage system allows for precise identification and maintenance of robotic picking stations, addressing operational challenges and enhancing system efficiency.

JP2025538690APending Publication Date: 2025-11-28OCADO INNOVATION LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025531853
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-30
Filing Date
2023-11-30
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing grid-based storage systems face challenges in efficiently identifying and maintaining robotic picking stations within a grid-based storage system, which is crucial for effective operation and maintenance.

Method used

A computer-implemented method and detection system using neural networks to detect and identify markers on picking stations within a grid-based storage system, utilizing ultra-wide-angle cameras and calibration processes to accurately map image data to the grid coordinates, enabling precise identification and location of picking stations.

Benefits of technology

Enables efficient identification and maintenance of robotic picking stations, facilitating timely resolution of failures and optimizing the operation of the entire grid-based storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025538690000001_ABST
    Figure 2025538690000001_ABST
Patent Text Reader

Abstract

A detection system and method for detecting identification markers on picking stations on a grid comprising a plurality of grid cells forming part of a grid-based storage system in which one or more picking stations are mounted on the grid. Each picking station comprises a robotic manipulator for transferring items between containers received in the respective grid cell adjacent to the picking station. The method involves acquiring image data representing an image portion including a picking station of the one or more picking stations. The image data is processed sequentially using a first neural network and a second neural network. The first neural network is trained to detect instances of identification markers on the picking stations in the image. The second neural network is trained to recognize marker information in the image associated with the identification markers. Marker data representing the marker information determined by the second neural network is output.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to the field of grid-based storage systems, and more particularly to detecting identification markers on picking stations in a grid that forms part of the grid-based storage system. [Background technology]

[0002] Online retail businesses that sell multiple product lines, such as online grocery stores and supermarkets, need systems that can store tens or even hundreds of thousands of different product lines. In such cases, the use of single-product stacks may be impractical because a huge amount of floor space would be required to accommodate all of the required stacks. Additionally, it may be desirable to store small quantities of some items, such as perishable or infrequently ordered goods, making single-product stacks an inefficient solution.

[0003] PCT Publication No. WO2015 / 185628A (Ocado) describes a further known storage and fulfillment system in which stacks of containers are arranged within a grid framework structure. The containers are accessed by one or more load handling devices, sometimes known as robots or "bots," operable on tracks on top of the grid framework structure. A system of this type is shown schematically in Figures 1 to 3 of the accompanying drawings.

[0004] As shown in FIGS. 1 and 2, stackable containers 10, also known as "bins," are stacked on top of each other to form a stack 12. The stacks 12 are arranged in a grid framework structure 14, for example, in a warehousing or manufacturing environment. The grid framework structure 14 consists of a plurality of storage rows or grid rows. Each grid in the grid framework structure has at least one grid row for storing a stack of containers. FIG. 1 is a schematic perspective view of the grid framework structure 14, and FIG. 2 is a schematic top-down view showing a stack 12 of bins 10 arranged within the framework structure 14. Each bin 10 typically holds multiple product items (not shown). The product items in the bins 10 can be of the same or different product types, depending on the application.

[0005] The grid framework structure 14 includes a plurality of upright members 16 supporting horizontal members 18, 20. A first set of parallel horizontal grid members 18 are arranged in a grid pattern, perpendicular to a second set of parallel horizontal members 20, to form a horizontal grid structure 15 supported by the upright members 16. The members 16, 18, 20 are typically fabricated from metal. The bins 10 are stacked between the members 16, 18, 20 of the grid framework structure 14 such that the grid framework structure 14 guards against horizontal movement of the stack 12 of bins 10 and guides vertical movement of the bins 10.

[0006] The top level of the grid framework structure 14 comprises a grid or grid structure 15 including rails 22 arranged in a grid pattern across the top of the stacks 12. Referring to FIG. 3 , the rails or tracks 22 guide a plurality of load handling devices 30. A first set 22a of parallel tracks or rails 22 guides movement of the robotic load handling devices 30 in a first direction (e.g., the X direction) across the top of the grid framework structure 14. A second set 22b of parallel tracks or rails 22, positioned orthogonal to the first set 22a, guides movement of the load handling devices 30 in a second direction (e.g., the Y direction) orthogonal to the first direction. In this manner, the tracks or rails 22 allow the robotic load handling devices 30 to move laterally in two dimensions in the horizontal XY plane. The load handling devices 30 can be moved to a position above any of the stacks 12.

[0007] A known form of load handling device 30, shown in Figures 4 and 5, is described in PCT Patent Publication No. WO2015 / 019055 (Ocado), which is incorporated herein by reference, with each load handling device 30 covering a single grid space 17 of the grid framework structure 14. This configuration allows for a higher density of load handlers and therefore a higher throughput for a storage system of a given size.

[0008] The exemplary load handling device 30 includes vehicles 32 positioned to roll on the rails 22 of the grid framework structure 14. A first set of wheels 34, consisting of a pair of wheels 34 at the front of the vehicle 32 and a pair of wheels 34 at the rear of the vehicle 32, are positioned to engage two adjacent rails of the first set 22a of rails 22. Similarly, a second set of wheels 36, consisting of a pair of wheels 36 on each side of the vehicle 32, are positioned to engage two adjacent rails of the second set 22b of rails 22. At any time during the movement of the load handling device 30, each set of wheels 34, 36 can be raised and lowered so that either the first set of wheels 34 or the second set of wheels 36 is engaged with the respective set of rails 22a, 22b. For example, when a first set of wheels 34 is engaged with a first set of rails 22a and a second set of wheels 36 is lifted off the rails 22, the first set of wheels 34 can be driven by a drive mechanism (not shown) housed in the vehicle 32 to move the load handling device 30 in the X direction. To achieve movement in the Y direction, the first set of wheels 34 is lifted off the rails 22 and the second set of wheels 36 is lowered to engage with a second set 22b of rails 22. The drive mechanism can then be used to drive the second set of wheels 36 to move the load handling device 30 in the Y direction.

[0009] The load handling device 30 is equipped with a lifting mechanism, e.g., a crane mechanism, for lifting a storage container from above. The lifting mechanism includes a winch tether or cable 38 wound on a spool or reel (not shown) and a gripper device 39. The lifting mechanism shown in FIGS. 4 and 5 includes a set of four vertically extending lifting tethers 38. The tethers 38 are connected at or near each of the four corners of the gripper device 39, e.g., a lifting frame, for releasable connection to the storage container 10. For example, each tether 38 is positioned at or near each of the four corners of the lifting frame 39. The gripper device 39 is configured to releasably grip the top of the storage container 10 to lift it from a stack of containers in a storage system 1 of the type shown in FIGS. 1 and 2. For example, the lifting frame 39 may include pins (not shown) that mate with corresponding holes (not shown) in a rim forming the top surface of the bin 10 and sliding clips (not shown) that are engageable with the rim to grip the bin 10. The clips are housed within a lifting frame 39 and are driven into engagement with the bins 10 by a suitable drive mechanism powered and controlled by signals carried through the cable 38 itself or a separate control cable (not shown).

[0010] To remove a bin 10 from the top of the stack 12, the load handling device 30 is first moved in the X and Y directions to position the gripper device 39 above the stack 12. The gripper device 39 is then lowered vertically in the Z direction to engage the bin 10 at the top of the stack 12, as shown in FIGS. 4 and 6B. The gripper device 39 grasps the bin 10 and is then pulled upward by the cable 38 with the bin 10 attached. At the top of its vertical travel, the bin 10 is held on the rail 22 and housed within the vehicle body 32. In this manner, the load handling device 30, carrying the bin 10 therewith, can be moved to different positions in the XY plane to transport the bin 10 to another location. Upon arriving at the target location (e.g., another stack 12, an access point in a storage system, or a conveyor belt), the bin or container 10 can be lowered from the container receiving portion and released from the gripper device 39. The cable 38 is long enough to allow the load handling device 30 to pick and place bins from any level of the stack 12, including, for example, floor level.

[0011] As shown in Figure 3, multiple load handling devices 30 are provided so that each load handling device 30 can operate simultaneously to increase system throughput. The system shown in Figure 3 may include specific locations, known as ports, where bins 10 can be transferred into or out of the system. An additional conveyor system (not shown) is associated with each port so that bins 10 transported to a port by a load handling device 30 can be transferred by that conveyor system to another location, such as a picking station (not shown). Similarly, bins 10 can be moved by the conveyor system from an external location to the port, for example, to a bin filling station (not shown), and transported by the load handling device 30 to the stacks 12 to replenish stock in the system.

[0012] Each load handling device 30 is capable of lifting and moving one bin 10 at a time. The load handling device 30 has a container receiving cavity or recess 40 in its lower portion. The recess 40 is sized to accommodate the container 10 when it is lifted by the lifting mechanisms 38, 39, as shown in Figures 6A and 6B. When in the recess, the container 10 is lifted off the lower rail 22 to allow the vehicle 32 to move laterally to different grid locations.

[0013] When it is necessary to remove a bin 10b that is not at the top of a stack 12 (a "target bin"), the bins 10a above (a "non-target bin") must first be moved to allow access to the target bin 10b. This is accomplished by an operation hereinafter referred to as "digging." Referring to FIG. 3, during a digging operation, one of the load handling devices 30 sequentially lifts each non-target bin 10a from the stack 12 containing the target bin 10b and places it in a vacant position in another stack 12. The target bin 10b can then be accessed by the load handling device 30 and moved to a port for further transport.

[0014] Each load handling device 30 is remotely operable under the control of a central computer, e.g., a master controller. Also, each individual bin 10 in the system is tracked so that the appropriate bin 10 can be removed, transported, and replaced as needed. For example, during a dig operation, each non-target bin location is logged so that the non-target bin 10a can be tracked.

[0015] Wireless communications and networks may be used to provide a communications infrastructure from a master controller, e.g., via one or more base stations, to one or more load handling devices 30 operable on the grid structure 15. In response to receiving instructions from the master controller, a controller in the load handling device 30 is configured to control various drive mechanisms to control movement of the load handling device. For example, the load handling device 30 may be instructed to retrieve a container from a target storage row at a specific location on the grid structure 15. The instructions may include various movements in the XY plane of the grid structure 15. As previously described, upon reaching the target storage row, the lifting mechanisms 38, 39 may be operated to grasp and lift the storage container 10. Once the container 10 is received in the container receiving space 40 of the load handling device 30, the container 10 is then transported to another location on the grid structure 15, e.g., a “drop-off port.” At the drop-off port, the container 10 is lowered to a suitable pick station to enable retrieval of any items in the storage container. Movement of the load handling device 30 on the grid structure 15 may also involve the load handling device 30 being commanded to move to a charging station, typically located on the periphery of the grid structure 15 .

[0016] To move the load handling devices 30 on the grid structure 15, each load handling device 30 is equipped with a motor for driving the wheels 34, 36. The wheels 34, 36 may be driven via one or more belts connected to the wheels or may be individually driven by motors integrated into the wheels. In the case of single-cell load handling devices (where the footprint of the load handling device 30 occupies a single grid cell 17), the motors for driving the wheels may be integrated into the wheels due to the limited availability of space within the vehicle body. For example, the wheels of a single-cell load handling device 30 are driven by respective hub motors. Each hub motor includes an outer rotor with multiple permanent magnets arranged to rotate about a wheel hub with coils forming an inner stator.

[0017] 1-5 has many advantages and is suitable for a wide range of storage and retrieval operations. In particular, it allows for very high density storage of products, providing a very economical way of storing a wide range of different items in the bins 10, while also allowing reasonably economical access to all of the bins 10 when needed for picking.

[0018] Referring to FIG. 6 , the system may further include a robotic picking station 50 mounted on the storage and retrieval structure 1, for example, beside the load handling device 30 (not shown). The robotic picking station 50 includes a number of designated grid cells 60, 62, as well as a robotic manipulator 52 including a robotic arm 54 and an end effector 56 for releasably engaging a product to be manipulated. The end effector 56 may be a suction device 64 connected to a negative pressure source by a vacuum line 66. The robotic manipulator 52 is mounted on a pedestal 58 above a single grid cell 60 and, depending on its location on the structure 1, may be surrounded by up to eight other grid cells 62 as shown in FIG. 6 . Generally, the robotic manipulator 52 is configured to pick up an item or product from any one of the containers in one of the designated grid cells 62 and place it into a container in another of the designated grid cells 62. The load handling device collects containers from and delivers containers to the designated grid cells 62 as needed. In this manner, the robotic picking stations 50 and the load handling devices 30 work in conjunction to fulfill customer orders or redeploy products throughout the storage and retrieval system 1. Summary of the Invention

[0019] A computer-implemented method is provided for detecting identification markers on a picking station in a grid comprising a plurality of grid cells, the grid forming part of a grid-based storage system with one or more picking stations mounted on the grid, each picking station comprising a robotic manipulator for transferring items between containers received in a respective grid cell adjacent the picking station, the method comprising: acquiring image data representing an image portion including a picking station of the one or more picking stations; processing the image data sequentially with a first neural network and a second neural network; wherein the first neural network is trained to detect instances of identification markers on the picking stations in the image and the second neural network is trained to recognize marker information in the image associated with the identification markers; and outputting marker data representing marker information determined by the second neural network.

[0020] Further provided is a data processing apparatus comprising a processor configured to perform the method. Also provided is a computer program comprising instructions that, when executed by a computer, cause the computer to perform the method. Similarly provided is a computer readable storage medium comprising instructions that, when executed by a computer, cause the computer to perform the method.

[0021] Further provided is a detection system for detecting identification markers on picking stations in a grid comprising a plurality of grid cells, the grid forming part of a grid-based storage system with one or more picking stations mounted on the grid, each picking station comprising a robotic manipulator for transferring items between containers received in a respective grid cell adjacent the picking station, the detection system comprising: an image sensor for capturing an image of at least a portion of the grid; an interface for acquiring image data; and one or more processors configured to implement a first neural network trained to detect instances of identification markers on the picking stations in the image; and a second neural network trained to recognize marker information in the scene associated with the identification markers, wherein the detection system is configured to: acquire image data representing an image portion including a picking station of the one or more picking stations; process the image data using the first neural network and the second neural network in series to generate marker data representing marker information present on the identification markers on the picking stations; and output the marker data.

[0022] Broadly speaking, this specification introduces a system and method for detecting identifiers on robotic pick stations installed on the grid of a grid-based automated storage and retrieval system (ASRS) so that the pick station can be distinguished from other pick stations on the grid. The system and method enable independent identification confirmation of robotic pick stations on the grid. For example, if a failure or unresponsive instance of a given pick station is detected, the pick station can be efficiently identified to aid in determining how to resolve its operation and the operation of a broader group of robotic pick stations. For example, service history for an identified pick station can be looked up to potentially diagnose the problem, and / or pick station service, which may involve inspection of the pick station on the grid or removal of the pick station from the grid, can be arranged or confirmed.

[0023] Thus, information regarding the identity and location of affected pick stations can aid in efficient maintenance of transport devices.

[0024] Embodiments will now be described, by way of example only, with reference to the accompanying drawings, in which like reference numerals designate the same or corresponding parts and in which: [Brief explanation of the drawings]

[0025] [Figure 1] Schematic diagram of the automatic storage and retrieval structure. [Figure 2] 2 is a schematic diagram of a plan view of a section of a track structure forming part of the storage structure of FIG. 1; [Figure 3] 2 is a schematic diagram of a plurality of load handling devices moving over the storage structure of FIG. 1; [Figure 4] Schematic of a load handling device interacting with a container. [Figure 5] Schematic of a load handling device interacting with a container. [Figure 6]1 is a schematic diagram of a known robotic picking station. [Figure 7A] 1 is a schematic diagram of a storage system with a camera located above a grid framework structure as part of a detection system, according to an embodiment. [Figure 7B] 1 is a schematic diagram of a storage system with a camera located above a grid framework structure as part of a detection system, according to an embodiment. [Figure 8] 1 is a diagram of a schematic representation of an image captured by a camera positioned above a grid framework structure, according to certain embodiments. [Figure 9] Schematic diagram of a neural network. [Figure 10A] Schematic of the generated model of the track in the grid framework structure. [Figure 10B] Schematic of the generated model of the track in the grid framework structure. [Figure 11] Schematic diagram showing flattening of a captured image of a grid framework structure. [Figure 12] 1 is a schematic diagram illustrating processing of a captured image of a grid framework structure according to an embodiment. [Figure 13] 1 is a flow chart illustrating a method for detecting mobile picking stations on a grid forming part of a grid-based storage system, according to an embodiment. [Figure 14] 10 is a flowchart illustrating a method for detecting identification markers on a picking station that is located on a grid, according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0026] In the following description, some specific details are included to provide a thorough understanding of the disclosed examples. However, those skilled in the art will recognize that other examples may be practiced without one or more of these specific details, or with other components, materials, etc., and that structural changes may be made without departing from the scope of the present invention as defined in the appended claims. Furthermore, references in the following description to any terms having an implied orientation are not intended to be limiting, but merely to refer to the orientation of features as shown in the accompanying drawings. In some examples, well-known features or systems, such as processors, sensors, storage devices, network interfaces, fasteners, electrical connectors, etc., are not shown or described in detail to avoid unnecessarily obscuring the description of the disclosed embodiments.

[0027] Unless the context otherwise requires, throughout this specification and the appended claims, the word "comprise" and variations thereof, such as "comprises" and "comprising," are to be interpreted in their open-ended, inclusive sense as "including, but not limited to."

[0028] References throughout this specification to "one," "an," or "another" as applied to an "embodiment," "example," or "implementation" mean that a particular referent feature, structure, or characteristic described in connection with an embodiment, example, or implementation is included in at least one embodiment, example, or implementation. Thus, the appearances of phrases such as "in one embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments, examples, or implementations.

[0029] It should be noted that, as used in this specification and the appended claims, the user forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. It should also be noted that the word "or" is generally employed in its sense including "and / or" unless the content clearly dictates otherwise.

[0030] 7A and 7B show schematic diagrams of a detection system for detecting mobile picking stations 50 on a grid 15 forming part of a grid-based storage system 1, according to one embodiment. The grid-based storage system 1 is of a type previously described, e.g., an automated storage and retrieval system (or "ASRS"). In this embodiment, there are multiple robotic picking stations 50 mounted on the grid-based storage system 1, e.g., mounted on a grid structure (or simply "grid") 15 as previously described with respect to FIG. 6. Each picking station 50 includes a robotic manipulator 52 for transferring items between containers received in designated grid cells adjacent the respective picking station 50. For example, the robotic manipulator 52 includes an end effector for manipulating and releasably engaging items to be transferred between containers. The end effector may be a suction device connected to a negative pressure source, as in the embodiment shown in FIG. 7B, or another type of end effector, such as a jaw gripper or finger gripper.

[0031] In the embodiment shown in FIG. 7B, each robotic manipulator 52 is mounted on a pedestal above a single grid cell and is surrounded by eight grid cells. In other embodiments, a given robotic manipulator 52 may be surrounded by fewer grid cells or on fewer sides, depending on its location on storage system 1. Similarly, while FIG. 7B shows robotic picking stations 50 arranged along both orthogonal directions of grid 15, in other embodiments, picking stations 50 may be arranged along only one axis of grid 15, for example, in a row or line. In some cases, there may be clusters of robotic picking stations 50 arranged at selected locations on grid 15 of storage system 1.

[0032] A camera 71 is disposed above the grid 15 as part of the detection system. In an example, the camera 71 is an ultra-wide-angle camera, i.e., it has an ultra-wide-angle lens (also referred to as a "super wide-angle" lens or a "fisheye" lens). The camera 71 includes an image sensor for receiving incident light focused through a lens, e.g., a fisheye lens. The camera 71 has a field of view 72 that includes at least a section of the grid 15. Multiple cameras may be used to observe the entire grid 15, e.g., each camera 71 having a respective field of view 72 that covers a section of the grid 15. The ultra-wide-angle lens may be selected because of its relatively large field of view 72, e.g., up to a 180-degree solid angle, compared to other lens types, which means that fewer cameras are needed to cover the grid 15. Also, space may be limited between the top of the grid 15 and surrounding structures, e.g., a warehouse roof, thus constraining the height of the camera 71 above the grid 15. An ultra-wide angle camera can provide a relatively large field of view at a relatively low height above the grid 15 compared to other camera types.

[0033] One or more cameras 71 may be used to monitor grid 15, including robotic picking stations 50. For example, image feeds from one or more cameras 71 may be displayed on one or more computer monitors remote from grid 15 to monitor picking stations 50 (and load handling devices) operating on grid 15.

[0034] The detection system also includes an object detection model trained to detect instances of mobile picking stations on the grid. For example, the object detection model is a trained neural network configured to process images to detect mobile picking stations in the images. Further details regarding neural network and non-neural approaches are described below. The detection system is configured to obtain image data representing a series of images captured by the camera 71 and process the image data with the object detection model. Based on the processing, the detection system can then determine whether the series of images includes a mobile picking station of one or more picking stations. In response to determining that the images include a mobile picking station, the detection system outputs annotation data indicating the mobile picking station in the image. Further details regarding processing images to detect mobile picking stations and scenarios involved in determining the location and even identification (ID) information (e.g., a unique ID label) of the detected picking station 50 are described in the embodiments below.

[0035] Calibration Process A monitoring or surveillance system for grid 15 may incorporate calibration of one or more cameras 71 located above grid 15, particularly in embodiments comprising wide-angle or ultra-wide-angle cameras. Accurate calibration of the (ultra-)wide-angle camera may allow interaction with the image captured by the (ultra-)wide-angle camera, distorted by the (ultra-)wide-angle lens, to be correctly mapped to the workspace. Thus, an operator may select an area of ​​pixels in the distorted image, which is mapped to a corresponding area in grid space, for example. In other scenarios, the (distorted) image from camera 71 may be processed to detect picking stations on the grid and output their corresponding locations on the grid and even identifying information, such as a unique ID label.

[0036] An exemplary calibration process for an ultra-wide-angle camera includes acquiring a section of a grid, i.e., an image of the grid section, captured by the camera. Acquiring the image includes obtaining, e.g., receiving, image data representing the image, e.g., in a processor. For example, the image data may be received via an interface, e.g., a camera serial interface (CSI). An image signal processor (ISP) may perform initial processing of the image data, e.g., saturation correction, renormalization, white balancing, and / or demosaicing, to prepare the image data for display.

[0037] Initial values ​​for several parameters corresponding to the ultra-wide-angle camera are also obtained, including the focal length of the ultra-wide-angle camera, a translation vector representing the position of the ultra-wide-angle camera above the grid section, and a rotation vector representing the tilt and rotation of the ultra-wide-angle camera. These parameters can be used in a mapping algorithm to map pixels in the image distorted by the camera's ultra-wide-angle lens onto a plane oriented relative to the storage system's Cartesian grid 15. The mapping algorithm is described in more detail below.

[0038] The calibration process involves processing images using a neural network trained to detect / predict tracks in images of grid sections captured by an ultra-wide-angle camera. Neural Networks 9 illustrates an example of a neural network architecture. The exemplary neural network 90 is a convolutional neural network (CNN). One example of a CNN is the U-Net architecture developed by the Department of Computer Science at the University of Freiburg, although other CNNs, such as the VGG-16 CNN, can be used. The input 91 to the CNN 90 comprises image data in this example. The input image data 91 is a given number of pixels wide and a given number of pixels high and includes one or more color channels (e.g., red, green, and blue channels).

[0039] The convolutional layers 92, 94 of the CNN 90 may generally extract specific features from the input data 91 and operate on small portions of the image to create feature maps. The fully connected layer 96 uses the feature maps to determine an output 97, e.g., classification data specifying the classes of objects predicted to be present in the input image 91.

[0040] In the example of FIG. 9 , the output of the first convolutional layer 92 undergoes pooling in a pooling layer 93 before being input to a second convolutional layer 94. Pooling, for example, allows values ​​for a region of an image or feature map to be aggregated or combined, e.g., by taking the highest value within the region. For example, in 2×2 max pooling, rather than transferring the entire output, the highest value of the output of the first convolutional layer 92 within a 2×2 pixel patch of the feature map output from the first convolutional layer 92 is used as input to the second convolutional layer 94. Pooling can therefore reduce the amount of computation for subsequent layers of the neural network 90. ​​The effect of pooling is shown schematically in FIG. 9 as a reduction in the size of the frames in the relevant layers. Further pooling is performed in a second pooling layer 95 between the second convolutional layer 94 and the fully connected layer 96. It should be appreciated that the schematic representation of neural network 90 in FIG. 9 is greatly simplified for ease of explanation, and that typical neural networks can be significantly more complex.

[0041] Generally, a neural network, such as the neural network 90 of FIG. 9, may undergo what is called a "training phase," during which the neural network is trained for a specific purpose. A neural network generally includes layers of interconnected artificial neurons that form a directed, weighted graph, where the graph's vertices (corresponding to neurons) or edges (corresponding to connections) are each associated with a weight. The weights may be adjusted throughout training, changing the output of individual neurons and, therefore, the neural network as a whole. In a CNN, a fully connected layer 96 generally connects every neuron in one layer to every neuron in another layer and may thus be used to identify global properties of an image, such as whether the image contains a particular class of object or a particular instance of a particular class.

[0042] In the present context, neural network 90 is trained to perform object identification by processing image data, e.g., to determine whether an object of a predetermined class of objects is present in the image (although in other examples, neural network 90 may instead be trained to identify other image characteristics of the image). For example, training neural network 90 in this manner generates weight data representing weights to be applied to the image data (e.g., different weights associated with different layers of a multi-layer neural network architecture). Each of these weights is multiplied by the corresponding pixel value of the image patch, e.g., to convolve the weight kernel with the image patch.

[0043] Specific to the context of ultra-wide-angle camera calibration, neural network 90 is trained using a training set of input images of grid sections captured by an ultra-wide-angle camera to detect tracks 22 of grid 15 in a given image of the grid section. In an example, the training set includes mask images that show only extracted track features corresponding to the input images. For example, the mask images are manually created. Thus, the mask images can serve as a desired result for neural network 90 to be trained using the training set of images. Once trained, neural network 90 can be used to detect tracks 22 in images of at least a portion of grid structure 15 captured by an ultra-wide-angle camera.

[0044] The calibration process includes processing images of the grid sections captured by the ultra-wide-angle camera 71 with a trained neural network to detect tracks 22 in the images. At least one processor (e.g., a neural network accelerator) may be used to perform the processing. The image processing generates models of the tracks, in particular a first set and a second set of parallel tracks, captured in the images of the grid sections. For example, the models comprise representations of predictions of the tracks in the distorted images of the grid sections determined by the neural network. The track models correspond, in examples, to masks or probability maps.

[0045] Selected pixels in the determined track model are then mapped to corresponding points on the grid 15 using a mapping, e.g., a mapping algorithm, that incorporates multiple parameters corresponding to the ultra-wide-angle camera. The obtained initial values ​​are used as input to the mapping algorithm.

[0046] An error function (or "loss function") is determined based on the discrepancy between the mapped grid coordinates and the true, e.g., known, grid coordinates of the points corresponding to the selected pixels. For example, a selected pixel at the center of the X-direction track 22a should correspond to a grid coordinate with a half-integer value in the Y-direction, e.g., (x, y.5), where x is an unknown number and y is an unknown integer. Similarly, a selected pixel at the center of the Y-direction track 22b should correspond to a grid coordinate with a half-integer value in the X-direction, e.g., (x'.5, y'), where x' is an unknown integer and y' is an unknown number. In an example, the width and length of a grid cell (or their ratio) are used in the loss function, for example, to calculate cell x,y coordinates for keypoints and determine whether they are on the track (e.g., coordinate values ​​of n.5, where n is an integer).

[0047] The initial values ​​of the parameters corresponding to the ultra-wide-angle camera are then updated to updated values ​​based on the determined error function. For example, a Broyden-Fletcher-Goldfarb-Shanno (BFGS) algorithm is applied using the error function and the initial parameter values ​​as input. In an example, the updated values ​​of the parameters are determined iteratively, and the error function is recalculated with each update. The iterations may continue, for example, until the error function is reduced by less than a predetermined threshold between successive iterations or compared to the initial error function, or until the absolute value of the error function falls below a predetermined threshold. Other iterative algorithms, such as sequential quadratic programming (SQP) or sequential least-squares quadratic programming (SLSQP), may be used with the initial values ​​to generate a sequence of improving approximate solutions for the parameters, with a given approximation in the sequence being derived from the previous one. In some cases, the iterative algorithm is used to optimize the values ​​of the parameters. For example, the updated values ​​are optimized values ​​of the parameters.

[0048] Updating the initial values ​​of a plurality of parameters corresponding to the ultra-wide-angle camera involves applying one or more respective boundary values ​​for the plurality of parameters. For example, boundary values ​​for a rotation angle associated with a rotation vector are substantially 0 degrees and substantially +5 degrees. Additionally or alternatively, boundary values ​​for a planar component of a translation vector are ±0.6 of the length of a grid cell. Additionally or alternatively, boundary values ​​for a height component of a translation vector are 1800 mm and 2100 mm, or 1950 mm and 2550 mm, or 2000 mm and 2550 mm above the grid. For example, a lower limit for the camera height is within the range of 1800 to 2000 mm. For example, an upper limit for the camera height is within the range of 2100 to 2600 mm. Additionally or alternatively, boundary values ​​for the camera focal length are 0.23 and 0.26 cm. Applying one or more respective boundary values ​​for a plurality of parameters can mean that the updating, e.g., optimization, process is performed in a feasible region or solution space, i.e., a set of all possible values ​​that satisfy one or more boundary conditions.

[0049] The updated values ​​of the plurality of parameters are electronically stored for future mapping of pixels in a grid section image captured by ultra-wide-angle camera 71 to corresponding points on grid 15 via a mapping algorithm. For example, the stored values ​​of the plurality of parameters are retrieved from data storage and used in a mapping algorithm to calculate grid coordinates corresponding to a given pixel in a given image of the grid section captured by ultra-wide-angle camera 71. In an example, the updated values ​​are stored in a storage location associated with ultra-wide-angle camera 71, for example, in a database. For example, a lookup function or table may be used in conjunction with the database to find the stored parameter values ​​associated with any given ultra-wide-angle camera employed in storage system 1 on grid 15.

[0050] Following calibration of a given camera 71 disposed over the grid 15, an image (e.g., a “snapshot”) of a grid section captured by the camera 71 may be flattened, i.e., undistorted, for interaction by an operator. For example, using the described image-to-grid mapping function, a distorted image 81 of the grid section may be converted into a flattened image 111 of the grid section, as shown in the example of FIG. 11 . Flattening involves selecting an area of ​​grid cells to flatten in the distorted image 81 and inputting the grid coordinates corresponding to those cells into the mapping function, which determines which respective pixel values ​​from the distorted image 81 should be copied into the flattened image 111 for each grid coordinate. For example, a target resolution, in pixels per grid cell, may be set for the flattened image 111, the target resolution having a ratio corresponding to the ratio of the grid cell dimensions. Once all pixel values ​​needed in the flattened image (according to the target resolution and the selected number of grid cells) have been determined, the flattened image 111 can be generated.

[0051] Snapshots may be captured by the camera 71 at predetermined intervals, e.g., every 10 seconds, and converted to corresponding flattened images 111. The most recent flattened images 111 may be stored in storage for viewing on a display, e.g., by an operator wishing to view the grid section covered by the camera 71. The operator may instead choose to re-take a snapshot of the grid section and flatten it. Thus, the operator may select regions, e.g., pixels, in the flattened image 111 and convert those selected regions to grid coordinates based on the image-to-grid mapping functionality described herein. In some cases, the flattened image 111 includes grid coordinate annotations for the grid space viewable in the flattened image 111. The flattened images 111 corresponding to each camera 71 may be more user-friendly for monitoring the grid 15 compared to the distorted image 81. Grid to image mapping A computational algorithm maps real-world points on grid 15 to pixels in an image captured by the camera. The grid points are first projected onto a plane corresponding to ultra-wide-angle camera 71. For example, at least one of a rotation using a rotation matrix and a planar translation in the X and Y directions is applied to points having x, y, and z coordinates in grid framework structure 14. The focal length f of the ultra-wide-angle camera may be used to project points with three-dimensional coordinates relative to grid 15 onto a two-dimensional plane relative to ultra-wide-angle camera 71. For example, the coordinates of a mapped point q in the plane of ultra-wide-angle camera 71 are given by q=f·p [x,y] ÷p z It is calculated as, where p [x,y] and p z are the planar xy coordinates and the third z coordinate of point p relative to grid 15, respectively.

[0052] A point q projected onto the ultra-wide-angle camera plane may be aligned with a Cartesian coordinate system in that plane to determine a first Cartesian coordinate of the point. For example, aligning a point with a Cartesian coordinate system involves rotating the point or its position vector in the plane (e.g., a vector from the origin to the point). Thus, the rotation is, for example, to align with a typical grid orientation in an image captured by the camera, but may not be necessary if the X and Y directions of the grid are already aligned with the captured image. In the example, the rotation is substantially 90 degrees. As shown in FIGS. 8A and 8B, the X and Y directions of the grid are offset by 90 degrees with respect to the horizontal and vertical axes of the image; therefore, the rotation "corrects" for this offset so that the X and Y directions of the grid are aligned with the horizontal and vertical axes of the captured image.

[0053] A grid-to-image mapping algorithm continues by converting the first Cartesian coordinate to a first polar coordinate using standard trigonometry. A distortion model is then applied to the first polar coordinate of the point to generate a second, e.g., "distorted," polar coordinate. In an example, the distortion model comprises a tangent model of distortion given by r' = f arctan(r / f), where r and r' are the undistorted and distorted radial coordinates of the point, respectively, and f is the focal length of the ultra-wide-angle camera.

[0054] The second polar coordinates are then converted back to (second) Cartesian coordinates using the same standard trigonometric methods inversely. Image coordinates of pixels in the image are then determined based on the second Cartesian coordinates. In examples, this determination includes at least one of de-centering or rescaling the second Cartesian coordinates. Additionally or alternatively, the ordinate (y-coordinate) of the second Cartesian coordinates is flipped, e.g., mirrored on the x-axis. Image to grid mapping Mapping pixels in the image captured by camera 71 to real-world points on grid 15 is done by different computational algorithms. For example, the image-to-grid mapping algorithm is the inverse of the grid-to-image mapping algorithm described above, with each mathematical operation reversed.

[0055] For a given pixel in the image, a (second) Cartesian coordinate of the mapped point is determined based on the image coordinate of the pixel in the image. For example, this determination may involve initializing the pixel in the image, including, for example, at least one of centering or normalizing the image coordinate. As mentioned above, the ordinate coordinate is inverted in some instances. The second Cartesian coordinate is converted to a second polar coordinate using standard trigonometry as described above. The use of the label "second" is used for consistency with the conversion performed in the grid-to-image mapping algorithm described, but is arbitrary.

[0056] An inverse distortion model is applied to the second polar coordinates to generate first, e.g., "undistorted," polar coordinates. In an example, the inverse distortion model is based on a tangent model of distortion given by r=f·tan(r' / f), where again r' is the distorted radial coordinate of the point, r is the undistorted radial coordinate of the point, and f is the focal length of the ultra-wide-angle camera. Thus, in an example, the inverse distortion model used in the image-to-grid mapping is the inverse, or "anti-function," of the distortion model used in the grid-to-image mapping.

[0057] The image-to-grid mapping algorithm continues by converting the first polar coordinate to a first Cartesian coordinate. The first Cartesian coordinate may be disaligned or misaligned with a Cartesian coordinate system in a plane corresponding to the ultra-wide-angle camera. For example, disaligning a point with a Cartesian coordinate system involves applying a rotation transformation to the point or its position vector in the plane (e.g., a vector from the origin to the point). The rotation is substantially 90 degrees in the example. This rotation may therefore "undo" any "correction" to the offset between the X and Y directions of the grid and the horizontal and vertical axes of the captured image, as previously described in the grid-to-image mapping.

[0058] Finally, the point is projected from the (second) plane corresponding to the camera 71 onto the (first) plane corresponding to the grid 15 to determine the grid coordinates of the point relative to the grid.

[0059] In the example, the projection of the point onto the plane corresponding to grid 15 is p=B -1 This involves calculating (f tq z), where B=q R 3,[1,2] -f·R [1,2],[1,2] In these equations, p comprises the point coordinate in the grid plane, q comprises the Cartesian coordinate in the camera plane, and f is the focal length of the ultra-wide-angle camera as described above. Furthermore, t is a planar translation vector, z is the distance (e.g., height) between the ultra-wide-angle camera and the grid, and R is a three-dimensional rotation matrix related to the rotation vector. The rotation vector comprises a direction representing the axis of rotation and a magnitude representing the angle of rotation. The rotation matrix R corresponding to the angle-axis rotation vector can be determined from the vector using, for example, Rodriguez's rotation formula.

[0060] Next, a mathematical derivation of the function for projecting an undistorted 2D point q from the camera plane is provided for completeness: starting from the projection from the grid onto the image from above, q = f p' [x,y] ÷p' zwhere p' is the rotated and translated grid point p: p'=R·p+(t x ,t y ,z) T The goal is to derive p from q. After rearranging and substituting for p', we get the following:

[0061]

number

[0062] Since the desired distance of a point p on the grid from the camera is given by the height parameter z, in translating the point p z = 0. Therefore, for all p z Terms can be eliminated to give:

[0063]

number

[0064] Matrix B=(q·R 3,[1,2] By defining -f·R), the equation becomes B·p [x,y] = f tz q, which can be further simplified to -1 Using

[0065] Returning to the calibration process, in some cases, grid cell coordinate data encoded in grid cell markers positioned with respect to the grid 15 may be used to calibrate calculated grid coordinates corresponding to pixels in the captured image. For example, the grid cell markers may be signboards, e.g., placed in predetermined grid cells 17, with corresponding cell coordinate data marked on each signboard. The process may involve, for example, processing the captured image to detect the grid cell markers in the image and then extracting the grid cell coordinate data encoded in the grid cell markers for use in calibrating the mapped grid coordinates. Each grid cell marker is in a respective grid cell, e.g., under and within the field of view 72 of a respective camera 71.

[0066] The image processing may involve using an object detection model, e.g., a neural network, trained to detect instances of grid cell markers in images of the grid section. A computer vision platform, e.g., Cloud Vision API (Application Programming Interface) by Google®, may be used to implement the object detection model. The object detection model may be trained using images of the grid section including the grid cell markers. In examples where the object detection model includes a neural network, e.g., a CNN, the description with reference to FIG. 9 applies accordingly.

[0067] Grid coordinates generated by mapping pixels in an image to points on a grid section represented in a captured image can be calibrated to the entire grid based on the extracted cell coordinate data. For example, a mapped grid point corresponding to a given pixel has coordinates in units of grid cells, e.g., (x, y), with x being the number of grid cells in the X direction and y being the number of grid cells in the Y direction. However, the grid cells captured by camera 71 are for a grid section, i.e., a section of grid 15, and therefore not necessarily the entire grid 15. Therefore, the mapped grid coordinates (x, y) for a grid section captured in an image can be calibrated to grid coordinates (x', y') for the entire grid based on the relative location of the grid section with respect to the entire grid. The location of the grid section with respect to the entire grid can be determined by extracting the grid cell coordinate data encoded in the grid cell markers captured in the image, as described.

[0068] 10A shows an exemplary model 101 of tracks generated by processing an image 81 of a grid section captured by an ultra-wide-angle camera 71 with a neural network 90 trained to detect tracks 22 in the image. Model 101 comprises a representation of a prediction of tracks 22 a, 22 b in a distorted image of the grid section determined by neural network 90. ​​Mapping pixels from track model 101 to corresponding points on grid 15 may be performed to calibrate camera 71 as described. For example, calibration may involve updating, e.g., optimizing, multiple parameters associated with camera 71 used for mapping between pixels in captured image 81 and points on grid 15.

[0069] In an example, the model 101 of the grid sections may be refined to represent only the centerlines of the first set 22a and the second set 22b of parallel tracks. Thus, the pixels to be mapped from the track model 101 to corresponding points on the grid 15 are, for example, pixels that lie on the centerlines of the first set 22a and the second set 22b of parallel tracks in the generated model 101. Refining involves, for example, filtering the model with horizontal and vertical line detection kernels. These kernels allow the centerlines of the tracks to be identified in the model 101, for example, in the same way that other kernels may be used to identify other features of an image, such as edges in edge detection. Each kernel is a given size, for example a 3x3 matrix, that may be convolved with the image data in the model 101 with a given stride. For example, the horizontal line detection kernel may be the matrix

[0070]

number

[0071] It can be expressed as:

[0072] Similarly, the vertical line detection kernel can be, for example, the matrix

[0073]

number

[0074] It can be expressed as:

[0075] In the example, filtering involves at least one of eroding and dilating pixel values ​​of the model 101 using horizontal and vertical line detection kernels. For example, at least one of an erosion function and a dilation function is applied to the model 101 using the kernel. The erosion function effectively "erodes" foreground objects, in this case, the boundaries of tracks 22 a, 22 b in the generated model 101, by convolving the kernel with the model. During erosion, a pixel value (either "1" or "0") in the original model is updated to a value of "1" only if all pixels convolved under the kernel are equal to "1"; otherwise, it is eroded (updated to a value of "0"). Effectively, all pixels near the boundaries of tracks 22 a, 22 b in the model 101 are discarded, depending on the size of the kernel used in the erosion, such that the thickness of each of tracks 22 a, 22 b is reduced substantially to its centerline. A dilation function, which is the opposite of the erosion function, may be applied after erosion to effectively "inflate" or widen the centerlines remaining after erosion. This dilation may stabilize the centerlines of the tracks 22a, 22b in the improved model 101. During dilation, if at least one pixel convolved under the kernel is equal to "1," the pixel value is updated to a value of "1." The erosion and dilation functions are each applied to the original generated model 101, and the resulting horizontal and vertical centerline "skeleton," for example, are combined to produce the improved model.

[0076] In some cases, the generated model 101 may have missing sections of the tracks 22a, 22b, for example, obscuring one or more areas of the grid section viewable by the camera 71. Objects on the grid 15, such as a conveying device 30, a pillar, or other structure, may obscure portions of the tracks in the captured image. Thus, the generated model 101 may have the same missing areas of the tracks. Similarly, false positive predictions of the tracks may be present in the generated model 101.

[0077] To assist with these problems, the tracks 22a, 22b (e.g., their centerlines) present in the generated model can be fitted to respective quadratic equations, for example, to create a secondary track for tracks 22a, 22b. FIG. 10B shows an example of a track among the first set 22a of tracks in model 101 being fitted to the first quadratic track 102 and a track among the second set 22b of tracks in model 101 being fitted to the second quadratic track 103. Then, the quadratic track centerline can be created based on the quadratic track by extrapolating pixel values along the quadratic track, for example, to fill gaps in model 101 or remove false positives. For example, if a sub-line generated from the predicted grid model 101 cannot be fitted to a given quadratic curve along with at least one other line, that sub-line is very likely not to be part of the grid and should be excluded.

[0078] The quadratic equation, y = ax 2 + bx + c used to fit the tracks in model 101 can also have specified boundary conditions, for example,

[0079]

Number

[0080] 、 -9.9×10 -4 <a < 9.9×10 -4 、 -5 < b < 5, and 0 < c < 3200.

[0081] In an example, a predetermined number of pixels are extracted from the improved model 101 of the track, for example, to reduce the storage requirements for storing the model. For example, a random subset of pixels is extracted to give the final improved model 101 of the track.

[0082] Calibrating the ultra-wide-angle camera 71 using the systems and methods described herein allows, for example, images captured by the camera 71 with a wide field of view of the grid 15 to be used to detect and locate a carrier device thereon, despite the relatively high distortion present in the image compared to other camera types.

[0083] The automated calibration process outlined above can also reduce the time it takes to calibrate each camera 71 installed on the grid 15 of a storage system compared to manual methods of tuning parameters associated with each camera 71. For example, combining a neural network model, e.g., U-Net, with a customized optimizer to implement the described calibration pipeline can eliminate over 80% of the error compared to standard calibration methods. Furthermore, the calibration systems and methods described herein have been found to be versatile and consistent enough to calibrate cameras in multiple warehousing storage systems, e.g., having different dimensions, scales, and layouts.

[0084] Additionally, the output flattened calibrated image 111 of the grid allows for easier interaction with the image 111 by both humans and machines to monitor the grid 15 and the picking stations 50 and transport devices 30 moving thereon.

[0085] Detecting a mobile picking station Provided herein are methods and systems for processing images, e.g., distorted images 81, captured by one or more cameras 71 to detect mobile picking stations 50 on a grid 15 of a grid-based storage system 1. For example, the location of the detected picking stations 50 relative to the grid 15 may be output. In some examples, identification (ID) information, e.g., a unique ID label, of the detected picking stations 50 may be output. Such examples are now described in more detail.

[0086] 13 shows a computer-implemented method 130 for detecting mobile picking stations 50 on grid 15. Method 130 involves obtaining 131 image data representing a series of images of at least a portion of grid 15 and processing 132 the image data with an object detection model trained to detect instances of picking stations on the grid. For example, images are captured by a camera 71 with a field of view 72 covering at least a portion of grid 15, and the image data is transferred to a computer for implementing detection method 130. The image data is received, for example, at an interface of the computer, e.g., a CSI.

[0087] The object detection model may be a neural network, e.g., a convolutional neural network, trained to perform object detection of picking stations 50 on the workspace grid 15. Accordingly, the neural network description with respect to FIG. 9 applies in these specific examples. In the present context, the object detection model, e.g., a CNN 90, is trained to perform object identification by processing acquired image data to determine whether an object (i.e., a picking station) of a predetermined class of objects is present in the image. Training the neural network 90 involves, for example, providing the neural network 90 with training images of the workspace section in which the picking station is located. Weight data is generated for each (convolutional) layer 92, 94 of the multi-layer neural network architecture and stored for use in implementing the trained neural network. In the example, the object detection model comprises a “You Only Look Once” (YOLO) object detection model, e.g., YOLOv4 or Scaled-YOLOv4, having a CNN-based architecture. Other exemplary object detection models include neural-based approaches, such as RetinaNet or R-CNN (Regions with CNN features), and non-neural approaches, such as support vector machines (SVMs) for object classification based on determined features, e.g., Haar-like features or Histogram of Oriented Gradients (HOG) features.

[0088] Method 130 involves determining 133 whether the image includes a mobile picking station 50 based on process 132. For example, an object detection model is configured, e.g., trained or learned, to detect whether a mobile picking station 50 is present in the series of captured images of grid 15. In an example, the object detection model makes determination 133 using a confidence level, e.g., a probability score, corresponding to the likelihood that the image includes a mobile picking station 50. Thus, a positive determination may correspond to a confidence level above a predetermined threshold, e.g., 90% or 95%. In response to determining 133 that the image includes a mobile picking station, annotation data (e.g., predicted data or inferred data) indicating the predicted picking station in the image is output 134. An updated version of one or more images in the series of images including the annotation data may be output, for example, as part of method 130.

[0089] In an example, an object detection model may receive as input a series of grid images and be trained with an additional time dimension to detect mobile picking stations in the series of images, e.g., an object detection model (such as a CNN) may be trained based on a dataset of multiple image series comprising mobile and stationary picking stations on a grid.

[0090] In another example, an object detection model is configured, e.g., trained, to detect instances of the picking station 50 in each image of a sequence of images. The trajectory model may be employed to determine varying object detection model parameters. Alternatively, the object detection model is configured to detect the picking station and its movement in a single object trajectory parameterized model, e.g., from a sequence of sets of detected edges indicating edge movement in a sequence of image frames.

[0091] Generally, moving object detection involves segmenting a non-stationary object of interest with respect to a surrounding area or region from a given sequence of images (e.g., video frames). Thus, as described, moving object detection and any further object tracking involves detecting a foreground moving object in every frame in an image sequence (e.g., video) or at the first instance of the moving object.

[0092] Detecting foreground objects, i.e., mobile picking stations 50, on grid 15 may involve a background subtraction method, in which a background model is initialized before a given frame and the difference between the given frame and the background model is obtained by pixel-by-pixel comparison of a color map of the given frame and the background model. For example, if the difference between each pixel in the color map is greater than a predetermined threshold, the corresponding pixel in the given frame is deemed to belong to the foreground. Exemplary background subtraction techniques include image variation matching, intrinsic background, mixture of Gaussians, kernel density estimation (KDE), moving Gaussian mean, sequential kernel density approximation, and temporal median filter.

[0093] In an example, the object detection model is configured to identify the presence of a mobile picking station by frame differencing. Frame differencing involves calculating the difference between at least two image frames, e.g., consecutive frames, in a sequence of images. For example, an image subtraction operator may be used to obtain an output image by subtracting a second image frame from a first image frame in corresponding frames of the sequence of images.

[0094] In some cases, frame differencing is determined on a pixel-by-pixel basis, e.g., the difference is calculated in a pixel-by-pixel manner. For example, temporal differencing involves detecting moving objects by employing a pixel-by-pixel differencing method between at least two consecutive frames.

[0095] An alternative "optical flow" approach to moving object detection involves computing the optical flow field of an image (or video frame). Clustering can be performed based on the optical flow distribution information obtained from the image.

[0096] In an example, annotation data output as part of the detection method comprises bounding box data. FIG. 12 shows an example of an updated version 83 of an image captured in a sequence of images by camera 71, annotated with a bounding box 120 based on the bounding box data. The bounding box 120 corresponds to a picking station 50 detected by an object detection model. A given bounding box may comprise, for example, a rectangle surrounding a detected object and may specify one or more of an image location, an identified class (e.g., a picking station), and a confidence score (e.g., how likely the object will be present within the box). The bounding box data defining a given bounding box may include coordinates of two corners of the box or center coordinates with width and height parameters for the box in image 83. In an example, the detection method 130 involves generating annotation data for output that can be represented, for example, as a bounding box 120.

[0097] In some cases, the object detection model is further trained to detect instances of faulty picking stations in the workspace, for example, picking stations that are unresponsive to communications from the master controller and / or have activated warning signals.

[0098] For example, method 130 may involve processing image data with an object detection model and, based on the processing, determining whether the image includes a picking station with an activated warning signal. The picking station's warning signal comprises a predetermined light, or color of light, emitted by a light source on the picking station, such as a light-emitting diode (LED). For example, the picking station includes an LED configured to emit a first wavelength (color) of light when responsive to communication from the master controller and to emit a second, different wavelength (color) of light when unresponsive to communication from the master controller. The picking station may become unresponsive, e.g., activate a warning signal, when communication with the master controller is lost. Other types of warning signals from the light source are possible, e.g., predetermined patterns of emission, such as blinking. In response to determining that the image includes an unresponsive picking station, method 130 may include outputting at least one of annotation data or an alert. The annotation data indicates a predicted picking station in the image with an activated warning signal. For example, the annotation data may comprise a bounding box in the image that surrounds a predicted picking station with an activated warning signal. Similarly, the output alert may signal that the image contains a picking station with an activated warning signal. Examples of output alerts include, for example, a text or other visual message to be displayed on a screen for viewing by an operator. Picking station localisation The detection method, in some examples, may involve determining the location of the detected picking station on a grid. For example, the method may involve acquiring a target image portion including at least a portion of the mobile picking station and mapping the target image portion to a target location on the grid. The location of the mobile picking station on the grid may then be determined based on the mapped target location. In some cases, the target image portion (of a given image in the sequence of images) is determined based on annotation data acquired as part of the detection method.

[0099] In an example, mapping a target image portion (e.g., one or more pixels in an image) to a target location (e.g., a point on a grid structure) involves inverting a distortion of the image of the workspace. For example, if an image sensor is used in combination with an ultra-wide-angle lens, the lens distorts the view of the workspace. Therefore, the distortion is reversed, e.g., as part of the mapping between image pixels and grid points. An inverse distortion model may be applied to the target image portion for this purpose. The description of the image-to-grid mapping algorithm in the previous example applies correspondingly here. For example, mapping the target image portion to a target grid location involves applying an image-to-grid mapping algorithm described herein.

[0100] The target image portion may be acquired via an object detection system configured to detect the mobile picking station in a series of images of the grid. For example, the process involves the object detection system acquiring images of the grid captured by one or more image sensors and using a motion detection algorithm to determine that the mobile picking station is present in the image data.

[0101] In another example, a user viewing the image representation of grid 15 selects a target image portion via an interface configured to acquire the target image portion. The interface may be, for example, a user interface with which the user interacts. The user interface may include a display screen for displaying the image representation of the workspace captured by the image sensor. The user interface may also include input means, for example, a touchscreen display, a keyboard, a mouse, or other suitable means, with which the user can select the target image portion.

[0102] In an example, the target image portion includes at least a portion of a picking station 50 located on grid 15. For example, the target image portion is a subset of one or more pixels selected from an image of grid 15 captured by an image sensor. The one or more pixels correspond to at least a portion of a picking station detected as moving that is depicted in the image of grid 15. For example, the target image portion includes the entire picking station 50 depicted in the image. In another example, the target image portion is only a single pixel that corresponds to a portion of the picking station depicted in the image.

[0103] In some examples, the target image portion corresponds to a given grid cell in grid 15 to which the picking station is mounted, e.g., on a pedestal. For example, the target image portion is a subset of one or more pixels that corresponds to at least a portion of a given cell. In some cases, the target image portion includes the entire cell, and in other cases, the target image portion is only a single pixel that corresponds to a portion of a cell.

[0104] In an example, determining the location of the detected mobile picking station 50 on the grid as part of the detection method 130 includes generating further annotation data corresponding to a plurality of virtual picking stations in respective grid spaces in the captured image of the grid. The location of the detected picking station 50 on the grid 15 may be determined by comparing the annotation data indicating the detected picking station in the image with the further annotation data corresponding to the plurality of virtual picking stations. For example, the comparison may include calculating an Intersection of Union (IoU) value based on the annotation data. The grid space corresponding to the further annotation data associated with the highest IoU value may then be selected as the grid location of the detected picking station.

[0105] In an example, the further annotation data comprises bounding box data corresponding to multiple bounding boxes associated with multiple virtual picking stations. Thus, calculating the IoU value may involve dividing the area of ​​the overlap, or “intersection,” between two bounding boxes by the area of ​​the union of the two bounding boxes (e.g., the total area covered by the two boxes). For example, the area overlap between the bounding box of the detected picking station and a given bounding box corresponding to a given virtual picking station is calculated and divided by the area of ​​the union for the same two bounding boxes. This calculation is repeated for the bounding box of the detected picking station and each bounding box corresponding to each virtual picking station to provide a set of IoU values. The highest IoU value in the set of IoU values ​​may then be selected, and the grid location of the corresponding bounding box is inferred as the grid location of the detected picking station.

[0106] The detection system described previously may be configured to implement any of the detection methods described herein. For example, the detection system includes an image sensor for capturing an image of at least a portion of the grid and an interface for acquiring image data. The detection system includes a trained object detection model, implemented, for example, on a graphics processing unit (GPU) or a dedicated neural processing unit (NPU), for performing the processing and decision steps of the computer-implemented method 130 for detecting mobile picking stations 50.

[0107] The described systems and methods for detecting a mobile picking station 50 on the grid 15 of a grid-based ASRS allow the mobile picking station to be distinguished from other, e.g., stationary, picking stations on the grid 15. For example, a given camera 71 in an array of cameras with a view of the grid 15 may have multiple picking stations 50 in its view. Thus, detecting which picking station 50 is moving allows that picking station to be distinguished from other stations, e.g., to confirm that it is the correct station being moved pursuant to a test or inspection. Motion detection can be correlated with other information, e.g., a picking station identifier (as described in other examples herein), as further confirmation that the mobile picking station is considered to be moving.

[0108] Further, in response to determining that the image feed captured by the camera 71 above the grid 15 includes the mobile picking station 50, it may be determined whether one or more persons are in proximity to the mobile picking station 50. For example, the image data may be processed by a further object detection model trained to detect human instances to determine whether one or more persons are near the mobile picking station 50 based on the processing. The proximity determination may be based on a distance threshold, e.g., a positive determination is made if a person is detected on the grid within a predetermined distance of the mobile picking station, e.g., on the grid cell in which the mobile picking station is located. Additionally or alternatively, the proximity determination may be made based on whether one or more persons are in any of the designated grid cells adjacent to, e.g., surrounding, the mobile picking station. In response to a positive determination that one or more persons are in proximity to the mobile picking station, the detection system may stop, e.g., turn off, the mobile picking station 50.

[0109] In a further example, the detection system may cause an exclusion zone to be established in response to detecting a moving picking station. For example, if a selected number of picking stations on the grid are targeted to be stopped, e.g., to allow a transport device to utilize designated grid cells adjacent to the selected picking stations, the detection method may be used to determine whether any of the selected picking stations are still active. In response to a positive determination, an exclusion zone corresponding to the designated grid cells adjacent to the detected picking station is determined, and may be established such that transport devices are prohibited from entering the exclusion zone. Similarly, the detection method may be used to detect unscheduled movement of a picking station 50 on the grid 15 in response to which exclusion zone data is determined to represent an exclusion zone that may be implemented around the picking station 50.

[0110] In other examples, an exclusion zone may be determined in response to detecting a faulty picking station, e.g., a picking station that is unresponsive to communications from the master controller and / or has an activated warning signal. For example, an exclusion zone corresponding to a designated grid cell adjacent to the detected picking station may be determined and then implemented such that transport devices are prohibited from entering the exclusion zone while the picking station is faulty, e.g., unresponsive.

[0111] In the described example in which an exclusion zone is determined, the determined exclusion zone may comprise a discrete number of grid spaces. For example, the exclusion zone may be determined to extend into each grid space adjacent to the grid space of the detected picking station 50 such that other transport devices are prohibited from entering those grid spaces. For example, therefore, collisions with picking stations by transport devices moving on the grid 15 may be prevented. In some cases, the exclusion zone may be set as an area of ​​grid cells, e.g., a 5×5 cell area, centered on the grid cell in which the defective picking station 50 is located. Thus, the exclusion zone includes a buffer area around the affected picking station 50. The size of the buffer area may be predetermined, for example, as a set area of ​​grid cells to be applied around the determined grid cell of the detected picking station 50. Additionally or alternatively, the size of the buffer area is a selectable parameter when implementing the exclusion zone in a control system.

[0112] A control system, e.g., a master controller, remotely controlling the movement of transport devices 30 operating on the grid 15 can implement the exclusion zone based on the exclusion zone data output as part of the detection process. For example, each of the one or more transport devices can be remotely operated under the control of a control system, e.g., a central computer. To control the movement of one or more transport devices 30 on the grid 15, instructions can be sent from the control system to one or more transport devices 30 via a wireless communication network, e.g., implementing one or more base stations. A separate controller in each transport device 30 is configured to control various drive mechanisms of the transport device, e.g., vehicle 32, to control its movement. For example, the instructions include various movements in the XY plane of the grid structure 15, which can be encapsulated in a defined trajectory for a given transport device. Thus, a given exclusion zone can be implemented by the central control system, e.g., a master controller, such that the defined trajectory avoids the exclusion zone represented by the exclusion zone data. For example, when an exclusion zone is implemented, one or more respective trajectories corresponding to one or more transport devices 30 on the grid are updated to avoid the exclusion zone.

[0113] Detecting identification markers on picking stations 14 shows a computer-implemented method 140 for detecting identification markers on picking stations in a grid 15. The method involves acquiring 141 image data representing an image portion including the picking station. The image portion may be, for example, a portion, e.g., at least a portion, of an image 81, 83 captured by a camera 71 positioned above the grid as shown in FIGS. 7A and 7B.

[0114] 12 , an exemplary image portion 121 including a detected picking station 50 is shown. The image portion 121 may be extracted from the image 83 based on annotation data, e.g., represented as a bounding box 120, corresponding to the detected picking station 50 in the image 83. For example, output annotation data of the method 130 for detecting mobile picking stations on the grid 15 is used to obtain, e.g., extract, the image portion 121 from the image 83. If the annotation data represents one or more bounding boxes, for example, one or more image portions 121 corresponding to the image data contained in the one or more bounding boxes 120 overlaid on the image 83 are extracted from the image 83. For example, the detection method 140 involves obtaining annotated image data 83 including annotation data 120 indicating one or more picking stations in the image and cropping the annotated image data 83 to produce one or more image portions 121 including the respective one or more picking stations 50.

[0115] In the alternative, the image portion comprises the entire image 83 including one or more picking stations 50 captured by the camera 71. In other words, the image portion comprises, for example, at least a portion of the image 83 captured by the camera 71.

[0116] The detection method 140 further involves processing 142 the acquired image data with a first neural network and a second neural network in series. The first neural network is trained to detect instances of identification markers on the picking station in the image. The second neural network is trained to recognize marker information associated with the identification marker in the image. An identification (“ID”) marker is, for example, a text label or other code (such as a barcode, QR code, or the like) on the picking station. The ID marker includes marker information, e.g., text or a QR code, associated with that marker. The marker information corresponds to ID information for the picking station, e.g., corresponds to, e.g., a name or other descriptor of the picking station in a broader system. The marker information is encoded in the ID marker, e.g., as text or other code, and the corresponding ID information can be used to distinguish a given picking station from other picking stations operating in the system.

[0117] In an example, a first neural network is configured, e.g., trained or learned, to receive image portions as first input data and produce feature vectors as intermediate data, e.g., for passing as input to a second neural network. For example, the first neural network comprises a CNN90 configured to use convolutions to extract visual features, e.g., of different sizes, and produce feature vectors. An Efficient and Accurate Scene Text (EAST) detector may be used as the first neural network to identify instances of identifying markers, e.g., text labels, on the picking station.

[0118] In some cases, the first neural network outputs further annotation data, e.g., defining a bounding box, corresponding to the detected identification marker in the image portion. For example, image processing 142 involves determining whether the image portion includes an identification marker on a picking station based on processing with the first neural network. If the determination is positive, further annotation data corresponding to the location of the identification marker in the image portion is generated and output as part of method 140. The further annotation data may comprise image coordinates for image 83 or image portion 121. For example, the image coordinates correspond to at least two corners of a bounding box for the identification marker in the image portion. The bounding box may be defined, for example, by the coordinates of two opposite corners.

[0119] In an example, image processing 142 involves extracting a sub-portion of the image portion, the sub-portion corresponding to a detected identification marker on picking station 50. For example, the image portion is cropped to generate a sub-portion that includes the identification marker. FIG. 12 shows an example sub-portion 122 extracted from image portion 121 and corresponding to a detected identification marker on picking station 50. Sub-portion 122 may be rotated such that a longitudinal axis of the identification marker is substantially horizontal with respect to sub-portion 122, as shown in the example of FIG. 12. Method 140 may then include processing sub-portion 122 with a second neural network configured, e.g., trained or learned, to recognize marker information in the image.

[0120] The detection method 140 concludes with outputting 143 marker data representing the marker information determined by the second neural network. For example, the second neural network is configured to derive marker data from an image sub-portion including an ID marker. In an example where the ID marker comprises a text label, the second neural network may be configured to transcribe the image sub-portion including the label into marker data comprising label sequence data, e.g., a sequence (or "string") of letters, numbers, punctuation, or other symbols. For the example sub-portion 122 shown in FIG. 12, the second neural network would output the marker data as, e.g., label sequence data "AA-Z82" for the identification label of picking station 50. In an alternative example, the ID marker is a code on the picking station, e.g., a QR ("quick response") code or barcode, applied to the picking station, e.g., on a label. The second neural network is configured, e.g., trained or learned, to determine the code from an image of the ID marker on the picking station. The code, e.g., marker data, may then be output. For example, the code may be further processed to decode the ID information encoded therein. In other words, the detected QR code or barcode is decoded to determine, for example, the ID information of the picking station, such as its name.

[0121] In an example, the second neural network comprises a convolutional recurrent neural network (CRNN) configured to apply convolutions to extract visual features from image subportions and arrange the features in a sequence. The CRNN comprises two neural networks, e.g., a CNN and an additional neural network. In some cases, the second neural network includes a bidirectional recurrent neural network (RNN), e.g., a bidirectional long short-term memory (LSTM) model. For example, the bidirectional RNN is configured to process the feature sequence output of the CNN to predict the ID sequence encoded in the marker by applying sequential cues learned from patterns in the feature sequence, e.g., in the example of FIG. 12, that the ID sequence is highly likely to start with the letter "A" and end with a digit. Thus, the second neural network may comprise a pipeline of two or more neural networks, e.g., a CNN piped into a deep bidirectional LSTM, such that the feature sequence output of the CNN is passed to the biLSTM that receives it as input. In other examples, the second neural network comprises a different type of deep learning architecture, for example, a deep neural network.

[0122] The detection system described previously may be configured to implement any of the detection methods described herein. For example, the detection system includes an image sensor for capturing an image of at least a portion of the grid and an interface for acquiring the image data. The detection system includes a trained object detection model, implemented, for example, on a graphics processing unit (GPU) or a dedicated neural processing unit (NPU), for performing the processing and decision steps of the computer-implemented method 140 for detecting identification markers on the picking station 50.

[0123] The described systems and methods for detecting identification markers on picking stations 50 on a grid 15 enable the picking stations to be distinguished from one another, for example, in a captured image of a grid 15 that includes multiple on-grid picking stations 50.

[0124] Furthermore, automatic detection of identification markers on a picking station means that a camera 71 having a view of a selected pick station can be identified from multiple cameras 71 located above the grid 15, each having a different view 72 of the grid 15 and therefore different picking stations 50. For example, if there is known to be a problem with a particular picking station 50, a detection system can be used to find the one or more cameras 71 having the particular picking station 50 in its field of view so that a video feed from said one or more cameras 71 can be displayed to an operator. Such a system or method might, for example, obtain ID information (e.g., an identifier) ​​for a given picking station on the grid 15 and obtain multiple images of the grid 15 captured by each of the multiple cameras mounted on the grid 15. Image data representing the multiple images is processed (e.g., as described above using first and second neural networks) to determine marker information associated with the identification markers on the picking station. The obtained ID information for a given picking station is compared to the determined marker information (e.g., a set of identifiers recognized on the picking station in the images) to determine one or more cameras 71 with a view of the given picking station 50 on the grid 15, e.g., a field of view that includes at least a portion of the given picking station 50. In some examples where more than one camera 71 is identified as having a view of the given picking station 50, it may be determined which camera 71 has a field of view with the largest portion of the picking station 50, e.g., the image feed of the camera 71 with the majority of the picking station 50 in its field of view is selected for display to the operator. An operator viewing images from the cameras 71 can, for example, set one or more exclusion zones in the grid 15 to avoid collisions between the transport device 30 and the picking station 50.

[0125] In some examples, it is determined which cameras 71 have which picking stations 50 in their field of view. For example, by recognizing marker information contained in ID markers on the picking stations, the detection system can associate the marker information extracted from the image with the camera that captured the image. Thus, a record of which picking stations are viewable by which cameras can be determined and stored for future lookup.

[0126] In a further embodiment, the described detection of an identification (ID) marker on a picking station occurs in response to the described detection of a (mobile) picking station. For example, the image portion acquired as part of ID detection method 140 may be determined based on annotation data output as part of detection method 130 for detecting a mobile picking station, as described in the example above.

[0127] The above examples should be understood as illustrative examples. Further examples are contemplated. For example, camera 71 disposed above grid 15 has been described in many examples as an ultra-wide-angle camera. However, camera 71 could be a wide-angle camera, which includes a wide-angle lens that has a relatively longer focal length than an ultra-wide-angle lens, but still introduces distortion compared to a normal lens that reproduces a field of view that appears "natural" to a human observer.

[0128] Similarly, additional techniques for moving object detection are envisioned for application to mobile picking stations. For example, the Canny edge detection algorithm can be combined with a multi-frame subtraction technique to obtain more complete information about moving objects. Alternatively, the use of a combined version of a multi-image subtraction algorithm and a background subtraction algorithm is envisioned to provide a more complete outline of the moving object.

[0129] Further, in a described example involving detecting ID markers on a picking station, the image data is processed using a first neural network and a second neural network in series. However, in an alternative example, the first neural network and the second neural network are merged in an end-to-end ID marker detection pipeline or architecture, e.g., as a single neural network. For example, a method of detecting identification markers on a picking station on a grid comprising a plurality of grid cells is also provided, the grid forming part of a grid-based storage system with one or more picking stations mounted on the grid, each picking station comprising a robotic manipulator for transferring items between containers received in a respective grid cell adjacent to the picking station. The method comprises obtaining image data representing an image portion including a picking station of the one or more picking stations; processing the image data with at least one neural network trained to detect instances of identification markers on the picking station in the image and to recognize marker information associated with the identification marker in the image; and outputting marker data representing the marker information determined by the at least one neural network. In the context of the detection system according to this alternative, the one or more processors are configured to implement at least one neural network trained to detect instances of identification markers on the picking station in the image and to recognize marker information associated with the identification markers in the image.The detection system is configured to process the acquired image data using at least one neural network to generate marker data (e.g., text strings or codes) representing marker information (e.g., picking station descriptors) present on (e.g., encoded in) the identification markers on the picking stations, and to output the marker data.

[0130] Further, in the described example of picking station location identification, the location of the detected picking station 50 on the grid 15 can be determined by comparing annotation data indicating the detected picking station in the image with further annotation data corresponding to multiple virtual picking stations. In an alternative example, the location of the detected picking station 50 on the grid 15 can be determined in a two-step process. First, the multiple virtual picking stations are filtered, which includes, for example, calculating an Intersection of Union (IoU) value using the predicted / inferred data of the picking station 50 and the annotation data of all possible locations of the multiple virtual picking stations. For example, virtual picking stations with calculated IoU values ​​smaller than a predetermined threshold are filtered out. Second, all remaining virtual picking stations are sorted (e.g., in ascending order) by closest distance to the center of the camera's field of view, and the first (e.g., closest) virtual picking station is taken as the mapping. The grid location of the picking station 50 is then set to, for example, the grid location from which the annotation data of the mapped virtual picking station is created.

[0131] Further embodiments are also envisioned in which collisions between transport devices 30 (e.g., bots) and robotic picking stations 50 are detected based on images captured by a camera having a view of the grid 15. For example, image data representing a series of images (e.g., video data) of at least a portion of the grid is acquired and processed by an object detection model trained to detect instances of the transport device colliding with a picking station on the grid. Based on the processing, it is determined whether the series of images includes the transport device colliding with one or more picking stations. The object detection model generally comprises a neural network in the example described with reference to FIG. 9 , which is to be construed as applying accordingly. For example, the object detection model is trained using a training set of videos of the transport device colliding with a picking station on the grid to classify videos subsequently captured by the camera as including or not including a collision. The image data may be acquired based on, for example, in response to, a protective stop (or “p-stop”) alert from a given picking station 50 on the grid. For example, a protective stop may be initiated by a robot controller upon detection of an obstruction in a given robotic manipulator 52. Exemplary causes of a protective stop include too much force being applied to a given robot manipulator 52, e.g., exceeding its operating specifications, and the robot manipulator or its attached end effector, peripheral, or workpiece unexpectedly coming into contact with (e.g., "hitting into") something. Thus, in response to a p-stop being initiated by the robot controller, video data corresponding to the time period leading up to the p-stop may be acquired to determine whether a transport device collided with the robot manipulator (thereby causing a p-stop warning to be issued).

[0132] In response to determining that the image includes a transport device colliding with a picking station, annotation data indicating at least one of the transport device or the picking station in the image is generated and, in an example, output. In some cases, a method for recognizing an identifier of the picking station is employed in response to a positive determination of a transport device colliding with a picking station. For example, identifier information related to an identifier on the picking station is determined and, optionally, output as part of the method. In some cases, an exclusion zone centered on a grid cell of the picking station is determined in response to determining that the transport device collided with the picking station. Additionally or alternatively, an exclusion zone centered on a grid cell in which the crashed bot is located is determined. In some cases, the exclusion zone may be determined as an area of ​​grid cells centered on the grid cell, e.g., a 3x3 cell area. Thus, the exclusion zone may include a buffer area around the affected cell in which the crashed bot is located. In some cases, the crashed transport device spans more than one grid cell, for example, if it is located between grid cells, has fallen, or is misaligned with grid 15. In such cases, a buffer area around the mapped grid cell can improve the effectiveness of the exclusion zone versus excluding only the mapped grid cell. The determined exclusion zone can be implemented, for example, by a master controller, to prohibit transport devices 30 from entering the exclusion zone on grid 15.

[0133] In examples employing storage to store data, the storage may be random access memory (RAM) such as DDR-SDRAM (double data rate synchronous dynamic random access memory). In other examples, the storage may include non-volatile memory such as read-only memory (ROM) or a solid-state drive (SSD) such as flash memory. The storage may in some cases include other storage media, for example, magnetic, optical, or tape media, compact discs (CDs), digital versatile discs (DVDs), or other data storage media. The storage may be removable or non-removable from the associated system.

[0134] In examples employing data processing, a processor may be employed as part of the system involved. The processor may be a general-purpose processor such as a central processing unit (CPU), microprocessor, graphics processing unit (GPU), digital signal processor (DSP), application specific integrated circuit (ASIC), field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any suitable combination thereof designed to perform the data processing functions described herein.

[0135] In examples involving neural networks, a dedicated processor may be employed as part of the system involved. The dedicated processor may be an NPU, a neural network accelerator (NNA), or other version of a hardware accelerator specialized for neural network functions. Additionally or alternatively, the neural network processing workload may be at least partially shared by one or more standard processors, e.g., a CPU or GPU.

[0136] The term “annotation data” is used throughout the description, and this term is assumed to coincide with predicted data or inferred data in alternative names. For example, an object detection model (e.g., comprising a neural network) may be trained using annotated images, e.g., images with annotations such as bounding boxes, that serve as ground truth for the model, e.g., predictions or inferences with a confidence of 100% or 1 when normalized. These annotations may be made by humans, e.g., for purposes of training the model. Thus, object detection of the present disclosure may be interpreted as outputting predicted or inferred data (e.g., instead of “annotation data”) to indicate a prediction or inference of a carrying device in an image. The predicted or inferred data may be expressed as annotations, e.g., bounding boxes and / or labels, applied to the image. The predicted or inferred data may include, for example, a confidence associated with the prediction or inference of a carrying device in the image. An annotation may be applied to the image, for example, based on the generated predicted or inferred data. For example, the image may be updated to include a bounding box surrounding the predicted carrying device with a label indicating the confidence level of the prediction, e.g., as a percentage value or a normalized value between 0 and 1.

[0137] It should also be understood that features described with respect to any one example may be used alone or in combination with other features described, and may also be used in combination with one or more features of any other of the examples, or in any combination of any other of the examples. Moreover, equivalents and modifications not described above may also be employed without departing from the scope of the appended claims.

Claims

1. 1. A computer-implemented method for detecting identification markers on a picking station in a grid comprising a plurality of grid cells, said grid forming part of a grid-based storage system with one or more picking stations mounted on said grid, each picking station comprising a robotic manipulator for transferring items between containers received in a respective grid cell adjacent said picking station; The method comprises: acquiring image data representing an image portion including a picking station of the one or more picking stations; processing the image data using a first neural network and a second neural network in series; wherein the first neural network is trained to detect instances of identification markers on a picking station in an image, and the second neural network is trained to recognize marker information associated with the identification markers in the image. outputting marker data representing the marker information determined by the second neural network; A computer-implemented method comprising:

2. acquiring annotated image data representing an image of at least a portion of the grid including the picking stations, wherein the annotated image data includes annotation data indicating the picking stations in the image; cropping the annotated image data to produce the image portion including the picking station; The method of claim 1 , comprising:

3. The process comprises: determining whether the image portion includes an identification marker on the picking station based on the processing with the first neural network; generating and outputting further annotation data corresponding to the location of the identifying marker in the image portion in response to determining that the image portion includes the identifying marker; The method of claim 1 or 2, comprising:

4. The method of claim 3 , wherein the further annotation data comprises image coordinates relative to the image.

5. The method of claim 4 , wherein the image coordinates correspond to at least two corners of a bounding box for the identification marker in the image portion.

6. Processing the image portion comprises: extracting a sub-portion of the image portion, the sub-portion corresponding to the identification marker on the picking station; rotating the sub-portion so that the longitudinal axis of the identification marker is substantially horizontal with respect to the sub-portion; processing the subportion using the second neural network; and The method of any one of claims 1 to 5, comprising:

7. 7. The method of claim 1, wherein the first neural network is configured to receive the image portion as first input data and to produce a feature vector as intermediate data.

8. The method of claim 7 , wherein the second neural network is configured to receive the intermediate data and to produce the marker data as output data.

9. The method of claim 1 , wherein the first neural network comprises a convolutional neural network.

10. The method of claim 1 , wherein the second neural network comprises a bidirectional recurrent neural network.

11. A detection system for detecting identification markers on a picking station in a grid comprising a plurality of grid cells, said grid forming part of a grid-based storage system with one or more picking stations mounted on said grid, each picking station comprising a robotic manipulator for transferring items between containers received in respective grid cells adjacent said picking station; the detection system comprising: an image sensor for capturing an image of at least a portion of the grid; an interface for acquiring image data; one or more processors configured to implement a first neural network trained to detect instances of identification markers on the picking station in the image and a second neural network trained to recognize marker information associated with the identification markers in the scene; Equipped with wherein the detection system comprises: acquiring image data representing an image portion including a picking station of the one or more picking stations; processing the image data using the first neural network and the second neural network in series to generate marker data representative of marker information present on an identification marker on the picking station; outputting the marker data; A detection system configured to:

12. the one or more processors: acquiring annotated image data representing an image of at least a portion of the grid including the picking stations, wherein the annotated image data includes annotation data indicating the picking stations in the image; cropping the annotated image data to produce the image portion including the picking station; The detection system of claim 11 configured to:

13. the one or more processors: determining whether the image portion includes an identification marker on the picking station based on the processing with the first neural network; generating and outputting further annotation data corresponding to the location of the identifying marker in the image portion in response to determining that the image portion includes the identifying marker; 13. The detection system of claim 11 or 12, configured to:

14. The detection system of claim 13 , wherein the further annotation data comprises image coordinates relative to the image.

15. The detection system of claim 14 , wherein the image coordinates correspond to four corners of a bounding box for the identification marker in the image portion.

16. the one or more processors: extracting a sub-portion of the image portion, the sub-portion corresponding to the identification marker on the picking station; rotating the sub-portion so that the longitudinal axis of the identification marker is substantially horizontal with respect to the sub-portion; processing the subportion using the second neural network; and 16. The detection system of claim 11, configured to:

17. 17. The detection system of claim 11, wherein the first neural network is configured to receive the image portion as first input data and to produce a feature vector as intermediate data.

18. 18. The detection system of claim 17, wherein the second neural network is configured to receive the intermediate data and produce the marker data as output data.

19. 19. The detection system of claim 11, wherein the first neural network comprises a convolutional neural network.

20. 20. The detection system of claim 11, wherein the second neural network comprises a bidirectional recurrent neural network.

Citation Information

Patent Citations

  • Picking system and method

    JP2018533536A

  • Object recognition system, location information acquisition method, and program

    JP2022079775A

  • Storage containers, bins and devices

    US20180319590A1