Detecting debris on the grid of the storage system
The system uses ultra-wide-angle cameras and neural networks to detect and map debris on grid frameworks, addressing inefficiencies and transport device obstructions, thereby maintaining system efficiency and throughput.
Patent Information
- Application Number
- JP2025531858
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-30
- Filing Date
- 2023-11-30
- Publication Date
- 2025-12-05
AI Technical Summary
Existing storage and fulfillment systems face inefficiencies in managing large numbers of product lines and small quantities of items, particularly perishables, and debris on the grid framework can obstruct transport devices, causing derailment or control issues.
A system using ultra-wide-angle cameras and neural networks for debris detection on the grid framework, calibrating images to accurately map debris locations, allowing for manual or robotic clearing and preventing transport device interference.
Enhances system efficiency by detecting and preventing debris-related transport device issues, ensuring smooth operation and maintaining high throughput.
Smart Images

Figure 2025539479000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to the field of storage or fulfillment systems in which stacks of bins or containers are arranged within a grid framework structure, and more particularly to detecting debris on the grid framework. [Background technology]
[0002] Online retail businesses that sell multiple product lines, such as online grocery stores and supermarkets, need systems that can store tens or even hundreds of thousands of different product lines. In such cases, the use of single-product stacks may be impractical because a huge amount of floor space would be required to accommodate all of the required stacks. Additionally, it may be desirable to store small quantities of some items, such as perishable or infrequently ordered goods, making single-product stacks an inefficient solution.
[0003] PCT Publication No. WO2015 / 185628A (Ocado) describes a further known storage and fulfillment system in which stacks of containers are arranged within a grid framework structure. The containers are accessed by one or more load handling devices, sometimes known as robots or "bots," operable on tracks on top of the grid framework structure. A system of this type is shown schematically in Figures 1 to 3 of the accompanying drawings.
[0004] As shown in FIGS. 1 and 2, stackable containers 10, also known as "bins," are stacked on top of each other to form a stack 12. The stacks 12 are arranged in a grid framework structure 14, for example, in a warehousing or manufacturing environment. The grid framework structure 14 consists of a plurality of storage rows or grid rows. Each grid in the grid framework structure has at least one grid row for storing a stack of containers. FIG. 1 is a schematic perspective view of the grid framework structure 14, and FIG. 2 is a schematic top-down view showing a stack 12 of bins 10 arranged within the framework structure 14. Each bin 10 typically holds multiple product items (not shown). The product items in the bins 10 can be of the same or different product types, depending on the application.
[0005] The grid framework structure 14 includes a plurality of upright members 16 supporting horizontal members 18, 20. A first set of parallel horizontal grid members 18 are arranged in a grid pattern, perpendicular to a second set of parallel horizontal members 20, to form a horizontal grid structure 15 supported by the upright members 16. The members 16, 18, 20 are typically fabricated from metal. The bins 10 are stacked between the members 16, 18, 20 of the grid framework structure 14 such that the grid framework structure 14 guards against horizontal movement of the stack 12 of bins 10 and guides vertical movement of the bins 10.
[0006] The top level of the grid framework structure 14 comprises a grid or grid structure 15 including rails 22 arranged in a grid pattern across the top of the stacks 12. Referring to FIG. 3 , the rails or tracks 22 guide a plurality of load handling devices 30. A first set 22a of parallel tracks or rails 22 guides movement of the robotic load handling devices 30 in a first direction (e.g., the X direction) across the top of the grid framework structure 14. A second set 22b of parallel tracks or rails 22, positioned orthogonal to the first set 22a, guides movement of the load handling devices 30 in a second direction (e.g., the Y direction) orthogonal to the first direction. In this manner, the tracks or rails 22 allow the robotic load handling devices 30 to move laterally in two dimensions in the horizontal XY plane. The load handling devices 30 can be moved to a position above any of the stacks 12.
[0007] A known form of load handling device 30, shown in Figures 4 and 5, is described in PCT Patent Publication No. WO2015 / 019055 (Ocado), which is incorporated herein by reference, with each load handling device 30 covering a single grid space 17 of the grid framework structure 14. This configuration allows for a higher density of load handlers and therefore a higher throughput for a storage system of a given size.
[0008] The exemplary load handling device 30 includes vehicles 32 positioned to roll on the rails 22 of the grid framework structure 14. A first set of wheels 34, consisting of a pair of wheels 34 at the front of the vehicle 32 and a pair of wheels 34 at the rear of the vehicle 32, are positioned to engage two adjacent rails of the first set 22a of rails 22. Similarly, a second set of wheels 36, consisting of a pair of wheels 36 on each side of the vehicle 32, are positioned to engage two adjacent rails of the second set 22b of rails 22. At any time during the movement of the load handling device 30, each set of wheels 34, 36 can be raised and lowered so that either the first set of wheels 34 or the second set of wheels 36 is engaged with the respective set of rails 22a, 22b. For example, when a first set of wheels 34 is engaged with a first set of rails 22a and a second set of wheels 36 is lifted off the rails 22, the first set of wheels 34 can be driven by a drive mechanism (not shown) housed in the vehicle 32 to move the load handling device 30 in the X direction. To achieve movement in the Y direction, the first set of wheels 34 is lifted off the rails 22 and the second set of wheels 36 is lowered to engage with a second set 22b of rails 22. The drive mechanism can then be used to drive the second set of wheels 36 to move the load handling device 30 in the Y direction.
[0009] The load handling device 30 is equipped with a lifting mechanism, e.g., a crane mechanism, for lifting a storage container from above. The lifting mechanism includes a winch tether or cable 38 wound on a spool or reel (not shown) and a gripper device 39. The lifting mechanism shown in FIGS. 4 and 5 includes a set of four vertically extending lifting tethers 38. The tethers 38 are connected at or near each of the four corners of the gripper device 39, e.g., a lifting frame, for releasable connection to the storage container 10. For example, each tether 38 is positioned at or near each of the four corners of the lifting frame 39. The gripper device 39 is configured to releasably grip the top of the storage container 10 to lift it from a stack of containers in a storage system 1 of the type shown in FIGS. 1 and 2. For example, the lifting frame 39 may include pins (not shown) that mate with corresponding holes (not shown) in a rim forming the top surface of the bin 10 and sliding clips (not shown) that are engageable with the rim to grip the bin 10. The clips are housed within a lifting frame 39 and are driven into engagement with the bins 10 by a suitable drive mechanism powered and controlled by signals carried through the cable 38 itself or a separate control cable (not shown).
[0010] To remove a bin 10 from the top of the stack 12, the load handling device 30 is first moved in the X and Y directions to position the gripper device 39 above the stack 12. The gripper device 39 is then lowered vertically in the Z direction to engage the bin 10 at the top of the stack 12, as shown in FIGS. 4 and 6B. The gripper device 39 grasps the bin 10 and is then pulled upward by the cable 38 with the bin 10 attached. At the top of its vertical travel, the bin 10 is held on the rail 22 and housed within the vehicle body 32. In this manner, the load handling device 30, carrying the bin 10 therewith, can be moved to different positions in the XY plane to transport the bin 10 to another location. Upon arriving at the target location (e.g., another stack 12, an access point in a storage system, or a conveyor belt), the bin or container 10 can be lowered from the container receiving portion and released from the gripper device 39. The cable 38 is long enough to allow the load handling device 30 to pick and place bins from any level of the stack 12, including, for example, floor level.
[0011] As shown in Figure 3, multiple load handling devices 30 are provided so that each load handling device 30 can operate simultaneously to increase system throughput. The system shown in Figure 3 may include specific locations, known as ports, where bins 10 can be transferred into or out of the system. An additional conveyor system (not shown) is associated with each port so that bins 10 transported to a port by a load handling device 30 can be transferred by that conveyor system to another location, such as a picking station (not shown). Similarly, bins 10 can be moved by the conveyor system from an external location to the port, for example, to a bin filling station (not shown), and transported by the load handling device 30 to the stacks 12 to replenish stock in the system.
[0012] Each load handling device 30 is capable of lifting and moving one bin 10 at a time. The load handling device 30 has a container receiving cavity or recess 40 in its lower portion. The recess 40 is sized to accommodate the container 10 when it is lifted by the lifting mechanisms 38, 39, as shown in Figures 6A and 6B. When in the recess, the container 10 is lifted off the lower rail 22 to allow the vehicle 32 to move laterally to different grid locations.
[0013] When it is necessary to remove a bin 10b that is not at the top of a stack 12 (a "target bin"), the bins 10a above (a "non-target bin") must first be moved to allow access to the target bin 10b. This is accomplished by an operation hereinafter referred to as "digging." Referring to FIG. 3, during a digging operation, one of the load handling devices 30 sequentially lifts each non-target bin 10a from the stack 12 containing the target bin 10b and places it in a vacant position in another stack 12. The target bin 10b can then be accessed by the load handling device 30 and moved to a port for further transport.
[0014] Each load handling device 30 is remotely operable under the control of a central computer, e.g., a master controller. Also, each individual bin 10 in the system is tracked so that the appropriate bin 10 can be removed, transported, and replaced as needed. For example, during a dig operation, each non-target bin location is logged so that the non-target bin 10a can be tracked.
[0015] Wireless communications and networks may be used to provide a communications infrastructure from a master controller, e.g., via one or more base stations, to one or more load handling devices 30 operable on the grid structure 15. In response to receiving instructions from the master controller, a controller in the load handling device 30 is configured to control various drive mechanisms to control movement of the load handling device. For example, the load handling device 30 may be instructed to retrieve a container from a target storage row at a specific location on the grid structure 15. The instructions may include various movements in the XY plane of the grid structure 15. As previously described, upon reaching the target storage row, the lifting mechanisms 38, 39 may be operated to grasp and lift the storage container 10. Once the container 10 is received in the container receiving space 40 of the load handling device 30, the container 10 is then transported to another location on the grid structure 15, e.g., a “drop-off port.” At the drop-off port, the container 10 is lowered to a suitable pick station to enable retrieval of any items in the storage container. Movement of the load handling device 30 on the grid structure 15 may also involve the load handling device 30 being commanded to move to a charging station, typically located on the periphery of the grid structure 15 .
[0016] To move the load handling devices 30 on the grid structure 15, each load handling device 30 is equipped with a motor for driving the wheels 34, 36. The wheels 34, 36 may be driven via one or more belts connected to the wheels or may be individually driven by motors integrated into the wheels. In the case of single-cell load handling devices (where the footprint of the load handling device 30 occupies a single grid cell 17), the motors for driving the wheels may be integrated into the wheels due to the limited availability of space within the vehicle body. For example, the wheels of a single-cell load handling device 30 are driven by respective hub motors. Each hub motor includes an outer rotor with multiple permanent magnets arranged to rotate about a wheel hub with coils forming an inner stator.
[0017] 1-5 has many advantages and is suitable for a wide range of storage and retrieval operations. In particular, it allows for very high density storage of products, providing a very economical way of storing a wide range of different items in the bins 10, while also allowing reasonably economical access to all of the bins 10 when needed for picking.
[0018] Referring to FIG. 6 , the system may further include a robotic picking station 50 mounted on the storage and retrieval structure 1, for example, beside the load handling device 30 (not shown). The robotic picking station 50 includes a number of designated grid cells 60, 62, as well as a robotic manipulator 52 including a robotic arm 54 and an end effector 56 for releasably engaging a product to be manipulated. The end effector 56 may be a suction device 64 connected to a negative pressure source by a vacuum line 66. The robotic manipulator 52 is mounted on a pedestal 58 above a single grid cell 60 and, depending on its location on the structure 1, may be surrounded by up to eight other grid cells 62 as shown in FIG. 6 . Generally, the robotic manipulator 52 is configured to pick up an item or product from any one of the containers in one of the designated grid cells 62 and place it into a container in another of the designated grid cells 62. The load handling device collects containers from and delivers containers to the designated grid cells 62 as needed. In this manner, the robotic picking stations 50 and the load handling devices 30 work in conjunction to fulfill customer orders or redeploy products throughout the storage and retrieval system 1. Summary of the Invention
[0019] A method for detecting debris in a workspace having a grid formed by a first set of tracks extending in a first direction and a second set of tracks extending in a second direction orthogonal to the first direction is provided, the method comprising: obtaining image data representing an image of at least a portion of the workspace; processing the image data with an object detection model trained to detect instances of debris on the grid; determining based on the processing whether the image includes debris on the grid; and outputting annotation data indicative of the debris in the image in response to determining that the image includes debris on the grid.
[0020] Also provided is a data processing apparatus comprising a processor configured to perform the method. Also provided is a computer program comprising instructions that, when executed by a computer, cause the computer to perform the method. Similarly, provided is a computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to perform the method.
[0021] Further provided is a system for detecting debris in a workspace comprising a grid formed by a first set of tracks extending in a first direction and a second set of tracks extending in a second direction perpendicular to the first direction, the detection system comprising: an image sensor for capturing an image of at least a portion of the workspace; and an object detection model trained to detect instances of debris on the grid, wherein the detection system is configured to: obtain image data representing the image; process the image data with the object detection model; determine based on the processing whether the image includes debris on the grid; and, in response to determining that the image includes debris on the grid, output annotation data indicative of the debris in the image.
[0022] Generally speaking, this specification introduces systems and methods for detecting debris on a grid structure of a grid-based storage system using a trained object detection model. This allows, for example, the grid structure to be monitored for debris and the location of detected debris to be determined. Thus, the systems and methods allow the location of debris on the grid structure to be determined so that action can be taken, for example, to limit movement of transport devices on the grid to avoid the detected debris and / or to clear the debris so that the storage system can return to full functionality.
[0023] Embodiments will now be described, by way of example only, with reference to the accompanying drawings, in which like reference numerals designate the same or corresponding parts and in which: [Brief explanation of the drawings]
[0024] [Figure 1] Schematic diagram of the automatic storage and retrieval structure. [Figure 2] 2 is a schematic diagram of a plan view of a section of a track structure forming part of the storage structure of FIG. 1; [Figure 3] 2 is a schematic diagram of a plurality of load handling devices moving over the storage structure of FIG. 1; [Figure 4] Schematic of a load handling device interacting with a container. [Figure 5] Schematic of a load handling device interacting with a container. [Figure 6] 1 is a schematic diagram of a known robotic picking station. [Figure 7A] 1 is a schematic diagram of a storage system with a camera located above a grid framework structure as part of a detection system, according to an embodiment. [Figure 7B] 1 is a schematic diagram of a storage system with a camera located above a grid framework structure as part of a detection system, according to an embodiment. [Figure 8] 1 is a diagram of a schematic representation of an image captured by a camera positioned above a grid framework structure, according to certain embodiments. [Figure 9] Schematic diagram of a neural network. [Figure 10A] Schematic of the generated model of the track in the grid framework structure. [Figure 10B] Schematic of the generated model of the track in the grid framework structure. [Figure 11] Schematic diagram showing flattening of a captured image of a grid framework structure. [Figure 12] 1 is a schematic diagram illustrating processing of a captured image of a grid framework structure according to an embodiment. [Figure 13]1 is a flowchart illustrating a method for detecting debris on a grid forming part of a grid-based storage system, according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0025] Monitoring the grid structure 15 of a grid-based storage system to detect debris (e.g., one or more discarded or scattered pieces of material) can reduce the likelihood that a transport device will encounter the debris. For example, debris on the grid can obstruct the transport device, potentially causing it to derail or otherwise lose control as it moves on the track. Liquid debris, for example, can cause the transport device to slip on the track. Alternatively, the debris can provide resistance to the movement of the transport device on the track, potentially creating problems with controlling its movement (e.g., by a master controller). Thus, detection and location of debris on the grid 15 can be used, for example, to aid in clearing the debris manually or by dedicated robotic equipment, or at least to prevent a transport device from contacting the debris while it is present on the grid.
[0026] FIG. 7A shows the previously described grid structure (or simply “grid”) 15 of the storage system. The grid is formed by a first set 22a of parallel tracks extending in the X direction and a second set 22b of parallel tracks extending in the Y direction, orthogonal to the first set in a substantially horizontal plane. The grid 15 has a plurality of grid spaces 17. One or more load handling devices, or “transport devices” 30, are arranged to selectively move in at least one of the X or Y directions on the tracks 22 and to handle containers 10 stacked below the tracks 22 within the footprint of a single grid space 17. In the example, the one or more transport devices 30 each have a footprint that occupies only a single grid space, such that a given transport device occupying one grid space does not interfere with another transport device occupying or crossing an adjacent grid space.
[0027] A camera 71 is disposed above the grid 15. In an example, the camera 71 is an ultra-wide-angle camera, i.e., it has an ultra-wide-angle lens (also referred to as a "super wide-angle" lens or a "fisheye" lens). The camera 71 includes an image sensor for receiving incident light focused through a lens, e.g., a fisheye lens. The camera 71 has a field of view 72 that includes at least a section of the grid 15. Multiple cameras may be used to observe the entire grid 15, e.g., each camera 71 having a respective field of view 72 that covers a section of the grid 15. The ultra-wide-angle lens may be selected because of its relatively large field of view 72, e.g., up to a 180-degree solid angle, compared to other lens types, which means that fewer cameras are needed to cover the grid 15. Also, space may be limited between the top of the grid 15 and surrounding structures, e.g., a warehouse roof, thus constraining the height of the camera 71 above the grid 15. An ultra-wide lens camera can provide a relatively large field of view at a relatively low height above the grid 15 compared to other camera types.
[0028] One or more cameras 71 may be used to monitor the workspace of transport device 30, the workspace including grid structure 15. For example, image feeds from one or more cameras 71 may be acquired for processing image data to detect debris in the workspace that may interfere with transport device 30 as it moves on grid 15. The image feeds may be simultaneously displayed on one or more remote computer monitors, for example, for manual monitoring of grid 15 by an operator.
[0029] Calibration Process A monitoring or surveillance system for grid 15 may incorporate calibration of one or more cameras 71 located above grid 15, particularly in embodiments comprising wide-angle or ultra-wide-angle cameras. Accurate calibration of the (ultra-)wide-angle camera may allow interaction with images captured by the (ultra-)wide-angle camera, distorted by the (ultra-)wide-angle lens, to be correctly mapped to the workspace. Thus, selected areas of pixels in the distorted image may be mapped to corresponding areas in grid space, for example.
[0030] An exemplary calibration process for an ultra-wide-angle camera includes acquiring a section of a grid, i.e., an image of the grid section, captured by the camera. Acquiring the image includes obtaining, e.g., receiving, image data representing the image, e.g., in a processor. For example, the image data may be received via an interface, e.g., a camera serial interface (CSI). An image signal processor (ISP) may perform initial processing of the image data, e.g., saturation correction, renormalization, white balancing, and / or demosaicing, to prepare the image data for display.
[0031] Initial values for several parameters corresponding to the ultra-wide-angle camera are also obtained, including the focal length of the ultra-wide-angle camera, a translation vector representing the position of the ultra-wide-angle camera above the grid section, and a rotation vector representing the tilt and rotation of the ultra-wide-angle camera. These parameters can be used in a mapping algorithm to map pixels in the image distorted by the camera's ultra-wide-angle lens onto a plane oriented relative to the storage system's Cartesian grid 15. The mapping algorithm is described in more detail below.
[0032] The calibration process involves processing images using a neural network trained to detect / predict tracks in images of grid sections captured by an ultra-wide-angle camera. Neural Networks 9 illustrates an example of a neural network architecture. The exemplary neural network 90 is a convolutional neural network (CNN). One example of a CNN is the U-Net architecture developed by the Department of Computer Science at the University of Freiburg, although other CNNs, such as the VGG-16 CNN, can be used. The input 91 to the CNN 90 comprises image data in this example. The input image data 91 is a given number of pixels wide and a given number of pixels high and includes one or more color channels (e.g., red, green, and blue channels).
[0033] The convolutional layers 92, 94 of the CNN 90 may generally extract specific features from the input data 91 and operate on small portions of the image to create feature maps. The fully connected layer 96 uses the feature maps to determine an output 97, e.g., classification data specifying the classes of objects predicted to be present in the input image 91.
[0034] In the example of FIG. 9 , the output of the first convolutional layer 92 undergoes pooling in a pooling layer 93 before being input to a second convolutional layer 94. Pooling, for example, allows values for a region of an image or feature map to be aggregated or combined, e.g., by taking the highest value within the region. For example, in 2×2 max pooling, rather than transferring the entire output, the highest value of the output of the first convolutional layer 92 within a 2×2 pixel patch of the feature map output from the first convolutional layer 92 is used as input to the second convolutional layer 94. Pooling can therefore reduce the amount of computation for subsequent layers of the neural network 90. The effect of pooling is shown schematically in FIG. 9 as a reduction in the size of the frames in the relevant layers. Further pooling is performed in a second pooling layer 95 between the second convolutional layer 94 and the fully connected layer 96. It should be appreciated that the schematic representation of neural network 90 in FIG. 9 is greatly simplified for ease of explanation, and that typical neural networks can be significantly more complex.
[0035] Generally, a neural network, such as the neural network 90 of FIG. 9, may undergo what is called a "training phase," during which the neural network is trained for a specific purpose. A neural network generally includes layers of interconnected artificial neurons that form a directed, weighted graph, where the graph's vertices (corresponding to neurons) or edges (corresponding to connections) are each associated with a weight. The weights may be adjusted throughout training, changing the output of individual neurons and, therefore, the neural network as a whole. In a CNN, a fully connected layer 96 generally connects every neuron in one layer to every neuron in another layer and may thus be used to identify global properties of an image, such as whether the image contains a particular class of object or a particular instance of a particular class.
[0036] In the present context, neural network 90 is trained to perform object identification by processing image data, e.g., to determine whether an object of a predetermined class of objects is present in the image (although in other examples, neural network 90 may instead be trained to identify other image characteristics of the image). For example, training neural network 90 in this manner generates weight data representing weights to be applied to the image data (e.g., different weights associated with different layers of a multi-layer neural network architecture). Each of these weights is multiplied by the corresponding pixel value of the image patch, e.g., to convolve the weight kernel with the image patch.
[0037] Specific to the context of ultra-wide-angle camera calibration, neural network 90 is trained using a training set of input images of grid sections captured by an ultra-wide-angle camera to detect tracks 22 of grid 15 in a given image of the grid section. In an example, the training set includes mask images that show only extracted track features corresponding to the input images. For example, the mask images are manually created. Thus, the mask images can serve as a desired result for neural network 90 to be trained using the training set of images. Once trained, neural network 90 can be used to detect tracks 22 in images of at least a portion of grid structure 15 captured by an ultra-wide-angle camera.
[0038] The calibration process 130 includes processing 133 the image of the grid section captured by the ultra-wide-angle camera 71 using the trained neural network 90 to detect tracks 22 in the image. At least one processor (e.g., a neural network accelerator) may be used to perform the processing 133. The image processing 133 generates models of the tracks, in particular a first set and a second set of parallel tracks, captured in the image of the grid section. For example, the models comprise representations of predictions of the tracks in the distorted image of the grid section determined by the neural network 90. The track models correspond, in examples, to masks or probability maps.
[0039] Selected pixels in the determined track model are then mapped 134 to corresponding points on the grid 15 using a mapping, e.g., a mapping algorithm, that incorporates multiple parameters corresponding to the ultra-wide-angle camera. The obtained initial values are used as input to the mapping algorithm.
[0040] An error function (or "loss function") is determined 135 based on the discrepancy between the mapped grid coordinates and the "true," e.g., known, grid coordinates of the points corresponding to the selected pixels. For example, a selected pixel at the center of X-direction track 22a should correspond to a grid coordinate with a half-integer value in the Y direction, e.g., (x, y.5), where x is an unknown number and y is an unknown integer. Similarly, a selected pixel at the center of Y-direction track 22b should correspond to a grid coordinate with a half-integer value in the X direction, e.g., (x'.5, y'), where x' is an unknown integer and y' is an unknown number. In an example, the width and length of a grid cell (or their ratio) are used in the loss function, for example, to calculate cell x,y coordinates for keypoints and determine whether they are on the track (e.g., coordinate values of n.5, where n is an integer).
[0041] The initial values of the multiple parameters corresponding to the ultra-wide-angle camera are then updated to updated values based on the determined error function 136. For example, a Broyden-Fletcher-Goldfarb-Schanno (BFGS) algorithm is applied using the error function and the initial parameter values as input. In an example, the updated values of the multiple parameters are determined iteratively, and the error function is recalculated with each update. The iterations may continue, for example, until the error function is reduced by less than a predetermined threshold between successive iterations or compared to the initial error function, or until the absolute value of the error function falls below a predetermined threshold. Other iterative algorithms, such as sequential quadratic programming (SQP) or sequential least-squares quadratic programming (SLSQP), may be used with the initial values to generate a sequence of improving approximate solutions for the multiple parameters, with a given approximation in the sequence being derived from the previous one. In some cases, the iterative algorithm is used to optimize the values of the multiple parameters. For example, the updated values are optimized values of the multiple parameters.
[0042] Updating 136 the initial values of the plurality of parameters corresponding to the ultra-wide-angle camera involves applying one or more respective boundary values for the plurality of parameters. For example, the boundary values for the rotation angle associated with the rotation vector are substantially 0 degrees and substantially +5 degrees. Additionally or alternatively, the boundary values for the planar component of the translation vector are ±0.6 of the length of the grid cell. Additionally or alternatively, the boundary values for the height component of the translation vector are 1800 mm and 2100 mm, or 1950 mm and 2550 mm, or 2000 mm and 2550 mm above the grid. For example, the lower boundary for the camera height is within the range of 1800 to 2000 mm. For example, the upper boundary for the camera height is within the range of 2100 to 2600 mm. Additionally or alternatively, the boundary values for the focal length of the camera are 0.23 and 0.26 cm. Applying one or more respective boundary values for multiple parameters can mean that the updating, e.g., optimization, process is performed in a feasible region or solution space, i.e., the set of all possible values that satisfy one or more boundary conditions.
[0043] The updated values of the plurality of parameters are electronically stored 137 for future mapping of pixels in a grid section image captured by the ultra-wide-angle camera 71 to corresponding points on the grid 15 via a mapping algorithm. For example, the stored values of the plurality of parameters are retrieved from data storage and used in the mapping algorithm to calculate grid coordinates corresponding to a given pixel in a given image of the grid section captured by the ultra-wide-angle camera 71. In an example, the updated values are stored in a storage location associated with the ultra-wide-angle camera 71, for example, in a database. For example, a lookup function or table may be used in conjunction with the database to find the stored parameter values associated with any given ultra-wide-angle camera employed in the storage system 1 on the grid 15.
[0044] Following calibration of a given camera 71 disposed over the grid 15, an image (e.g., a “snapshot”) of a grid section captured by the camera 71 may be flattened, i.e., undistorted, for interaction by an operator. For example, using the described image-to-grid mapping function, a distorted image 81 of the grid section may be converted into a flattened image 111 of the grid section, as shown in the example of FIG. 11 . Flattening involves selecting an area of grid cells to flatten in the distorted image 81 and inputting the grid coordinates corresponding to those cells into the mapping function, which determines which respective pixel values from the distorted image 81 should be copied into the flattened image 111 for each grid coordinate. For example, a target resolution, in pixels per grid cell, may be set for the flattened image 111, the target resolution having a ratio corresponding to the ratio of the grid cell dimensions. Once all pixel values needed in the flattened image (according to the target resolution and the selected number of grid cells) have been determined, the flattened image 111 can be generated.
[0045] Snapshots may be captured by the camera 71 at predetermined intervals, e.g., every 10 seconds, and converted to corresponding flattened images 111. The most recent flattened images 111 may be stored in storage for viewing on a display, e.g., by an operator wishing to view the grid section covered by the camera 71. The operator may instead choose to re-take a snapshot of the grid section and flatten it. Thus, the operator may select regions, e.g., pixels, in the flattened image 111 and convert those selected regions to grid coordinates based on the image-to-grid mapping functionality described herein. In some cases, the flattened image 111 includes grid coordinate annotations for the grid space viewable in the flattened image 111. The flattened images 111 corresponding to each camera 71 may be more user-friendly for monitoring the grid 15 compared to the distorted images 81, 82. Grid to image mapping A computational algorithm maps real-world points on grid 15 to pixels in an image captured by the camera. The grid points are first projected onto a plane corresponding to ultra-wide-angle camera 71. For example, at least one of a rotation using a rotation matrix and a planar translation in the X and Y directions is applied to points having x, y, and z coordinates in grid framework structure 14. The focal length f of the ultra-wide-angle camera may be used to project points with three-dimensional coordinates relative to grid 15 onto a two-dimensional plane relative to ultra-wide-angle camera 71. For example, the coordinates of a mapped point q in the plane of ultra-wide-angle camera 71 are given by q=f·p [x,y] ÷p z It is calculated as, where p [x,y] and p z are the planar xy coordinates and the third z coordinate of point p relative to grid 15, respectively.
[0046] A point q projected onto the ultra-wide-angle camera plane may be aligned with a Cartesian coordinate system in that plane to determine a first Cartesian coordinate of the point. For example, aligning a point with a Cartesian coordinate system involves rotating the point or its position vector in the plane (e.g., a vector from the origin to the point). Thus, the rotation is, for example, to align with a typical grid orientation in an image captured by the camera, but may not be necessary if the X and Y directions of the grid are already aligned with the captured image. In the example, the rotation is substantially 90 degrees. As shown in FIGS. 8A and 8B, the X and Y directions of the grid are offset by 90 degrees with respect to the horizontal and vertical axes of the image; therefore, the rotation "corrects" for this offset so that the X and Y directions of the grid are aligned with the horizontal and vertical axes of the captured image.
[0047] A grid-to-image mapping algorithm continues by converting the first Cartesian coordinate to a first polar coordinate using standard trigonometry. A distortion model is then applied to the first polar coordinate of the point to generate a second, e.g., "distorted," polar coordinate. In an example, the distortion model comprises a tangent model of distortion given by r' = f arctan(r / f), where r and r' are the undistorted and distorted radial coordinates of the point, respectively, and f is the focal length of the ultra-wide-angle camera.
[0048] The second polar coordinates are then converted back to (second) Cartesian coordinates using the same standard trigonometric methods inversely. Image coordinates of pixels in the image are then determined based on the second Cartesian coordinates. In examples, this determination includes at least one of de-centering or rescaling the second Cartesian coordinates. Additionally or alternatively, the ordinate (y-coordinate) of the second Cartesian coordinates is flipped, e.g., mirrored on the x-axis. Image to grid mapping Mapping pixels in the image captured by camera 71 to real-world points on grid 15 is done by different computational algorithms. For example, the image-to-grid mapping algorithm is the inverse of the grid-to-image mapping algorithm described above, with each mathematical operation reversed.
[0049] For a given pixel in the image, a (second) Cartesian coordinate of the mapped point is determined based on the image coordinate of the pixel in the image. For example, this determination may involve initializing the pixel in the image, including, for example, at least one of centering or normalizing the image coordinate. As mentioned above, the ordinate coordinate is inverted in some instances. The second Cartesian coordinate is converted to a second polar coordinate using standard trigonometry as described above. The use of the label "second" is used for consistency with the conversion performed in the grid-to-image mapping algorithm described, but is arbitrary.
[0050] An inverse distortion model is applied to the second polar coordinates to generate first, e.g., "undistorted," polar coordinates. In an example, the inverse distortion model is based on a tangent model of distortion given by r=f·tan(r' / f), where again r' is the distorted radial coordinate of the point, r is the undistorted radial coordinate of the point, and f is the focal length of the ultra-wide-angle camera. Thus, in an example, the inverse distortion model used in the image-to-grid mapping is the inverse, or "anti-function," of the distortion model used in the grid-to-image mapping.
[0051] The image-to-grid mapping algorithm continues by converting the first polar coordinate to a first Cartesian coordinate. The first Cartesian coordinate may be disaligned or misaligned with a Cartesian coordinate system in a plane corresponding to the ultra-wide-angle camera. For example, disaligning a point with a Cartesian coordinate system involves applying a rotation transformation to the point or its position vector in the plane (e.g., a vector from the origin to the point). The rotation is substantially 90 degrees in the example. This rotation may therefore "undo" any "correction" to the offset between the X and Y directions of the grid and the horizontal and vertical axes of the captured image, as previously described in the grid-to-image mapping.
[0052] Finally, the point is projected from the (second) plane corresponding to the camera 71 onto the (first) plane corresponding to the grid 15 to determine the grid coordinates of the point relative to the grid.
[0053] In the example, the projection of the point onto the plane corresponding to grid 15 is p=B -1 This involves calculating (f tq z), where B=q R 3,[1,2] -f·R [1,2],[1,2] In these equations, p comprises the point coordinate in the grid plane, q comprises the Cartesian coordinate in the camera plane, and f is the focal length of the ultra-wide-angle camera as described above. Furthermore, t is a planar translation vector, z is the distance (e.g., height) between the ultra-wide-angle camera and the grid, and R is a three-dimensional rotation matrix related to the rotation vector. The rotation vector comprises a direction representing the axis of rotation and a magnitude representing the angle of rotation. The rotation matrix R corresponding to the angle-axis rotation vector can be determined from the vector using, for example, Rodriguez's rotation formula.
[0054] Next, a mathematical derivation of the function for projecting an undistorted 2D point q from the camera plane is provided for completeness: starting from the projection from the grid onto the image from above, q = f p' [x,y] ÷p' zwhere p' is the rotated and translated grid point p, i.e., p'=R·p+(t x ,t y ,z) T The goal is to derive p from q. After rearranging and substituting for p', we get the following:
[0055]
number
[0056] Since the desired distance of a point p on the grid from the camera is given by the height parameter z, in translating the point p z = 0. Therefore, for all p z Terms can be eliminated to give:
[0057]
number
[0058] Matrix B=(q·R 3,[1,2] By defining -f·R), the equation becomes B·p [x,y] = f tz q, which can be further simplified to -1 Using
[0059] Returning to calibration process 130, in some cases, grid cell coordinate data encoded in grid cell markers positioned relative to grid 15 may be used to calibrate calculated grid coordinates corresponding to pixels in the captured image. For example, the grid cell markers may be signboards placed in predetermined grid cells 17, with corresponding cell coordinate data marked on each signboard. Process 130 may include, for example, processing the captured image to detect the grid cell markers in the image and then extracting the grid cell coordinate data encoded in the grid cell markers for use in calibrating the mapped grid coordinates. Each grid cell marker is in a respective grid cell, e.g., under and within the field of view 72 of a respective camera 71.
[0060] The image processing may involve using an object detection model, e.g., a neural network, trained to detect instances of grid cell markers in images of the grid section. A computer vision platform, e.g., Cloud Vision API (Application Programming Interface) by Google®, may be used to implement the object detection model. The object detection model may be trained using images of the grid section including the grid cell markers. In examples where the object detection model includes a neural network, e.g., a CNN, the description with reference to FIG. 9 applies accordingly.
[0061] Grid coordinates generated by mapping pixels in an image to points on a grid section represented in a captured image can be calibrated to the entire grid based on the extracted cell coordinate data. For example, a mapped grid point corresponding to a given pixel has coordinates in units of grid cells, e.g., (x, y), with x being the number of grid cells in the X direction and y being the number of grid cells in the Y direction. However, the grid cells captured by camera 71 are for a grid section, i.e., a section of grid 15, and therefore not necessarily the entire grid 15. Therefore, the mapped grid coordinates (x, y) for a grid section captured in an image can be calibrated to grid coordinates (x', y') for the entire grid based on the relative location of the grid section with respect to the entire grid. The location of the grid section with respect to the entire grid can be determined by extracting the grid cell coordinate data encoded in the grid cell markers captured in the image, as described.
[0062] 10A shows an exemplary model 101 of tracks generated by processing an image 81 of a grid section captured by an ultra-wide-angle camera 71 using a neural network 90 trained to detect tracks 22 in the image. Model 101 comprises a representation of a prediction of tracks 22 a, 22 b in a distorted image of the grid section determined by neural network 90. Mapping pixels from track model 101 to corresponding points on grid 15 may be performed to calibrate camera 71 as described. For example, calibration may involve updating, e.g., optimizing, multiple parameters associated with camera 71 used for mapping between pixels in captured images 81, 82 and points on grid 15.
[0063] In an example, the model 101 of the grid sections may be refined to represent only the centerlines of the first set 22a and the second set 22b of parallel tracks. Thus, the pixels to be mapped from the track model 101 to corresponding points on the grid 15 are, for example, pixels that lie on the centerlines of the first set 22a and the second set 22b of parallel tracks in the generated model 101. Refining involves, for example, filtering the model with horizontal and vertical line detection kernels. These kernels allow the centerlines of the tracks to be identified in the model 101, for example, in the same way that other kernels may be used to identify other features of an image, such as edges in edge detection. Each kernel is a given size, for example a 3x3 matrix, that may be convolved with the image data in the model 101 with a given stride. For example, the horizontal line detection kernel may be the matrix
[0064]
number
[0065] It can be expressed as:
[0066] Similarly, the vertical line detection kernel can be, for example, the matrix
[0067]
number
[0068] It can be expressed as:
[0069] In the example, filtering involves at least one of eroding and dilating pixel values of the model 101 using horizontal and vertical line detection kernels. For example, at least one of an erosion function and a dilation function is applied to the model 101 using the kernel. The erosion function effectively "erodes" foreground objects, in this case, the boundaries of tracks 22 a, 22 b in the generated model 101, by convolving the kernel with the model. During erosion, a pixel value (either "1" or "0") in the original model is updated to a value of "1" only if all pixels convolved under the kernel are equal to "1"; otherwise, it is eroded (updated to a value of "0"). Effectively, all pixels near the boundaries of tracks 22 a, 22 b in the model 101 are discarded, depending on the size of the kernel used in the erosion, such that the thickness of each of tracks 22 a, 22 b is reduced substantially to its centerline. A dilation function, which is the opposite of the erosion function, may be applied after erosion to effectively "inflate" or widen the centerlines remaining after erosion. This dilation may stabilize the centerlines of the tracks 22a, 22b in the improved model 101. During dilation, if at least one pixel convolved under the kernel is equal to "1," the pixel value is updated to a value of "1." The erosion and dilation functions are each applied to the original generated model 101, and the resulting horizontal and vertical centerline "skeleton," for example, are combined to produce the improved model.
[0070] In some cases, the generated model 101 may have missing sections of the tracks 22a, 22b, for example, obscuring one or more areas of the grid section viewable by the camera 71. Objects on the grid 15, such as a conveying device 30, a pillar, or other structure, may obscure portions of the tracks in the captured image. Thus, the generated model 101 may have the same missing areas of the tracks. Similarly, false positive predictions of the tracks may be present in the generated model 101.
[0071] To assist with these problems, the tracks 22a, 22b (e.g., their centerlines) present in the generated model can be fitted to their respective quadratic equations, for example, to create a secondary track for tracks 22a, 22b. FIG. 10B shows an example of a track among the first set 22a of tracks in model 101 being fitted to a first quadratic track 102 and a track among the second set 22b of tracks in model 101 being fitted to a second quadratic track 103. Then, a quadratic track centerline can be created based on the quadratic track by extrapolating pixel values along the quadratic track, for example, to fill gaps in model 101 or remove false positives. For example, if a sub-line generated from the predicted grid model 101 cannot be fitted to a given quadratic curve along with at least one other line, that sub-line is very likely not to be part of the grid and should be excluded.
[0072] The quadratic equation, y = ax 2 + bx + c used to fit the tracks in model 101 can also have specified boundary conditions, for example,
[0073]
Number
[0074] , -9.9×10 -4 < a < 9.9×10 -4 , -5 < b < 5, and 0 < c < 3200.
[0075] In an example, a predetermined number of pixels are extracted from the improved model 101 of the track, for example, to reduce the storage requirements for storing the model. For example, a random subset of pixels is extracted to give the final improved model 101 of the track.
[0076] Calibrating the ultra-wide-angle camera 71 using the systems and methods described herein allows, for example, images captured by the camera 71 with a wide field of view of the grid 15 to be used to detect and locate a carrier device thereon, despite the relatively high distortion present in the image compared to other camera types.
[0077] The automated calibration process outlined above can also reduce the time it takes to calibrate each camera 71 installed on the grid 15 of a storage system compared to manual methods of tuning parameters associated with each camera 71. For example, combining a neural network model, e.g., U-Net, with a customized optimizer to implement the described calibration pipeline can eliminate over 80% of the error compared to standard calibration methods. Furthermore, the calibration systems and methods described herein have been found to be versatile and consistent enough to calibrate cameras in multiple warehousing storage systems, e.g., having different dimensions, scales, and layouts.
[0078] Additionally, the output flattened calibrated image 111 of the grid allows for easier interaction with the image 111 by both humans and machines to monitor the grid 15 and the transport devices 30 moving thereon. Thus, instances of unresponsiveness of a given transport device on the grid may be more efficiently detected and / or acted upon to resolve fleet movements of the transport devices 30.
[0079] Detecting debris Provided herein are methods and systems for processing images, e.g., distorted images 81, captured by one or more cameras 71 to detect debris on the grid 15 of a grid-based storage system 1. For example, the location of the detected debris relative to the grid 15 may be output. Examples of debris include items stored in containers, spills of such items, and parts of a transport device. For example, items stored in containers may be dropped by a robotic manipulator at a picking station or fall from the container during transport of the container through the storage system. Items comprising liquid (e.g., cartons of milk) may cause spills onto the track with or without the container itself falling onto the track. For example, a leaking carton or bottle in a container may spill its contents onto the track. In another example, a transport device 30 may lose a part that fell onto the track following a collision with fixtures on the grid, such as another transport device or a picking station.
[0080] 13 shows a computer-implemented method 130 for detecting debris on grid 15. Method 130 involves obtaining 131 image data representing an image of at least a portion of grid 15 and processing 132 the image data with an object detection model trained to detect instances of debris on the grid. For example, the image is captured by a camera 71 with a field of view 72 covering at least a portion of grid 15, and the image data is transferred to a computer for implementing detection method 130. The image data is received, for example, at an interface of the computer, e.g., a CSI.
[0081] The object detection model may be a neural network, e.g., a convolutional neural network, trained to perform object detection of debris on the workspace grid 15. Accordingly, the neural network description with respect to FIG. 9 applies in these specific examples. In the present context, the object detection model, e.g., CNN 90, is trained to perform object identification by processing acquired image data to determine whether an object (i.e., debris) of a predetermined class of objects is present in the image. Training the neural network 90 involves, for example, providing the neural network 90 with training images of a workspace section in which a picking station resides. Weight data is generated for each (convolutional) layer 92, 94 of the multi-layer neural network architecture and stored for use in implementing the trained neural network. In the example, the object detection model comprises a “You Only Look Once” (YOLO) object detection model, e.g., YOLOv4 or Scaled-YOLOv4, having a CNN-based architecture. Other exemplary object detection models include neural-based approaches, such as RetinaNet or R-CNN (Regions with CNN features), and non-neural approaches, such as support vector machines (SVMs) for object classification based on determined features, e.g., Haar-like features or Histogram of Oriented Gradients (HOG) features.
[0082] Method 130 involves determining 133 whether the image includes debris on the grid based on process 132. For example, an object detection model is configured, e.g., trained or learned, to detect whether one or more pieces of debris are present in the captured image of grid 15. In an example, the object detection model makes determination 133 with a level of confidence, e.g., a probability score, corresponding to the likelihood that the image includes debris on the grid. Thus, a positive determination may correspond to a confidence level above a predetermined threshold, e.g., 90% or 95%. In response to determining 133 that the image includes debris, annotation data (e.g., predicted data or inferred data) indicative of the predicted debris in the image is output 134. An updated version of the image including the annotation data may be output, for example, as part of method 130.
[0083] In an example, annotation data output as part of the detection method comprises bounding box data. FIG. 12 shows an example of an updated version 83 of an image captured by camera 71 annotated with a bounding box 120 based on the bounding box data. The bounding box 120 corresponds to debris 122 detected by an object detection model. A given bounding box may comprise, for example, a rectangle surrounding a detected object and may specify one or more of an image location, an identified object class (e.g., debris), and a confidence score (e.g., how likely the object will be present within the box). The bounding box data defining a given bounding box may include coordinates of two corners of the box or center coordinates with width and height parameters for the box in image 83. In an example, the detection method 130 involves generating annotation data, for output, that can be represented, for example, as a bounding box 120.
[0084] In some cases, the object detection model is further trained to classify the debris into one of multiple classes of debris. For example, the debris may be classified as a storage item (e.g., a stockkeeping unit or "SKU"), a spill, or a bot part. Each class may have further subclasses; for example, a storage item may be classified as a class of storage item such as a carton, a bag, a can, etc. Thus, the detection method may involve, in response to determining that the image contains debris on the grid, processing the image data with one or more object classification models trained to classify debris to determine classification data representing the class of debris to which the detected debris belongs.
[0085] In an example, the detection system causes different responses to a determination that debris is present on the grid depending on the class of the detected debris. For example, in response to the classification data indicating that the detected debris belongs to a first class of debris (e.g., storage items), the detection system causes deployment of a servicing device arranged to move on a track and comprising a cleaning mechanism having means for removing debris present on the grid. Alternatively, in response to the classification data indicating that the detected debris belongs to a second class of debris (e.g., bot parts), the detection system causes, for example, the master controller to shut down any transport devices on the grid.
[0086] Determining exclusion zones Additionally or alternatively, in preparation for the deployment of service devices to remove detected debris on the grid (described further below), the detection system may cause exclusion zones to be established in the workspace. The exclusion zones may be implemented by the master controller of the transport device 30 and function, for example, to prohibit transport devices 30 operating in the workspace from entering the exclusion zone. For example, the exclusion zone may be determined around detected debris so that the debris can be later addressed, for example, cleaned or removed from the workspace. This allows the workspace to remain usable while reducing the risk of other transport devices coming into contact with the debris. In some cases, the determined exclusion zones may be proposed, for example to an operator, prior to implementation, which can help ensure that the determined exclusion zones will cover the actual location of the debris in the workspace.
[0087] In an example, the detection method involves determining a target image portion of the captured image based on annotation data corresponding to the detected debris on a grid. The target image portion is mapped to a target location in a workspace. Based on the mapping, an exclusion zone in the workspace is determined into which one or more transport devices will be prohibited from entering. The exclusion zone includes the target location mapped from the target image portion. Exclusion zone data representing the exclusion zone is output to a control system, for example, a master controller of the transport device, for implementing the exclusion zone in the workspace.
[0088] In an example, the target image portion includes at least a portion of a debris object in the workspace. For example, the target image portion is a subset of one or more pixels selected from an image of the workspace captured by an image sensor. The one or more pixels correspond to at least a portion of a debris object present in the captured image of the workspace. For example, the target image portion includes the entire debris object present in the image. In another example, the target image portion is only a single pixel corresponding to a portion of a debris object present in the image.
[0089] The target image portion is obtained from an object detection system configured to detect debris present on the grid from an image of the workspace. For example, the process involves the object detection system obtaining an image of the workspace captured by a camera and using an object classification model to determine that debris on the grid is present in the image data. The object classification model, e.g., an object classifier, generally comprises a neural network in the example described with reference to FIG. 9, which is to be taken to apply accordingly. For example, the object classifier is trained using a training set of images of debris in the workspace to classify images subsequently captured by the image sensor as either including or not including debris in the workspace.
[0090] In the case of a positive classification by the trained object classifier, the object detection system can then output the target image portion. For example, the object detection system may indicate the target image portion in the original image captured by the camera using annotation data, such as, for example, a bounding box. Alternatively, the object detection system outputs the target image portion as a cropped version of the original input image received from the image sensor, the cropped version including the identified debris in the workspace.
[0091] In an example, an object detection system includes a neural network trained to detect debris and its location in image data. For example, the object detection system determines regions of an input image in which debris is present. The regions may then be output, for example, as target image portions. In such a case, training the neural network involves using annotated images of the workspace that show debris in the workspace. Thus, the neural network is trained to both classify objects in the workspace as debris and to detect where the debris is in the image, i.e., to identify the location of the debris relative to the image of the workspace.
[0092] As described herein, the target image portion output by the object detection system may include at least a portion of a debris object in the workspace. For example, the target image portion is a subset of one or more pixels selected by the object detection system from an image captured by a camera based, for example, on a positive location identification of the debris object.
[0093] In an example, the determined exclusion zone comprises a discrete number of grid spaces, e.g., a plurality of grid cells adjacent to the detected debris on the grid. For example, if debris is detected at a junction of crossing tracks, the determined exclusion zone may comprise four grid cells adjacent to the junction. Alternatively, if debris is detected along a single portion of the track, the determined exclusion zone may comprise two grid cells on either side of the track portion. In either case, when implemented, transport devices operating on the grid may be prohibited from entering the exclusion zone, thus avoiding contact with debris on the portion of the track required to access the excluded grid cell.
[0094] In some examples, the exclusion zone may be increased to include a buffer area around the affected grid cell where the debris is located. In such cases, the buffer area may improve the effectiveness of the exclusion zone relative to excluding only grid cells immediately adjacent to the debris's mapped location on the grid. The size of the buffer area may be predetermined, for example, as a set area of grid cells that will apply once the immediate vicinity of the grid cell for exclusion is determined. Additionally or alternatively, the size of the buffer area is a selectable parameter when implementing the exclusion zone in the control system.
[0095] A control system, e.g., a master controller, remotely controlling movement of transport devices operating within the workspace can implement the exclusion zone based on the exclusion zone data output as part of the method. For example, each of the one or more transport devices 30 can be remotely operable under the control of a master control system, e.g., a central computer. To control movement of the one or more transport devices 30 on the grid 15, instructions can be sent from the master control system to the one or more transport devices 30 via a wireless communications network, e.g., implementing one or more base stations.
[0096] A controller in each transport device 30 is configured to control various drive mechanisms of the transport device, e.g., vehicle 32, to control its movement. For example, the instructions include various movements in the XY plane of the grid structure 15, which may be encapsulated in a defined trajectory for a given transport device. Thus, the exclusion zones may be implemented by a central control system, e.g., a master controller, such that the defined trajectory avoids the exclusion zone represented by the exclusion zone data. For example, when an exclusion zone is implemented, one or more respective trajectories corresponding to one or more transport devices 30 on the grid are updated to avoid the exclusion zone.
[0097] In examples, mapping a target image portion (e.g., one or more pixels in an image) to a target location (e.g., a point on a grid structure) involves inverting a distortion of the image of the workspace. For example, if a camera is equipped with a wide-angle or ultra-wide-angle lens, the lens distorts the view of the workspace. Thus, that distortion is reversed, e.g., as part of the mapping between image pixels and grid points. An inverse distortion model may be applied to the target image portion for this purpose. The description of the image-to-grid mapping algorithm in the previous example applies correspondingly here. For example, mapping the target image portion to a target grid location involves applying an image-to-grid mapping algorithm described herein.
[0098] In some examples, a check is made as to whether debris has been cleared from the grid so that the exclusion zone can be released, e.g., canceled, to allow free entry of the transport device into the corresponding grid cell. For example, further image data representing a further image of at least a portion of the workspace is acquired and processed using an object detection model trained to detect instances of debris on the grid. Based on the processing, it is determined whether the further image includes debris on the grid. In response to determining that the further image does not include debris on the grid, the exclusion zone is released, e.g., by a signal sent to the master controller.
[0099] The assistance system may be implemented to assist the control system in controlling the movement of the transport device in the workspace. An object detection system configured to detect debris in the workspace is, for example, part of the assistance system. An interface of the assistance system may obtain target image portions from the object detection system, as described in the examples. The assistance system is configured, for example, to map the target image portions to target locations in the workspace, determine exclusion zones, and output exclusion zone data. For example, the assistance system outputs exclusion zone data for a control system, e.g., a master controller, to receive as input and implement in the workspace. The exclusion zone data may be transferred directly between the assistance system and the control system or may be stored by the assistance system in storage accessible by the control system.
[0100] In embodiments employing an assistance system, the assistance system may be incorporated into a storage system 1, such as the example shown in FIG. 7A , including a workspace and a control system for controlling transport device movement within the workspace. As described in the example with reference to FIG. 7A , the workspace includes a grid 15 formed by a first set of parallel tracks 22 a extending in the X direction and a second set of parallel tracks 22 b extending in the Y direction, orthogonal to the first set in a substantially horizontal plane. The grid 15 includes a plurality of grid spaces 17, and one or more transport devices 30 are positioned to selectively move peripherally on the tracks 22 to handle containers 10 stacked below the tracks 22 within the footprint of a single grid space 17. Each transport device 30 may have a footprint that occupies only a single grid space 17, such that a given transport device occupying one grid space does not interfere with another transport device occupying or crossing an adjacent grid space.
[0101] In some cases, for example, instead of implementing an exclusion zone around the detected debris, a signal is sent to the master controller in response to determining that the image contains debris on the grid, causing the master controller to stop one or more transport devices. As described in the previous example, for example, if further classification of debris is performed, a shutdown option may be taken based on detecting a particular class of debris on the grid.
[0102] Deploying service devices As described in some examples, a robotic servicing device can be deployed to remove detected debris from the grid. For example, a detection method involves determining a target image portion of an image based on annotation data (as described in other examples) and mapping the target image portion to a target location in the workspace. A signal is output to deploy the servicing device to a target location in the workspace, e.g., on the grid. The servicing device is arranged to selectively move in at least one of a first or second direction on the track. For example, the servicing device, similar to a transport device, comprises a body mounted on two sets of wheels, where a first set of wheels is arranged to engage at least two tracks of the first set of tracks and a second set of wheels is arranged to engage at least two tracks of the second set of tracks. The first set of wheels is independently movable and drivable with respect to the second set of wheels such that only one set of wheels is engaged with the grid at any time, thereby enabling movement of the servicing device along the track to any point on the grid by driving only the set of wheels that engages with the rails. The robotic servicing device is provided with additional features in addition to those of the robotic transport device, namely, the servicing device comprises a cleaning mechanism comprising means for removing debris present on the grid, for example, the cleaning mechanism comprising at least one of a vacuum cleaning system (e.g., mounted adjacent each set of wheels), a brush mechanism (e.g., comprising one or more brushes), and a spray device capable of discharging a suitable detergent adapted to address contaminants on the grid.
[0103] Picking Station As described with respect to FIG. 6, the system may further include one or more robotic picking stations 50 mounted on the grid-based storage system 1.
[0104] FIG. 7B shows a schematic diagram of a detection system in this context. The grid-based storage system 1 is of a type previously described, e.g., an automated storage and retrieval system (or "ASRS"). In this embodiment, there are multiple robotic picking stations 50 mounted on the grid-based storage system 1, e.g., mounted on a grid structure (or simply "grid") 15 as previously described with respect to FIG. 6. Each picking station 50 includes a robotic manipulator 52 for transferring items between containers received in designated grid cells adjacent the respective picking station 50. For example, the robotic manipulator 52 includes an end effector for manipulating and releasably engaging items to be transferred between containers. The end effector may be a suction device connected to a negative pressure source, as in the embodiment shown in FIG. 7B, or another type of end effector, such as a jaw gripper or finger gripper.
[0105] In the embodiment shown in FIG. 7B , each robotic manipulator 52 is mounted on a pedestal above a single grid cell and is surrounded by eight grid cells. In other embodiments, a given robotic manipulator 52 may be surrounded by fewer grid cells or on fewer sides, depending on its location on storage system 1. Similarly, while FIG. 7B shows robotic picking stations 50 arranged along both orthogonal directions of grid 15, in other embodiments, picking stations 50 may be arranged along only one axis of grid 15, e.g., in a row or line. In some cases, there may be clusters of robotic picking stations 50 arranged at selected locations on grid 15 of storage system 1. As described in other examples, a camera 71 is disposed above grid 15, forming part of a detection system.
[0106] In this context, the detection system and method may determine whether detected debris on the grid is in a portion of the track adjacent to one or more designated grid cells of a robotic picking station 50 mounted on the grid 15. For example, a target image portion of the image is determined based on annotation data output from the initial detection and mapped to a target location in the workspace, as described in other examples. Based on the mapping, e.g., the determined location of the debris relative to the grid, it is determined whether the debris is in a portion of the track adjacent to one or more grid cells associated with the given picking station 50. In response to a positive determination, a signal is output to cause the robotic manipulator 52 of the given picking station 50 to remove the debris from the track.
[0107] The detection system described previously may be configured to perform any of the detection methods described herein. For example, the detection system includes an image sensor for capturing an image of at least a portion of the grid and an interface for acquiring the image data. The detection system includes a trained object detection model, implemented, for example, on a graphics processing unit (GPU) or a dedicated neural processing unit (NPU), for performing the processing and decision steps of the computer-implemented method 130 for detecting debris 122 on the grid 15.
[0108] The above examples should be understood as illustrative examples. Further examples are contemplated. For example, camera 71 disposed above grid 15 has been described in many examples as an ultra-wide-angle camera. However, camera 71 could be a wide-angle camera, which includes a wide-angle lens that has a relatively longer focal length than an ultra-wide-angle lens, but still introduces distortion compared to a normal lens that reproduces a field of view that appears "natural" to a human observer.
[0109] Similarly, described examples include acquiring and processing "images" or "image data." Such images may, in some cases, be video frames, e.g., selected from a video comprising a sequence of frames. The video may be captured by a camera positioned on a grid as described herein. Thus, acquiring and processing images should be interpreted as including acquiring and processing frames from a video, e.g., a video stream. For example, the described neural networks may be trained to detect instances of objects (e.g., debris) in a video stream comprising multiple images.
[0110] In examples employing storage to store data, the storage may be random access memory (RAM) such as DDR-SDRAM (double data rate synchronous dynamic random access memory). In other examples, the storage may include non-volatile memory such as read-only memory (ROM) or a solid-state drive (SSD) such as flash memory. The storage may in some cases include other storage media, for example, magnetic, optical, or tape media, compact discs (CDs), digital versatile discs (DVDs), or other data storage media. The storage may be removable or non-removable from the associated system.
[0111] In examples employing data processing, a processor may be employed as part of the system involved. The processor may be a general-purpose processor such as a central processing unit (CPU), microprocessor, graphics processing unit (GPU), digital signal processor (DSP), application specific integrated circuit (ASIC), field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any suitable combination thereof designed to perform the data processing functions described herein.
[0112] In examples involving neural networks, a dedicated processor may be employed as part of the system involved. The dedicated processor may be an NPU, a neural network accelerator (NNA), or other version of a hardware accelerator specialized for neural network functions. Additionally or alternatively, the neural network processing workload may be at least partially shared by one or more standard processors, e.g., a CPU or GPU.
[0113] Although the term “annotation data” is used throughout the description, it is assumed that this term is consistent with predicted or inferred data in alternative names. For example, an object detection model (e.g., comprising a neural network) may be trained using annotated images, e.g., images with annotations such as bounding boxes, that serve as ground truth for the model, e.g., predictions or inferences with a confidence of 100% or 1 when normalized. These annotations may be made by humans, e.g., for purposes of training the model. Thus, object detection of the present disclosure may be interpreted as outputting predicted or inferred data (e.g., instead of “annotation data”) to indicate a prediction or inference of debris in an image. The predicted or inferred data may be expressed as annotations, e.g., bounding boxes and / or labels, applied to the image. The predicted or inferred data may include, for example, a confidence associated with the prediction or inference of debris in the image. An annotation may be applied to the image, for example, based on the generated predicted or inferred data. For example, the image may be updated to include a bounding box surrounding the predicted debris with a label indicating the confidence level of the prediction, e.g., as a percentage value or a normalized value between 0 and 1.
[0114] It should also be understood that features described with respect to any one example may be used alone or in combination with other features described, and may also be used in combination with one or more features of any other of the examples, or in any combination of any other of the examples. Moreover, equivalents and modifications not described above may also be employed without departing from the scope of the appended claims.
Claims
1. 1. A computer-implemented method for detecting debris in a workspace comprising a grid formed by a first set of tracks extending in a first direction and a second set of tracks extending in a second direction orthogonal to the first direction, comprising: acquiring image data representing an image of at least a portion of the workspace; processing the image data with an object detection model trained to detect instances of debris on the grid; determining whether the image contains debris on the grid based on the processing; in response to determining that the image contains debris on the grid, outputting annotation data indicative of the debris in the image; A computer-implemented method comprising:
2. The method of claim 1 , comprising generating the annotation data.
3. The method of claim 1 or 2, comprising outputting an updated version of the image including the annotation data.
4. The method of claim 1 , wherein the annotation data comprises a bounding box.
5. The method of claim 1 , wherein the object detection model comprises a convolutional neural network.
6. one or more transport devices arranged to selectively move on said track in at least one of said first or second directions and to handle containers stacked below said track within the footprint of a single grid cell; The method comprises: determining a target image portion of the image based on the annotation data; mapping the target image portion to a target location in the workspace; determining an exclusion zone in the workspace comprising the target location where the one or more transport devices will be prohibited from entering based on the mapping; outputting exclusion zone data representing the exclusion zone to a control system for implementing the exclusion zone in the workspace; The method of any one of claims 1 to 5, comprising:
7. The method of claim 6 , wherein the exclusion zone comprises a plurality of grid cells adjacent to the debris detected on the grid.
8. acquiring further image data representing a further image of the at least a portion of the workspace; processing the further image data through the object detection model; and determining whether the further image contains debris on the grid based on the processing; disabling the exclusion zone in response to determining that the further image does not include debris on the grid. The method of claim 6 or 7, comprising:
9. one or more transport devices arranged to selectively move on said track in at least one of said first or second directions and to handle containers stacked below said track within the footprint of a single grid cell; The method comprises: and in response to determining that the image includes debris on the grid, outputting a signal to the master controller of the one or more transport devices to cause a master controller to stop the one or more transport devices. The method of any one of claims 1 to 5, comprising:
10. determining a target image portion of the image based on the annotation data; mapping the target image portion to a target location in the workspace; outputting a signal to deploy a servicing device at the target location, the servicing device being disposed to selectively move on the track in at least one of the first or second directions, the servicing device comprising a cleaning mechanism having means for removing debris present on the grid. The method of any of claims 1 to 9, comprising:
11. the workspace comprises one or more picking stations mounted on the grid, each picking station comprising a robotic manipulator for transferring items between containers received in respective grid cells adjacent the picking station; The method comprises: determining a target image portion of the image based on the annotation data; mapping the target image portion to a target location in the workspace; determining, based on the mapping, whether the debris detected on the grid is in a portion of the track adjacent to one or more grid cells associated with at least one of the one or more picking stations; in response to determining that the debris is in a portion of the track adjacent one or more grid cells associated with a given picking station of the one or more picking stations, outputting a signal to cause the robotic manipulator of the given picking station to remove the debris from the track; The method of any of claims 1 to 10, comprising:
12. responsive to determining that the image contains debris on the grid, processing the image data with one or more object classification models trained to classify debris; determining classification data representative of a class of debris to which the detected debris belongs based on said processing; and The method of any one of claims 1 to 5, comprising:
13. deploying a servicing device arranged to move on said track and comprising a cleaning mechanism having means for removing debris present on said grid in response to said classification data indicating that said detected debris belongs to a first class of debris; or and in response to the classification data indicating that the detected debris belongs to a second class of debris, shutting down any transport devices on the grid that are arranged to move on the truck to transport containers stacked below the truck between grid cells. The method of claim 12, comprising:
14. Data processing apparatus comprising means for performing the method according to any one of claims 1 to 13.
15. A computer program comprising instructions which, when said program is executed by a computer, cause said computer to perform the method of any one of claims 1 to 13.
16. 16. A computer readable data carrier storing a computer program according to claim 15.
17. 1. A detection system for detecting debris in a workspace, comprising: a grid formed by a first set of tracks extending in a first direction and a second set of tracks extending in a second direction orthogonal to the first direction; the detection system comprising: an image sensor for capturing an image of at least a portion of the workspace; an object detection model trained to detect instances of debris on the grid; Equipped with wherein the detection system comprises: obtaining image data representative of the image; processing the image data through the object detection model; determining whether the image contains debris on the grid based on the processing; in response to determining that the image contains debris on the grid, outputting annotation data indicative of the debris in the image; A detection system configured to:
18. 20. The detection system of claim 17, comprising a wide-angle or ultra-wide-angle camera comprising the image sensor.
19. 19. The detection system of claim 17 or 18, wherein the object detection model comprises a convolutional neural network.
Citation Information
Patent Citations
Route planning device, route planning method and traveling object
JP2009025898A
Obstacle detection device, obstacle detection system including the same, and obstacle detection method
JP2021174124A
Methods and systems for agriculture
JP2021510305A
Mobile device, method for moving mobile device, and program for controlling movement of mobile device
WO2009090807A1
Mechanical handling apparatus
WO2022243326A1