Detecting a transport device in a workspace

The use of ultra-wide-angle cameras and object detection models addresses the challenge of accurately locating load handling devices within grid frameworks, improving system efficiency and safety by enabling real-time monitoring and intervention for device misalignment or failure.

JP7778929B2Active Publication Date: 2025-12-02OCADO INNOVATION LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024531335
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-26
Filing Date
2022-11-25
Publication Date
2025-12-02
Estimated Expiration
2042-11-25

AI Technical Summary

Technical Problem

Existing storage and fulfillment systems face challenges in accurately determining the location of remotely operated load handling devices within a grid framework structure, particularly in cases of communication loss or device misalignment, which can lead to collisions and operational inefficiencies.

Method used

A method and system utilizing ultra-wide-angle cameras and a trained object detection model to monitor and detect the position of load handling devices on a grid framework, incorporating a calibration process to accurately map distorted images from the cameras to the physical grid, enabling independent assessment of device location and detecting defects or unresponsiveness.

Benefits of technology

Enables reliable detection and monitoring of load handling devices, allowing for timely intervention in cases of misalignment or failure, enhancing system efficiency and safety by providing an independent verification of device positions and trajectories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007778929000011
    Figure 0007778929000011
  • Figure 0007778929000012
    Figure 0007778929000012
  • Figure 0007778929000013
    Figure 0007778929000013
Patent Text Reader

Abstract

A method and system for detecting transport devices in a workspace comprising a grid, the grid comprising a plurality of grid spaces. One or more transport devices are arranged to selectively move in at least one of an X-direction or a Y-direction on a track and to handle containers stacked under the track within the footprint of a single grid space. Image data representative of an image of at least a portion of the workspace is acquired and processed by an object detection model trained to detect instances of the transport devices on the grid. Based on the processing, it is determined whether the image includes a transport device of the one or more transport devices. In response to determining that the image includes the transport device, annotation data indicative of the transport device in the image is output.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to the field of storage or fulfillment systems in which stacks of bins or containers are arranged within a grid framework structure, and more particularly to detecting transport devices in a workspace comprising a grid framework structure. [Background technology]

[0002] Online retail businesses that sell multiple product lines, such as online grocery stores and supermarkets, need systems that can store tens or even hundreds of thousands of different product lines. In such cases, the use of single-product stacks may be impractical because a huge amount of floor space would be required to accommodate all of the required stacks. Additionally, it may be desirable to store small quantities of some items, such as perishable or infrequently ordered goods, making single-product stacks an inefficient solution.

[0003] International Patent Application No. WO98 / 049076A (Autostore), the contents of which are incorporated herein by reference, describes a system in which a multi-product stack of containers is arranged within a frame structure.

[0004] PCT Publication No. WO2015 / 185628A (Ocado) describes a further known storage and fulfillment system in which stacks of containers are arranged within a grid framework structure. The containers are accessed by one or more load handling devices, sometimes known as "bots," operable on trucks on top of the grid framework structure. A system of this type is shown schematically in Figures 1 to 3 of the accompanying drawings.

[0005] As shown in FIGS. 1 and 2, stackable containers 10, also known as "bins," are stacked on top of each other to form a stack 12. The stacks 12 are arranged in a grid framework structure 14, for example, in a warehousing or manufacturing environment. The grid framework structure 14 consists of a plurality of storage rows or grid rows. Each grid in the grid framework structure has at least one grid row for storing a stack of containers. FIG. 1 is a schematic perspective view of the grid framework structure 14, and FIG. 2 is a schematic top-down view showing a stack 12 of bins 10 arranged within the framework structure 14. Each bin 10 typically holds multiple product items (not shown). The product items in the bins 10 can be of the same or different product types, depending on the application.

[0006] The grid framework structure 14 includes a plurality of upright members 16 supporting horizontal members 18, 20. A first set of parallel horizontal grid members 18 are arranged in a grid pattern, perpendicular to a second set of parallel horizontal members 20, to form a horizontal grid structure 15 supported by the upright members 16. The members 16, 18, 20 are typically fabricated from metal. The bins 10 are stacked between the members 16, 18, 20 of the grid framework structure 14 such that the grid framework structure 14 guards against horizontal movement of the stack 12 of bins 10 and guides vertical movement of the bins 10.

[0007] The top level of the grid framework structure 14 comprises a grid or grid structure 15 including rails 22 arranged in a grid pattern across the top of the stacks 12. Referring to FIG. 3 , the rails or tracks 22 guide a plurality of load handling devices 30. A first set 22a of parallel rails 22 guides movement of the robotic load handling devices 30 in a first direction (e.g., the X direction) across the top of the grid framework structure 14. A second set 22b of parallel rails 22, positioned orthogonal to the first set 22a, guides movement of the load handling devices 30 in a second direction (e.g., the Y direction) orthogonal to the first direction. In this manner, the rails 22 allow the robotic load handling devices 30 to move laterally in two dimensions in the horizontal XY plane. The load handling devices 30 can be moved to a position above any of the stacks 12.

[0008] A known form of load handling device 30, shown in Figures 4, 5, 6A and 6B, is described in PCT Patent Publication No. WO2015 / 019055 (Ocado), which is incorporated herein by reference, with each load handling device 30 covering a single grid space 17 of the grid framework structure 14. This configuration allows for a higher density of load handlers and therefore a higher throughput for a storage system of a given size.

[0009] The exemplary load handling device 30 includes vehicles 32 that are positioned to roll on the rails 22 of the frame structure 14. A first set of wheels 34, consisting of a pair of wheels 34 at the front of the vehicles 32 and a pair of wheels 34 at the rear of the vehicles 32, are positioned to engage two adjacent rails of the first set 22a of rails 22. Similarly, a second set of wheels 36, consisting of a pair of wheels 36 on each side of the vehicles 32, are positioned to engage two adjacent rails of the second set 22b of rails 22. At any time during the movement of the load handling device 30, each set of wheels 34, 36 can be raised and lowered so that either the first set of wheels 34 or the second set of wheels 36 is engaged with the respective set of rails 22a, 22b. For example, when a first set of wheels 34 is engaged with a first set of rails 22a and a second set of wheels 36 is lifted off the rails 22, the first set of wheels 34 can be driven by a drive mechanism (not shown) housed in the vehicle 32 to move the load handling device 30 in the X direction. To achieve movement in the Y direction, the first set of wheels 34 is lifted off the rails 22 and the second set of wheels 36 is lowered to engage with a second set 22b of rails 22. The drive mechanism can then be used to drive the second set of wheels 36 to move the load handling device 30 in the Y direction.

[0010] The load handling device 30 is equipped with a lifting mechanism, e.g., a crane mechanism, for lifting a storage container from above. The lifting mechanism includes a winch tether or cable 38 wound on a spool or reel (not shown) and a gripper device 39. The lifting mechanism shown in FIG. 5 includes a set of four vertically extending lifting tethers 38. The tethers 38 are connected at or near each of the four corners of the gripper device 39, e.g., a lifting frame, for releasable connection to the storage container 10. For example, each tether 38 is positioned at or near each of the four corners of the lifting frame 39. The gripper device 39 is configured to releasably grip the top of the storage container 10 to lift it from a stack of containers in a storage system 1 of the type shown in FIGS. 1 and 2. For example, the lifting frame 39 may include pins (not shown) that mate with corresponding holes (not shown) in a rim forming the top surface of the bin 10 and sliding clips (not shown) that are engageable with the rim to grip the bin 10. The clips are housed within a lifting frame 39 and are driven into engagement with the bins 10 by a suitable drive mechanism powered and controlled by signals carried through the cable 38 itself or a separate control cable (not shown).

[0011] To remove a bin 10 from the top of the stack 12, the load handling device 30 is first moved in the X and Y directions to position the gripper device 39 above the stack 12. The gripper device 39 is then lowered vertically in the Z direction to engage the bin 10 at the top of the stack 12, as shown in FIGS. 4 and 6B. The gripper device 39 grasps the bin 10 and is then pulled upward by the cable 38 with the bin 10 attached. At the top of its vertical travel, the bin 10 is held on the rails 22 housed within the vehicle body 32. In this way, the load handling device 30, carrying the bin 10 therewith, can be moved to different positions in the XY plane to transport the bin 10 to another location. Upon arriving at the target location (e.g., another stack 12, an access point in a storage system, or a conveyor belt), the bin or container 10 can be lowered from the container receiving portion and released from the grabber device 39. The cable 38 is long enough to allow the load handling device 30 to pick and place bins from any level of the stack 12, including, for example, floor level.

[0012] As shown in Figure 3, multiple identical load handling devices 30 are provided so that each load handling device 30 can operate simultaneously to increase system throughput. The system shown in Figure 3 may include specific locations, known as ports, where bins 10 can be transferred into or out of the system. An additional conveyor system (not shown) is associated with each port so that bins 10 transported to a port by a load handling device 30 can be transferred by that conveyor system to another location, such as a picking station (not shown). Similarly, bins 10 can be moved by the conveyor system from an external location to the port, for example, to a bin filling station (not shown), and transported by the load handling device 30 to the stacks 12 to replenish stock in the system.

[0013] Each load handling device 30 is capable of lifting and moving one bin 10 at a time. The load handling device 30 has a container receiving cavity or recess 40 in its lower portion. The recess 40 is sized to accommodate the container 10 when it is lifted by the lifting mechanisms 38, 39, as shown in Figures 6A and 6B. When in the recess, the container 10 is lifted off the lower rail 22 to allow the vehicle 32 to move laterally to different grid locations.

[0014] When it is necessary to remove a bin 10b that is not at the top of a stack 12 (a "target bin"), the bins 10a above (a "non-target bin") must first be moved to allow access to the target bin 10b. This is accomplished by an operation hereinafter referred to as "digging." Referring to FIG. 3, during a digging operation, one of the load handling devices 30 sequentially lifts each non-target bin 10a from the stack 12 containing the target bin 10b and places it in a vacant position in another stack 12. The target bin 10b can then be accessed by the load handling device 30 and moved to a port for further transport.

[0015] Each load handling device 30 is remotely operable under the control of a central computer, e.g., a master controller. Also, each individual bin 10 in the system is tracked so that the appropriate bin 10 can be removed, transported, and replaced as needed. For example, during a dig operation, each non-target bin location is logged so that the non-target bin 10a can be tracked.

[0016] Wireless communications and networks may be used to provide a communications infrastructure from a master controller, e.g., via one or more base stations, to one or more load handling devices 30 operable on the grid structure 15. In response to receiving instructions from the master controller, a controller in the load handling device 30 is configured to control various drive mechanisms to control movement of the load handling device. For example, the load handling device 30 may be instructed to retrieve a container from a target storage row at a specific location on the grid structure 15. The instructions may include various movements in the XY plane of the grid structure 15. As previously described, upon reaching the target storage row, the lifting mechanisms 38, 39 may be operated to grasp and lift the storage container 10. Once the container 10 is received in the container receiving space 40 of the load handling device 30, the container 10 is then transported to another location on the grid structure 15, e.g., a “drop-off port.” At the drop-off port, the container 10 is lowered to a suitable pick station to enable retrieval of any items in the storage container. Movement of the load handling device 30 on the grid structure 15 may also involve the load handling device 30 being commanded to move to a charging station, typically located on the periphery of the grid structure 15 .

[0017] To move the load handling devices 30 on the grid structure 15, each load handling device 30 is equipped with a motor for driving the wheels 34, 36. The wheels 34, 36 may be driven via one or more belts connected to the wheels or may be individually driven by motors integrated into the wheels. In the case of single-cell load handling devices (where the footprint of the load handling device 30 occupies a single grid cell 17), the motors for driving the wheels may be integrated into the wheels due to the limited availability of space within the vehicle body. For example, the wheels of a single-cell load handling device 30 are driven by respective hub motors. Each hub motor includes an outer rotor with multiple permanent magnets arranged to rotate about a wheel hub with coils forming an inner stator.

[0018] 1-6B has many advantages and is suitable for a wide range of storage and retrieval operations. In particular, it allows for very high density storage of products, providing a very economical way of storing a wide range of different items in the bins 10, while also allowing reasonably economical access to all of the bins 10 when needed for picking.

[0019] However, it is an object of the present disclosure to provide a method and system for reliably determining the correct location of a remotely operated load handling device in a storage system. Summary of the Invention

[0020] A method is provided for detecting transport devices in a workspace comprising a grid formed by a first set of parallel tracks extending in an X direction and a second set of parallel tracks extending in a Y direction orthogonal to the first set in a substantially horizontal plane, the grid comprising a plurality of grid spaces, wherein one or more transport devices are arranged to selectively move in at least one of the X or Y directions on the tracks and to handle containers stacked below the tracks within the footprint of a single grid space. The method comprises obtaining image data representing an image of at least a portion of the workspace, processing the image data with an object detection model trained to detect instances of transport devices on the grid, determining based on the processing whether the image includes a transport device of the one or more transport devices, and outputting annotation data indicative of the transport device in the image in response to determining that the image includes the transport device.

[0021] Also provided is a data processing apparatus comprising a processor configured to perform the method. Also provided is a computer program comprising instructions that, when executed by a computer, cause the computer to perform the method. Similarly, provided is a computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to perform the method.

[0022] Further provided is a system for detecting transport devices in a workspace having a grid formed by a first set of parallel tracks extending in an X direction and a second set of parallel tracks extending in a Y direction orthogonal to the first set in a substantially horizontal plane, the grid comprising a plurality of grid spaces, wherein one or more transport devices are arranged to selectively move on the tracks in at least one of the X direction or the Y direction and to handle containers stacked below the tracks within the footprint of a single grid space. The system includes an image sensor for capturing an image of at least a portion of the workspace, an interface for acquiring a target image portion of an image representation of the workspace, and an object detection model trained to detect instances of the transport devices on the grid. The system is configured to: acquire image data representing the image; process the image data with the object detection model; determine based on the processing whether the image includes a transport device of the one or more transport devices; and, in response to determining that the image includes the transport device, output annotation data indicative of the transport device in the image.

[0023] Broadly speaking, this description introduces a system and method for detecting transport devices operable in a workspace using a trained object detection model. This allows the workspace to be monitored and, for example, the location of detected transport devices to be determined. Thus, the system and method allow the position of a transport device in the workspace to be determined separately from information stored by a master controller that remotely controls the transport device. Having an independent assessment of the position of a given transport device relative to the workspace may also be useful for other technical purposes, such as monitoring a predetermined trajectory of a transport device in the workspace relative to its true position. Monitoring the workspace and detecting transport devices moving therein may also allow instances of transport device defects or unresponsiveness to be detected and / or acted upon to resolve the operation of a fleet of transport devices.

[0024] Embodiments will now be described, by way of example only, with reference to the accompanying drawings, in which like reference numerals designate the same or corresponding parts and in which: [Brief explanation of the drawings]

[0025] [Figure 1] 1 is a schematic diagram of a grid framework structure according to known systems; [Figure 2] Schematic of a top-down view showing a stack of bins arranged within the framework structure of FIG. 1. [Figure 3] 1 is a schematic diagram of a known storage system showing load handling devices operable on a grid framework structure. [Figure 4] 1 is a schematic perspective view of a load handling device on a portion of a grid framework structure. [Figure 5] 1 is a schematic perspective view of a load handling device showing a lifting mechanism for gripping a container from above. [Figure 6A]6 is a schematic perspective cutaway view of the load handling device of FIG. 5 showing the container receiving space of the load handling device and how it accommodates a container in use. [Figure 6B] 6 is a schematic perspective cutaway view of the load handling device of FIG. 5 showing the container receiving space of the load handling device and how it accommodates a container in use. [Figure 7] 1 is a schematic diagram of a storage system illustrating a load handling device operable on a grid framework structure along with a camera located above the grid framework structure, according to an embodiment. [Figure 8A] 1 is a diagram of a schematic representation of an image captured by a camera positioned above a grid framework structure, according to an embodiment. [Figure 8B] 1 is a diagram of a schematic representation of an image captured by a camera positioned above a grid framework structure, according to an embodiment. [Figure 9] 1 is a schematic diagram of a neural network, according to an embodiment. [Figure 10A] 1 is a schematic diagram of a generated model of a track of a grid framework structure, according to an embodiment. [Figure 10B] 1 is a schematic diagram of a generated model of a track of a grid framework structure, according to an embodiment. [Figure 11] 1 is a schematic diagram illustrating flattening of a captured image of a grid framework structure, according to an embodiment. [Figure 12] 10A-10C are schematic diagrams illustrating the modification of a captured image of a grid framework structure, according to an embodiment. [Figure 13] 10 is a flowchart illustrating a method for calibrating ultra-wide angle cameras disposed on a grid of a storage system, according to an embodiment. [Figure 14] 10 is a flowchart illustrating a method for detecting a transport device in a workspace comprising a grid, according to an embodiment. [Figure 15] 10 is a flowchart illustrating a method for detecting identification markers on a transport device in a workspace comprising a grid, according to an embodiment. [Figure 16]1 is a flow chart illustrating a method for assisting in controlling the movement of one or more transport devices operating in a workspace. DETAILED DESCRIPTION OF THE INVENTION

[0026] In a storage system of the type shown in FIGS. 1-3, it is useful to determine the location of a given load handling device 30 operating on the grid structure 15 independently of the master controller. Each load handling device 30 receives control signals from the master controller to move along a predetermined path from one location on the grid structure to another. For example, a given load handling device 30 may be commanded to move to a particular location on the grid structure 15 and lift a target container from a stack of containers at that particular location. When multiple such devices 30 move along respective tracks on the grid structure 15, it is useful to be able to determine the precise location of a given load handling device 30 relative to the grid structure 15, for example, in the event that communication between a given load handling device and the master controller is lost. For example, a collision between a load handling device and another object, such as another load handling device, on or around the grid structure may cause the, or each, load handling device to become unresponsive to communications from the master controller, for example, by losing connection thereto and / or disengaging from tracks 22 of grid structure 15. The collision may, for example, cause one or more load handling devices to become misaligned with tracks 22 or to fall over on the grid structure.

[0027] Monitoring the grid structure 15 and the load handling devices 30 moving thereon can allow for unresponsive instances of the load handling devices to be detected and / or acted upon to resolve the operation of multiple load handling devices. Having an independent assessment of the position of a given load handling device 30 relative to the grid structure 15 can also be useful for other technical purposes, such as monitoring a predetermined trajectory of a load handling device 30 on the grid structure 15 against its true position.

[0028] FIG. 7 shows the previously described grid structure (or simply "grid") 15 of the storage system. The grid is formed by a first set 22a of parallel tracks extending in the X direction and a second set 22b of parallel tracks extending in the Y direction, orthogonal to the first set in a substantially horizontal plane. The grid 15 has a plurality of grid spaces 17. One or more load handling devices, or "transport devices" 30, are arranged to selectively move in at least one of the X or Y directions on the tracks 22 and to handle containers 10 stacked below the tracks 22 within the footprint of a single grid space 17. In the example, the one or more transport devices 30 each have a footprint that occupies only a single grid space, such that a given transport device occupying one grid space does not interfere with another transport device occupying or crossing an adjacent grid space.

[0029] A camera 71 is disposed above the grid 15. In an example, the camera 71 is an ultra-wide-angle camera, i.e., it has an ultra-wide-angle lens (also referred to as a "super wide-angle" lens or a "fisheye" lens). The camera 71 includes an image sensor for receiving incident light focused through a lens, e.g., a fisheye lens. The camera 71 has a field of view 72 that includes at least a section of the grid 15. Multiple cameras may be used to observe the entire grid 15, e.g., each camera 71 having a respective field of view 72 that covers a section of the grid 15. The ultra-wide-angle lens may be selected because of its relatively large field of view 72, e.g., up to a 180-degree solid angle, compared to other lens types, which means that fewer cameras are needed to cover the grid 15. Also, space may be limited between the top of the grid 15 and surrounding structures, e.g., a warehouse roof, thus constraining the height of the camera 71 above the grid 15. An ultra-wide lens camera can provide a relatively large field of view at a relatively low height above the grid 15 compared to other camera types.

[0030] One or more cameras 71 may be used to monitor the workspace of transport device 30, which workspace includes grid structure 15. For example, image feed from one or more cameras 71 may be displayed on one or more computer monitors remote from grid 15 to monitor for instances of faulty, e.g., unresponsive, transport devices. An operator may thus detect such instances and act to resolve the issue, for example, by resetting a communication link between the transport device and the master controller or by requesting manual intervention for a mechanical problem.

[0031] An effective monitoring or surveillance system for a workspace incorporates calibration of one or more ultra-wide-angle cameras positioned above the workspace. Accurate calibration of the ultra-wide-angle camera allows interactions with images captured by the ultra-wide-angle camera, distorted by the ultra-wide-angle lens, to be correctly mapped to the workspace. Thus, an operator can select an area of ​​pixels in the distorted image, which is mapped to a corresponding area of ​​grid space in the workspace, for example. In another scenario, the distorted image from the camera 71 can be processed to detect a defective transport device 30 in the workspace and output its location in the workspace and, further, identification information for the detected transport device 30, such as a unique ID label. Such an example is described in the following embodiments.

[0032] Calibration Process Calibration process 130 includes acquiring 131 an image of a section of grid 15, i.e., a grid section, captured by ultra-wide-angle camera 71, according to the example shown in FIG. 13. Acquiring the image includes obtaining, e.g., receiving, image data representing the image, e.g., in a processor. For example, the image data may be received via an interface, e.g., a camera serial interface (CSI). An image signal processor (ISP) may perform initial processing of the image data, e.g., saturation correction, re-normalization, white balancing, and / or demosaicing, to prepare the image data for display.

[0033] Initial values ​​for multiple parameters corresponding to the ultra-wide-angle camera 71 are also obtained 132. These parameters include the focal length of the ultra-wide-angle camera, a translation vector representing the position of the ultra-wide-angle camera above the grid section, and a rotation vector representing the tilt and rotation of the ultra-wide-angle camera. These parameters can be used in a mapping algorithm to map pixels in an image distorted by the ultra-wide-angle lens of camera 71 onto a plane oriented relative to the Cartesian grid 15 of the storage system. The mapping algorithm is described in more detail below.

[0034] The calibration process 130 includes processing 133 the images using a neural network trained to detect / predict tracks in images of grid sections captured by the ultra-wide angle camera.

[0035] Neural Networks 9 illustrates an example of a neural network architecture. The exemplary neural network 90 is a convolutional neural network (CNN). One example of a CNN is the U-Net architecture developed by the Department of Computer Science at the University of Freiburg, although other CNNs, such as the VGG-16 CNN, can be used. The input 91 to the CNN 90 comprises image data in this example. The input image data 91 is a given number of pixels wide and a given number of pixels high and includes one or more color channels (e.g., red, green, and blue channels).

[0036] The convolutional layers 92, 94 of the CNN 90 may generally extract specific features from the input data 91 and operate on small portions of the image to create feature maps. The fully connected layer 96 uses the feature maps to determine an output 97, e.g., classification data specifying the classes of objects predicted to be present in the input image 91.

[0037] In the example of FIG. 9 , the output of the first convolutional layer 92 undergoes pooling in a pooling layer 93 before being input to a second convolutional layer 94. Pooling, for example, allows values ​​for a region of an image or feature map to be aggregated or combined, e.g., by taking the highest value within the region. For example, in 2×2 max pooling, rather than transferring the entire output, the highest value of the output of the first convolutional layer 92 within a 2×2 pixel patch of the feature map output from the first convolutional layer 92 is used as input to the second convolutional layer 94. Pooling can therefore reduce the amount of computation for subsequent layers of the neural network 90. ​​The effect of pooling is shown schematically in FIG. 9 as a reduction in the size of the frames in the relevant layers. Further pooling is performed in a second pooling layer 95 between the second convolutional layer 94 and the fully connected layer 96. It should be appreciated that the schematic representation of neural network 90 in FIG. 9 is greatly simplified for ease of explanation, and that typical neural networks can be significantly more complex.

[0038] Generally, a neural network, such as the neural network 90 of FIG. 9, may undergo what is called a "training phase," during which the neural network is trained for a specific purpose. A neural network generally includes layers of interconnected artificial neurons that form a directed, weighted graph, where the graph's vertices (corresponding to neurons) or edges (corresponding to connections) are each associated with a weight. The weights may be adjusted throughout training, changing the output of individual neurons and, therefore, the neural network as a whole. In a CNN, a fully connected layer 96 generally connects every neuron in one layer to every neuron in another layer and may thus be used to identify global properties of an image, such as whether the image contains a particular class of object or a particular instance of a particular class.

[0039] In the present context, neural network 90 is trained to perform object identification by processing image data, e.g., to determine whether an object of a predetermined class of objects is present in the image (although in other examples, neural network 90 may instead be trained to identify other image characteristics of the image). For example, training neural network 90 in this manner generates weight data representing weights to be applied to the image data (e.g., different weights associated with different layers of a multi-layer neural network architecture). Each of these weights is multiplied by the corresponding pixel value of the image patch, e.g., to convolve the weight kernel with the image patch.

[0040] Specific to the context of ultra-wide-angle camera calibration, neural network 90 is trained using a training set of input images of grid sections captured by an ultra-wide-angle camera to detect tracks 22 of grid 15 in a given image of the grid section. In an example, the training set includes mask images that show only extracted track features corresponding to the input images. For example, the mask images are manually created. Thus, the mask images can serve as a desired result for neural network 90 to be trained using the training set of images. Once trained, neural network 90 can be used to detect tracks 22 in images of at least a portion of grid structure 15 captured by an ultra-wide-angle camera.

[0041] The calibration process 130 includes processing 133 the image of the grid section captured by the ultra-wide-angle camera 71 using the trained neural network 90 to detect tracks 22 in the image. At least one processor (e.g., a neural network accelerator) may be used to perform the processing 133. The image processing 133 generates models of the tracks, in particular a first set and a second set of parallel tracks, captured in the image of the grid section. For example, the models comprise representations of predictions of the tracks in the distorted image of the grid section determined by the neural network 90. ​​The track models correspond, in examples, to masks or probability maps.

[0042] Selected pixels in the determined track model are then mapped 134 to corresponding points on the grid 15 using a mapping, e.g., a mapping algorithm, that incorporates multiple parameters corresponding to the ultra-wide-angle camera. The obtained initial values ​​are used as input to the mapping algorithm.

[0043] An error function (or "loss function") is determined 135 based on the discrepancy between the mapped grid coordinates and the "true," e.g., known, grid coordinates of the points corresponding to the selected pixels. For example, a selected pixel at the center of X-direction track 22a should correspond to a grid coordinate with a half-integer value in the Y direction, e.g., (x, y.5), where x is an unknown number and y is an unknown integer. Similarly, a selected pixel at the center of Y-direction track 22b should correspond to a grid coordinate with a half-integer value in the X direction, e.g., (x'.5, y'), where x' is an unknown integer and y' is an unknown number. In an example, the width and length of a grid cell (or their ratio) are used in the loss function, for example, to calculate cell x,y coordinates for keypoints and determine whether they are on the track (e.g., coordinate values ​​of n.5, where n is an integer).

[0044] The initial values ​​of the multiple parameters corresponding to the ultra-wide-angle camera are then updated to updated values ​​based on the determined error function 136. For example, a Broyden-Fletcher-Goldfarb-Schanno (BFGS) algorithm is applied using the error function and the initial parameter values ​​as input. In an example, the updated values ​​of the multiple parameters are determined iteratively, and the error function is recalculated with each update. The iterations may continue, for example, until the error function is reduced by less than a predetermined threshold between successive iterations or compared to the initial error function, or until the absolute value of the error function falls below a predetermined threshold. Other iterative algorithms, such as sequential quadratic programming (SQP) or sequential least-squares quadratic programming (SLSQP), may be used with the initial values ​​to generate a sequence of improving approximate solutions for the multiple parameters, with a given approximation in the sequence being derived from the previous one. In some cases, the iterative algorithm is used to optimize the values ​​of the multiple parameters. For example, the updated values ​​are optimized values ​​of the multiple parameters.

[0045] Updating 136 the initial values ​​of the plurality of parameters corresponding to the ultra-wide-angle camera involves applying one or more respective boundary values ​​for the plurality of parameters. For example, the boundary values ​​for the rotation angle associated with the rotation vector are substantially 0 degrees and substantially +5 degrees. Additionally or alternatively, the boundary values ​​for the planar component of the translation vector are ±0.6 of the length of the grid cell. Additionally or alternatively, the boundary values ​​for the height component of the translation vector are 1800 mm and 2100 mm, or 1950 mm and 2550 mm, or 2000 mm and 2550 mm above the grid. For example, the lower boundary for the camera height is within the range of 1800 to 2000 mm. For example, the upper boundary for the camera height is within the range of 2100 to 2600 mm. Additionally or alternatively, the boundary values ​​for the focal length of the camera are 0.23 and 0.26 cm. Applying one or more respective boundary values ​​for multiple parameters can mean that the updating, e.g., optimization, process is performed in a feasible region or solution space, i.e., the set of all possible values ​​that satisfy one or more boundary conditions.

[0046] The updated values ​​of the plurality of parameters are electronically stored 137 for future mapping of pixels in a grid section image captured by the ultra-wide-angle camera 71 to corresponding points on the grid 15 via a mapping algorithm. For example, the stored values ​​of the plurality of parameters are retrieved from data storage and used in the mapping algorithm to calculate grid coordinates corresponding to a given pixel in a given image of the grid section captured by the ultra-wide-angle camera 71. In an example, the updated values ​​are stored in a storage location associated with the ultra-wide-angle camera 71, for example, in a database. For example, a lookup function or table may be used in conjunction with the database to find the stored parameter values ​​associated with any given ultra-wide-angle camera employed in the storage system 1 on the grid 15.

[0047] Following calibration of a given camera 71 disposed over the grid 15, an image (e.g., a “snapshot”) of a grid section captured by the camera 71 may be flattened, i.e., undistorted, for interaction by an operator. For example, using the described image-to-grid mapping function, a distorted image 81 of the grid section may be converted into a flattened image 111 of the grid section, as shown in the example of FIG. 11 . Flattening involves selecting an area of ​​grid cells to flatten in the distorted image 81 and inputting the grid coordinates corresponding to those cells into the mapping function, which determines which respective pixel values ​​from the distorted image 81 should be copied into the flattened image 111 for each grid coordinate. For example, a target resolution, in pixels per grid cell, may be set for the flattened image 111, the target resolution having a ratio corresponding to the ratio of the grid cell dimensions. Once all pixel values ​​needed in the flattened image (according to the target resolution and the selected number of grid cells) have been determined, the flattened image 111 can be generated.

[0048] Snapshots may be captured by the camera 71 at predetermined intervals, e.g., every 10 seconds, and converted to corresponding flattened images 111. The most recent flattened images 111 may be stored in storage for viewing on a display, e.g., by an operator wishing to view the grid section covered by the camera 71. The operator may instead choose to retake a grid-shot snapshot and flatten it. Thus, the operator may select regions, e.g., pixels, in the flattened image 111 and convert those selected regions to grid coordinates based on the image-to-grid mapping functionality described herein. In some cases, the flattened image 111 includes grid coordinate annotations for the grid space viewable in the flattened image 111. The flattened images 111 corresponding to each camera 71 may be more user-friendly for monitoring the grid 15 compared to the distorted images 81, 82.

[0049] Grid to image mapping A computational algorithm maps real-world points on grid 15 to pixels in an image captured by the camera. The grid points are first projected onto a plane corresponding to ultra-wide-angle camera 71. For example, at least one of a rotation using a rotation matrix and a planar translation in the X and Y directions is applied to points having x, y, and z coordinates in grid framework structure 14. The focal length f of the ultra-wide-angle camera may be used to project points with three-dimensional coordinates relative to grid 15 onto a two-dimensional plane relative to ultra-wide-angle camera 71. For example, the coordinates of a mapped point q in the plane of ultra-wide-angle camera 71 are given by q=f·p [x,y] ÷p z It is calculated as, where p [x,y] and p z are the planar xy coordinates and the third z coordinate of point p relative to grid 15, respectively.

[0050] A point q projected onto the ultra-wide-angle camera plane may be aligned with a Cartesian coordinate system in that plane to determine a first Cartesian coordinate of the point. For example, aligning a point with a Cartesian coordinate system involves rotating the point or its position vector in the plane (e.g., a vector from the origin to the point). Thus, the rotation is, for example, to align with a typical grid orientation in an image captured by the camera, but may not be necessary if the X and Y directions of the grid are already aligned with the captured image. In the example, the rotation is substantially 90 degrees. As shown in FIGS. 8A and 8B, the X and Y directions of the grid are offset by 90 degrees with respect to the horizontal and vertical axes of the image; therefore, the rotation "corrects" for this offset so that the X and Y directions of the grid are aligned with the horizontal and vertical axes of the captured image.

[0051] A grid-to-image mapping algorithm continues by converting the first Cartesian coordinate to a first polar coordinate using standard trigonometry. A distortion model is then applied to the first polar coordinate of the point to generate a second, e.g., "distorted," polar coordinate. In an example, the distortion model comprises a tangent model of distortion given by r' = f arctan(r / f), where r and r' are the undistorted and distorted radial coordinates of the point, respectively, and f is the focal length of the ultra-wide-angle camera.

[0052] The second polar coordinates are then converted back to (second) Cartesian coordinates using the same standard trigonometric methods inversely. Image coordinates of pixels in the image are then determined based on the second Cartesian coordinates. In examples, this determination includes at least one of de-centering or rescaling the second Cartesian coordinates. Additionally or alternatively, the ordinate (y-coordinate) of the second Cartesian coordinates is flipped, e.g., mirrored on the x-axis.

[0053] Image to grid mapping Mapping pixels in the image captured by camera 71 to real-world points on grid 15 is done by different computational algorithms. For example, the image-to-grid mapping algorithm is the inverse of the grid-to-image mapping algorithm described above, with each mathematical operation reversed.

[0054] For a given pixel in the image, a (second) Cartesian coordinate of the mapped point is determined based on the image coordinate of the pixel in the image. For example, this determination may involve initializing the pixel in the image, including, for example, at least one of centering or normalizing the image coordinate. As mentioned above, the ordinate coordinate is inverted in some instances. The second Cartesian coordinate is converted to a second polar coordinate using standard trigonometry as described above. The use of the label "second" is used for consistency with the conversion performed in the grid-to-image algorithm described, but is arbitrary.

[0055] An inverse distortion model is applied to the second polar coordinates to generate first, e.g., "undistorted," polar coordinates. In an example, the inverse distortion model is based on a tangent model of distortion given by r=f·tan(r' / f), where again r' is the distorted radial coordinate of the point, r is the undistorted radial coordinate of the point, and f is the focal length of the ultra-wide-angle camera. Thus, in an example, the inverse distortion model used in the image-to-grid mapping is the inverse, or "anti-function," of the distortion model used in the grid-to-image mapping.

[0056] The image-to-grid mapping algorithm continues by converting the first polar coordinate to a first Cartesian coordinate. The first Cartesian coordinate may be disaligned or misaligned with a Cartesian coordinate system in a plane corresponding to the ultra-wide-angle camera. For example, disaligning a point with a Cartesian coordinate system involves applying a rotation transformation to the point or its position vector in the plane (e.g., a vector from the origin to the point). The rotation is substantially 90 degrees in the example. This rotation may therefore "undo" any "correction" to the offset between the X and Y directions of the grid and the horizontal and vertical axes of the captured image, as previously described in the grid-to-image mapping.

[0057] Finally, the point is projected from the (second) plane corresponding to the camera 71 onto the (first) plane corresponding to the grid 15 to determine the grid coordinates of the point relative to the grid.

[0058] In the example, the projection of the point onto the plane corresponding to grid 15 is p=B -1 This involves calculating (f tq z), where B=q R 3,[1,2] -f·R [1,2],[1,2]In these equations, p comprises the point coordinate in the grid plane, q comprises the Cartesian coordinate in the camera plane, and f is the focal length of the ultra-wide-angle camera as described above. Furthermore, t is a planar translation vector, z is the distance (e.g., height) between the ultra-wide-angle camera and the grid, and R is a three-dimensional rotation matrix related to the rotation vector. The rotation vector comprises a direction representing the axis of rotation and a magnitude representing the angle of rotation. The rotation matrix R corresponding to the angle-axis rotation vector can be determined from the vector using, for example, Rodriguez's rotation formula.

[0059] Next, a mathematical derivation of the function for projecting an undistorted 2D point q from the camera plane is provided for completeness: starting from the projection from the grid onto the image from above, q = f p' [x,y] ÷p' z where p' is the rotated and translated grid point p, i.e., p'=R·p+(t x ,t y ,z) T The goal is to derive p from q. After rearranging and substituting for p', we get the following:

[0060]

number

[0061]

number

[0062]

number

[0063]

number

[0064]

number

[0065] Since the desired distance of a point p on the grid from the camera is given by the height parameter z, in translating the point p z = 0. Therefore, for all p z Terms can be eliminated to give:

[0066]

number

[0067]

number

[0068] Matrix B=(q·R 3,[1,2] By defining -f·R), the equation becomes B·p [x,y] = f tz q, which can be further simplified to -1 Using

[0069] Returning to calibration process 130, in some cases, grid cell coordinate data encoded in grid cell markers positioned relative to grid 15 may be used to calibrate calculated grid coordinates corresponding to pixels in the captured image. For example, the grid cell markers may be signboards placed in predetermined grid cells 17, with corresponding cell coordinate data marked on each signboard. Process 130 may include, for example, processing the captured image to detect the grid cell markers in the image and then extracting the grid cell coordinate data encoded in the grid cell markers for use in calibrating the mapped grid coordinates. Each grid cell marker is in a respective grid cell, e.g., under and within the field of view 72 of a respective camera 71.

[0070] The image processing may involve using an object detection model, e.g., a neural network, trained to detect instances of grid cell markers in images of the grid section. A computer vision platform, e.g., Cloud Vision API (Application Programming Interface) by Google®, may be used to implement the object detection model. The object detection model may be trained using images of the grid section including the grid cell markers. In examples where the object detection model includes a neural network, e.g., a CNN, the description with reference to FIG. 9 applies accordingly.

[0071] Grid coordinates generated by mapping pixels in an image to points on a grid section represented in a captured image can be calibrated to the entire grid based on the extracted cell coordinate data. For example, a mapped grid point corresponding to a given pixel has coordinates in units of grid cells, e.g., (x, y), with x being the number of grid cells in the X direction and y being the number of grid cells in the Y direction. However, the grid cells captured by camera 71 are for a grid section, i.e., a section of grid 15, and therefore not necessarily the entire grid 15. Therefore, the mapped grid coordinates (x, y) for a grid section captured in an image can be calibrated to grid coordinates (x', y') for the entire grid based on the relative location of the grid section with respect to the entire grid. The location of the grid section with respect to the entire grid can be determined by extracting the grid cell coordinate data encoded in the grid cell markers captured in the image, as described.

[0072] 10A shows an exemplary model 101 of tracks generated by processing an image 81 of a grid section captured by an ultra-wide-angle camera 71 using a neural network 90 trained to detect tracks 22 in the image. Model 101 comprises a representation of a prediction of tracks 22 a, 22 b in a distorted image of the grid section determined by neural network 90. ​​Mapping pixels from track model 101 to corresponding points on grid 15 may be performed to calibrate camera 71 as described. For example, calibration may involve updating, e.g., optimizing, multiple parameters associated with camera 71 used for mapping between pixels in captured images 81, 82 and points on grid 15.

[0073] In an example, the model 101 of the grid sections may be refined to represent only the centerlines of the first set 22a and the second set 22b of parallel tracks. Thus, the pixels to be mapped from the track model 101 to corresponding points on the grid 15 are, for example, pixels that lie on the centerlines of the first set 22a and the second set 22b of parallel tracks in the generated model 101. Refining involves, for example, filtering the model with horizontal and vertical line detection kernels. These kernels allow the centerlines of the tracks to be identified in the model 101, for example, in the same way that other kernels may be used to identify other features of an image, such as edges in edge detection. Each kernel is a given size, for example a 3x3 matrix, that may be convolved with the image data in the model 101 with a given stride. For example, the horizontal line detection kernel may be the matrix

[0074]

number

[0075] It can be expressed as:

[0076] Similarly, the vertical line detection kernel can be, for example, the matrix

[0077]

number

[0078] It can be expressed as:

[0079] In the example, filtering involves at least one of eroding and dilating pixel values ​​of the model 101 using horizontal and vertical line detection kernels. For example, at least one of an erosion function and a dilation function is applied to the model 101 using the kernel. The erosion function effectively "erodes" foreground objects, in this case, the boundaries of tracks 22 a, 22 b in the generated model 101, by convolving the kernel with the model. During erosion, a pixel value (either "1" or "0") in the original model is updated to a value of "1" only if all pixels convolved under the kernel are equal to "1"; otherwise, it is eroded (updated to a value of "0"). Effectively, all pixels near the boundaries of tracks 22 a, 22 b in the model 101 are discarded, depending on the size of the kernel used in the erosion, such that the thickness of each of tracks 22 a, 22 b is reduced substantially to its centerline. A dilation function, which is the opposite of the erosion function, may be applied after erosion to effectively "inflate" or widen the centerlines remaining after erosion. This dilation may stabilize the centerlines of the tracks 22a, 22b in the improved model 101. During dilation, if at least one pixel convolved under the kernel is equal to "1," the pixel value is updated to a value of "1." The erosion and dilation functions are each applied to the original generated model 101, and the resulting horizontal and vertical centerline "skeleton," for example, are combined to produce the improved model.

[0080] In some cases, the generated model 101 may have missing sections of the tracks 22a, 22b, for example, obscuring one or more areas of the grid section viewable by the camera 71. Objects on the grid 15, such as a conveying device 30, a pillar, or other structure, may obscure portions of the tracks in the captured image. Thus, the generated model 101 may have the same missing areas of the tracks. Similarly, false positive predictions of the tracks may be present in the generated model 101.

[0081] To help with these issues, tracks 22 a, 22 b (e.g., their centerlines) present in the generated model can be fitted to respective quadratic equations, for example, to create secondary trajectories for tracks 22 a, 22 b. FIG. 10B shows an example of tracks from a first set of tracks 22 a in model 101 being fitted to a first secondary trajectory 102 and tracks from a second set of tracks 22 b in model 101 being fitted to a second secondary trajectory 103. Secondary track centerlines can then be created based on the secondary trajectories, for example, by extrapolating pixel values ​​along the secondary trajectories to fill gaps in model 101 or remove false positives. For example, if a sub-line generated from the predicted grid model 101 cannot be fitted to a given quadratic curve along with at least one other line, then the sub-line is highly unlikely to be part of the grid and should be excluded.

[0082] The quadratic equation used to fit the truck in Model 101, y=ax 2 +bx+c also depends on the specified boundary conditions, e.g.

[0083]

number

[0084] , -9.9×10 -4 <a<9.9×10 -4may have -5 < b < 5 and 0 < c < 3200.

[0085] In an example, a predetermined number of pixels are extracted from an improved model 101 of a track, for example, to reduce storage requirements for storing the model. For example, a random subset of pixels is extracted to provide the final improved model 101 of the track.

[0086] Calibrating the ultra-wide-angle camera 71 using the systems and methods described herein enables, for example, an image captured by the camera 71 having a wide field of view of the grid 15 to be used to detect the transport device thereon and identify its location. This is despite the relatively high distortion present in the image compared to those of other camera types.

[0087] The automatic calibration process outlined above can also reduce the time taken to calibrate each camera 71 installed on the grid 15 of the storage system compared to manual methods of tuning the parameters associated with each camera 71. For example, combining a neural network model, such as a U-Net, with a customized optimization function to implement the described calibration pipeline can remove over 80% of the error compared to standard calibration methods. Further, the calibration systems and methods described herein have proven to be sufficiently versatile and consistent to calibrate cameras in multiple warehouse storage systems, for example, having different dimensions, scales, and layouts.

[0088] Further, the flattened calibrated image 111 output by the grid enables easier interaction with the image 111 by both humans and machines for monitoring the grid 15 and the transport device 30 moving thereon. Thus, it may be more efficient to detect and / or act on an unresponsive instance of a given transport device on the grid to resolve the operation of a group of transport devices of the transport device 30.

[0089] Detecting a transport device in a workspace Methods and systems are provided herein for processing distorted images 82 captured by camera 71 to detect transport devices 30 on the grid. For example, the location of the detected transport devices 30 relative to the grid 15 may be output. In some examples, identification information, such as a unique ID label, of the detected transport devices 30 may be output. Such examples are now described in more detail.

[0090] 14 illustrates a computer-implemented method 140 for detecting a transport device 30 in a workspace comprising a grid 15. The method 140 involves obtaining 141 image data representing an image of at least a portion of the workspace and processing 142 the image data with an object detection model trained to detect instances of the transport device on the grid. For example, the image is captured by a camera 71 with a field of view covering at least a portion of the workspace, and the image data is transferred to a computer for implementing the detection method 140. The image data is received, for example, at an interface of the computer, e.g., a CSI.

[0091] The object detection model may be a neural network, e.g., a convolutional neural network, trained to perform object detection of the transport device 30 on the grid 15 of the workspace. Accordingly, the description of the neural network with respect to FIG. 9 applies in these specific examples. In the present context, the object detection model, e.g., a CNN 90, is trained to perform object identification by processing acquired image data to determine whether an object (i.e., the transport device) of a predetermined class of objects is present in the image. Training the neural network 90 involves, for example, providing the neural network 90 with training images of the workspace section in which the transport device is present. Weight data is generated for each (convolutional) layer 92, 94 of the multi-layer neural network architecture and stored for use in implementing the trained neural network. In the example, the object detection model comprises a “You Only Look Once” (YOLO) object detection model, e.g., YOLOv4 or Scaled-YOLOv4, having a CNN-based architecture. Other exemplary object detection models include neural-based approaches, such as RetinatNet or R-CNN (Regions with CNN features), and non-neural approaches, such as support vector machines (SVMs) for object classification based on determined features, e.g., Haar-like features or Histogram of Oriented Gradients (HOG) features.

[0092] Method 140 involves determining 143 whether the image includes the carrying device 30 based on process 142. For example, an object detection model is configured, e.g., trained or learned, to detect whether the carrying device 30 is present in a captured image of the workspace. In an example, the object detection model makes determination 143 with a level of confidence, e.g., a probability score, corresponding to the likelihood that the image includes the carrying device 30. Thus, a positive determination may correspond to a confidence level above a predetermined threshold, e.g., 90% or 95%. In response to determining 143 that the image includes the carrying device, annotation data (e.g., predicted data or inferred data) indicating the predicted carrying device in the image is output 144. An updated version of the image including the annotation data may be output, for example, as part of method 140.

[0093] In an example, the annotation data comprises a bounding box. FIG. 12 shows an example of an updated version 83 of an image 82 captured by the ultra-wide-angle camera 71, annotated with bounding boxes 120a, 120b. The bounding boxes 120a, 120b correspond to the first and second transport devices 30a, 30b, respectively, detected by the object detection model. A given bounding box may comprise, for example, a rectangle surrounding a detected object and may specify one or more of a location, an identified class (e.g., a transport device), and a confidence score (e.g., how likely the object will be present within the box). The bounding box data defining a given bounding box may include coordinates of two corners of the box or center coordinates with width and height parameters for the box. In an example, the detection method 140 involves generating the annotation data 120a, 120b for output.

[0094] In some cases, the object detection model is further trained to detect instances of faulty transport devices in the workspace, for example, transport devices that are unresponsive to communications from the master controller and / or are misaligned with the grid and / or have engaged warning signals.

[0095] For example, method 140 may involve processing image data with an object detection model and, based on the processing, determining whether the image includes a transport device misaligned with the grid. The object detection model may be the same or different from the one used to detect the transport device. The object detection model may be trained, for example, with a training set of images of transport devices misaligned with grid 15, e.g., at angles offset from orthogonal track 22. In response to determining that the image includes a misaligned transport device, the method may include outputting at least one of annotation data or an alert. The annotation data may, for example, indicate a predicted misaligned transport device in the image. As described above, the annotation may comprise a bounding box surrounding the predicted misaligned transport device on the grid. The output alert, for example, signals that the image includes a misaligned transport device.

[0096] Similarly, method 140 may involve processing image data with an object detection model and, based on the processing, determining whether the image includes a carrying device with an activated warning signal. The carrying device's warning signal may comprise a predetermined light or color of light emitted by a light source on the carrying device, such as a light-emitting diode (LED). For example, the carrying device may include an LED configured to emit a first wavelength (color) of light when responsive to communication from the master controller and a second, different wavelength (color) of light when unresponsive to communication from the master controller. The carrying device may become unresponsive when communication with the master controller is lost, thus activating the warning signal. Other types of warning signals from the light source are possible, such as predetermined patterns of emission, such as blinking. In response to determining that the image includes an inconsistent carrying device, method 140 may include outputting at least one of annotation data or an alert. The annotation data indicates the predicted carrying device in the image with the activated warning signal. For example, the annotation data may comprise a bounding box surrounding a predicted transport device with an activated warning signal in the image. Similarly, the output alert may signal that the image contains a transport device with an activated warning signal. Examples of output alerts include, for example, a text or other visual message to be displayed on a screen for viewing by an operator.

[0097] Localisation of transport devices The method 140 for detecting a transport device 30 in a workspace may include generating further annotation data corresponding to a plurality of virtual transport devices in a plurality of respective grid spaces 17 in the image 82. The location of the detected transport device 30 on the grid 15 may then be determined by comparing the annotation data indicating the detected transport device in the image with the further annotation data corresponding to the plurality of virtual transport devices. For example, the comparison may include calculating an intersection over union (IoU) value using the annotation data. The grid space corresponding to the further annotation data associated with the highest IoU value may then be selected as the grid location of the detected transport device.

[0098] In an example, the further annotation data comprises multiple bounding boxes corresponding to multiple virtual transport devices. Thus, calculating the IoU value may involve dividing the area of ​​the overlap, or "intersection," between two bounding boxes by the area of ​​the union of the two bounding boxes (e.g., the total area covered by the two boxes). For example, the area overlap between the bounding box of the detected transport device and a given bounding box corresponding to a given virtual transport device is calculated and divided by the area of ​​the union for the same two bounding boxes. This calculation is repeated for the bounding box of the detected transport device and each bounding box corresponding to each virtual transport device to provide a set of IoU values. The highest IoU value in the set of IoU values ​​may then be selected, and the grid location of the corresponding bounding box is inferred as the grid location of the detected transport device.

[0099] A detection system may be configured to implement any of the detection methods described herein. For example, the detection system includes an image sensor for capturing an image of at least a portion of the workspace and an interface for acquiring image data. The detection system includes a trained object detection model, implemented, for example, on a graphics processing unit (GPU) or a dedicated neural processing unit (NPU), for performing the processing and decision steps of the computer-implemented method 140 for detecting the transport device 30.

[0100] Detecting an identification marker on a delivery device 15 illustrates a computer-implemented method 150 for detecting identification markers on a transport device in a workspace comprising a grid 15. The method involves acquiring 151 image data representing an image portion including the transport device. The image portion may be a portion, e.g., at least a portion, of an image 82, e.g., shown in FIG. 8B, captured by a camera 71 positioned above the grid, e.g., as shown in FIG. 7.

[0101] 12 shows exemplary image portions 121a, 121b including respective transport devices 30a, 30b. The image portions 121a, 121b may be extracted from the image 82 based on annotation data, e.g., bounding boxes 120a, 120b, corresponding to the detected transport devices 30a, 30b in the image 82. For example, output annotation data of the method 140 for detecting transport devices in a workspace is used to obtain, e.g., extract, the image portions 121a, 121b from the image 82. If the annotation data represents one or more bounding boxes, for example, one or more image portions 121a, 121b corresponding to image data contained in the one or more bounding boxes 120a, 120b overlaid on the annotated image 83 are extracted from the image 82. For example, method 150 involves obtaining annotated image data 83 including annotation data 120a, 120b indicating one or more transport devices in the image, and cropping the annotated image data 83 to create one or more image portions 121a, 121b including the respective one or more transport devices 30a, 30b.

[0102] In the alternative, the image portion comprises the entire image 82 captured by the camera 71. The image portion may include one or more transport devices 30. In other words, the image portion comprises, for example, at least a portion of the image 82 captured by the camera 71.

[0103] Method 150 further involves processing 152 the acquired image data using a first neural network and a second neural network in series. The first neural network is trained to detect instances of identification markers on the transport device in the image. The second neural network is trained to recognize marker information associated with the identification marker in the image. An identification (“ID”) marker is, for example, a text label or other code (such as a barcode, QR code, or the like) on the transport device. The ID marker includes marker information, e.g., text or a QR code, associated with that marker. The marker information corresponds to ID information for the transport device, e.g., corresponds to, for example, a name or other descriptor of the transport device in a broader system. The marker information is encoded in the ID marker, e.g., as text or other code, and the corresponding ID information can be used to distinguish a given transport device from other transport devices operating in the system.

[0104] In an example, a first neural network is configured, e.g., trained or learned, to receive image portions as first input data and produce feature vectors as intermediate data, e.g., for passing as input to a second neural network. For example, the first neural network comprises a CNN90 configured to use convolutions to extract visual features, e.g., of different sizes, and produce feature vectors. An Efficient and Accurate Scene Text (EAST) detector may be used as the first neural network to identify instances of identifying markers, e.g., text labels, on a transport device.

[0105] In some cases, the first neural network outputs further annotation data, e.g., defining a bounding box, corresponding to the detected identification marker in the image portion. For example, process 152 involves determining whether the image portion includes an identification marker on a transport device based on processing with the first neural network. If the determination is positive, further annotation data corresponding to the location of the identification marker in the image portion is generated and output as part of method 150. The further annotation data may comprise image coordinates for the image or image portion. For example, the image coordinates correspond to at least two corners of a bounding box for the identification marker in the image portion. The bounding box may be defined, for example, by the coordinates of two opposite corners.

[0106] In an example, processing 152 the image data involves extracting a subportion of the image portion, the subportion corresponding to a detected identification marker on the transport device. For example, the image portion is cropped to generate a subportion including the identification marker. FIG. 12 shows an example subportion 122 extracted from image portion 121a and corresponding to a detected identification marker on transport device 30a. Subportion 122 may be rotated such that a longitudinal axis of the identification marker is substantially horizontal with respect to subportion 122, as shown in the example of FIG. 12. Method 150 may then include processing subportion 122 with a second neural network configured, e.g., trained or learned, to recognize marker information in the image.

[0107] Method 150 concludes with outputting 153 marker data representing the marker information determined by the second neural network. For example, the second neural network is configured to derive marker data from an image sub-portion including an ID marker. In an example where the ID marker comprises a text label, the second neural network may be configured to transcribe the image sub-portion including the label into marker data comprising label sequence data, e.g., a sequence (or "string") of letters, numbers, punctuation, or other symbols. For the example sub-portion 122 shown in FIG. 12, the second neural network would output the marker data as, e.g., label sequence data "AA-Z82" for the identification label of carrier device 30a. In an alternative example, the ID marker is a code on the carrier device, e.g., a QR ("quick response") code or barcode applied to the carrier device, e.g., on a label. The second neural network is configured, e.g., trained or learned, to determine the code from an image of the ID marker on the carrier device. The code, e.g., marker data, may then be output. For example, the code may be further processed to decode the ID information encoded therein. In other words, the detected QR code or barcode is decoded to determine, for example, the ID information of the carrying device, for example, the name.

[0108] In an example, the second neural network comprises a convolutional recurrent neural network (CRNN) configured to apply convolutions to extract visual features from image subportions and arrange the features in a sequence. The CRNN comprises two neural networks, e.g., a CNN and an additional neural network. In some cases, the second neural network includes a bidirectional recurrent neural network (RNN), e.g., a bidirectional long short-term memory (LSTM) model. For example, the bidirectional RNN is configured to process the feature sequence output of the CNN to predict the ID sequence encoded in the marker by applying sequential cues learned from patterns in the feature sequence, e.g., in the example of FIG. 12, that the ID sequence is highly likely to start with the letter "A" and end with a digit. Thus, the second neural network may comprise a pipeline of two or more neural networks, e.g., a CNN piped into a deep bidirectional LSTM, such that the feature sequence output of the CNN is passed to the biLSTM that receives it as input. In other examples, the second neural network comprises a different type of deep learning architecture, for example, a deep neural network.

[0109] A detection system may be configured to perform any of the detection methods described herein. For example, the detection system may include an image sensor for capturing an image of at least a portion of the workspace and an interface for acquiring image data. The detection system may include a trained object detection model, implemented, for example, on a graphics processing unit (GPU) or a dedicated neural processing unit (NPU), for performing the processing and decision steps of the computer-implemented method 150 for detecting identification markers on the transport device 30.

[0110] Determining the exclusion zone A method and system for determining an exclusion zone in a workspace is provided herein. The exclusion zone can be implemented by a master controller of transport devices and functions to prohibit transport devices operating in the workspace from entering the exclusion zone. For example, the exclusion zone can be determined around a defective transport device that has fallen and / or lost communication with the master controller, so that the defective transport device can later be attended to and, for example, removed from the workspace. This allows the workspace to remain usable while reducing the risk of other transport devices colliding with the defective transport device. In some cases, the determined exclusion zone can be proposed, for example, to an operator, prior to implementation, which can help ensure that the determined exclusion zone will cover the actual location of the defective transport device in the workspace.

[0111] FIG. 16 illustrates a computer-implemented method 160 for assisting in controlling the movement of one or more transport devices 30 operating in a workspace, for example, a workspace comprising the grid 15 described with reference to FIG.

[0112] Method 160 begins with obtaining 161 an image representation of the workspace captured by one or more image sensors. For example, the one or more image sensors are part of one or more cameras 71 having a view of the workspace. The cameras 71 may be arranged on a grid 15 of the workspace as shown in FIG. 7. The image of the workspace is received, for example, at an interface, e.g., a camera interface or CSI, communicatively coupled to the one or more image sensors.

[0113] A target image portion of an image representation of the workspace is acquired 162 at an interface, e.g., a different interface than that used to receive the image. The target image portion is mapped 163 to a target location in the workspace. Based on the mapping 163, an exclusion zone in the workspace is determined 164, into which one or more transport devices will be prohibited from entering. The exclusion zone includes the target location mapped from the target image portion. Exclusion zone data representing the exclusion zone is output 165 to a control system, e.g., a master controller, for implementing the exclusion zone in the workspace.

[0114] For example, a user viewing the graphical representation of the workspace selects a target image portion via an interface configured to acquire the target image portion. The interface may be, for example, a user interface with which the user interacts. The user interface may include a display screen for displaying the graphical representation of the workspace captured by the image sensor. The user interface may also include input means, for example, a touchscreen display, a keyboard, a mouse, or other suitable means, with which the user can select the target image portion.

[0115] In an example, the target image portion includes at least a portion of a defective transport device in the workspace. For example, the target image portion is a subset of one or more pixels selected from an image of the workspace captured by an image sensor. The one or more pixels correspond to at least a portion of the defective transport device shown in the image of the workspace. For example, the target image portion includes the entire defective transport device shown in the image. In another example, the target image portion is only a single pixel corresponding to a portion of the defective transport device shown in the image.

[0116] In other examples, for example, where the workspace comprises a grid 15 of cells 17, the target image portion corresponds to a given cell in the grid of cells. For example, the target image portion is a subset of one or more pixels that correspond to at least a portion of a given cell. In some cases, the target image portion includes the entire cell, and in other cases, the target image portion is only a single pixel that corresponds to a portion of the cell.

[0117] As described above, in some examples, a user selects a target image portion via an interface, e.g., a user interface. However, in other examples, the target image portion is obtained from an object detection system configured to detect defective transport devices from images of a workspace. For example, method 160 involves the object detection system obtaining images of the workspace captured by an image sensor and using an object classification model to determine that a defective transport device is present in the image data.

[0118] The object classification model, e.g., object classifier, generally comprises a neural network in the example described with reference to Figure 9, which is to be taken to apply accordingly. For example, the object classifier is trained using a training set of images of defective transport devices in the workspace to classify images subsequently captured by the image sensor as either including or not including defective transport devices in the workspace.

[0119] In the case of a positive classification by the trained object classifier, the object detection system can then output the target image portion. For example, the object detection system may indicate the target image portion in the original image from the image sensor using annotation data, such as, for example, a bounding box. Alternatively, the object detection system outputs the target image portion as a cropped version of the original input image received from the image sensor, the cropped version including the identified defective transport device in the workspace.

[0120] In an example, the object detection system includes a neural network trained to detect defective transport devices and their locations in image data. For example, the object detection system determines regions of an input image in which defective transport devices are present. The regions may then be output, for example, as target image portions. In such cases, training the neural network involves using annotated images of the workspace showing defective transport devices in the workspace. Thus, the neural network is trained to both classify objects in the workspace as defective transport devices and to detect where the defective transport devices are in the image, i.e., to identify the location of the defective transport devices relative to the image of the workspace.

[0121] As described herein, the target image portion output by the object detection system may include at least a portion of the defective transport device in the workspace. For example, the target image portion may be a subset of one or more pixels selected by the object detection system from an image captured by an image sensor based on, for example, a positive location identification of the defective transport device.

[0122] In an example, the determined exclusion zone includes a discrete number of grid spaces. For example, it may be determined that a defective transport device is located within a single grid space 17 on the grid 15. Therefore, the exclusion zone may be determined to extend into that grid space, for example, so that other transport devices are prohibited from entering that single grid space. Thus, collisions between other transport devices and the defective transport device may be prevented. Alternatively, the exclusion zone may be set as an area of ​​grid cells, for example, a 3×3 cell area, centered on the grid cell in which the defective transport device is located. Thus, the exclusion zone includes a buffer area around the affected grid cell in which the defective transport device is located. In some cases, the defective transport device spans more than one grid cell, for example, if it is located between grid cells, has fallen, or is misaligned with the track 22. In such cases, a buffer area around the mapped grid cell (containing the target location) can improve the effectiveness of the exclusion zone relative to excluding only the mapped grid cell. The size of the buffer area may be predetermined, for example, as a set area of ​​grid cells that will be applied once mapped grid cells for exclusion are determined. Additionally or alternatively, the size of the buffer area is a selectable parameter when implementing the exclusion zone in a control system.

[0123] A control system, e.g., a master controller, remotely controlling the movement of transport devices operating in the workspace can implement the exclusion zones based on the exclusion zone data output 165 as part of method 160. For example, each of one or more transport devices 30 can be remotely operable under the control of a control system, e.g., a central computer. To control the movement of one or more transport devices 30 on grid 15, instructions can be sent from the control system to one or more transport devices 30 via a wireless communications network, e.g., implementing one or more base stations.

[0124] A controller in each transport device 30 is configured to control various drive mechanisms of the transport device, e.g., vehicle 32, to control its movement. For example, the instructions include various movements in the XY plane of the grid structure 15, which may be encapsulated in a defined trajectory for a given transport device. Thus, the exclusion zones may be implemented by a central control system, e.g., a master controller, such that the defined trajectory avoids the exclusion zone represented by the exclusion zone data. For example, when an exclusion zone is implemented, one or more respective trajectories corresponding to one or more transport devices 30 on the grid are updated to avoid the exclusion zone.

[0125] In an example, mapping a target image portion (e.g., one or more pixels in an image) to a target location (e.g., a point on a grid structure) involves inverting a distortion of the image of the workspace. For example, if an image sensor is used in combination with an ultra-wide-angle lens, the lens distorts the view of the workspace. Therefore, the distortion is reversed, e.g., as part of the mapping between image pixels and grid points. An inverse distortion model may be applied to the target image portion for this purpose. The description of the image-to-grid mapping algorithm in the previous example applies correspondingly here. For example, mapping the target image portion to a target grid location involves applying an image-to-grid mapping algorithm described herein.

[0126] In an embodiment, method 160 of assisting a control system to control transport device movement through a workspace is implemented by the assistance system. For example, the assistance system includes one or more image sensors for capturing an image representation of the workspace and an interface for obtaining a target image portion of the image. The assistance system is configured to perform the mapping 163, determining 164, and outputting 165 steps of method 160. For example, the assistance system outputs exclusion zone data that a control system, e.g., a master controller, receives as input and implements in the workspace. The exclusion zone data can be transferred directly between the assistance system and the control system or can be stored by the assistance system in storage accessible by the control system.

[0127] In embodiments employing an assistance system, an object detection system configured to detect defective transport devices in the workspace is, for example, part of the assistance system, and an interface of the assistance system may obtain target image portions from the object detection system, as described in the examples.

[0128] The support system may be incorporated into a storage system 1, such as the example shown in FIG. 7, that includes a workspace and a control system for controlling transport device movement within the workspace. As described in the example with reference to FIG. 7, the workspace includes a grid 15 formed by a first set 22a of parallel tracks extending in an X direction and a second set 22b of parallel tracks extending in a Y direction that are orthogonal to the first set in a substantially horizontal plane. The grid 15 includes a plurality of grid spaces 17, and one or more transport devices 30 are positioned to selectively move peripherally on the tracks 22 to handle containers 10 stacked below the tracks 22 within the footprint of a single grid space 17. Each transport device 30 may have a footprint that occupies only a single grid space 17, such that a given transport device occupying one grid space does not interfere with another transport device occupying or crossing an adjacent grid space.

[0129] As described in the examples, the exclusion zone can correspond to a discrete number of grid spaces 17. For example, the assistance system determines a target location on the grid 15 corresponding to a defective transport device 30 detected in an image captured by a camera 71 on the grid 15. The target location is converted to grid space coordinates for the entire grid, for example, based on the calibration method described in the previous example, and the exclusion zone is determined based on the grid space of the target location. For example, the exclusion zone includes at least the grid space of the target location but may include additional surrounding grid spaces, for example, as a buffer area, as described in the examples. A control system of the storage system, for example, a master controller, is configured to implement the exclusion zone in the workspace based on the exclusion zone data determined by the assistance system such that transport devices 30 operating in the workspace are prohibited from entering the exclusion zone.

[0130] The above examples should be understood as illustrative examples. Further examples are contemplated. For example, camera 71 disposed above grid 15 has been described in many examples as an ultra-wide-angle camera. However, camera 71 could be a wide-angle camera, which includes a wide-angle lens that has a relatively longer focal length than an ultra-wide-angle lens, but still introduces distortion compared to a normal lens that reproduces a field of view that appears "natural" to a human observer.

[0131] Similarly, the described examples include acquiring and processing “images” or “image data.” Such images, in some cases, may be video frames, e.g., selected from a video comprising a sequence of frames. The video may be captured by a camera positioned on a grid as described herein. Thus, acquiring and processing images should be interpreted as including acquiring and processing a video, e.g., a frame from a video stream. For example, the described neural networks may be trained to detect instances of objects (e.g., a transport device, an ID marker thereon, etc.) in a video stream comprising multiple images. Furthermore, in the described example involving detecting an ID marker on a transport device, the image data is processed using a first neural network and a second neural network in series. However, in alternative examples, the first neural network and the second neural network are merged in an end-to-end ID marker detection pipeline or architecture, e.g., as a single neural network. Also provided is a method for detecting identification markers on a transport device in a workspace comprising a grid formed by, for example, a first set of parallel tracks extending in an X direction and a second set of parallel tracks extending in a Y direction orthogonal to the first set in a substantially horizontal plane, the grid comprising a plurality of grid spaces, wherein one or more transport devices are arranged to selectively move in at least one of the X or Y directions on the tracks and to handle containers stacked below the tracks within the footprint of a single grid space.The method comprises acquiring image data representing an image portion including one or more of the transport devices, processing the image data with at least one neural network trained to detect instances of identification markers on the transport devices in the image and to recognize marker information associated with the identification markers in the image, and outputting marker data representing the marker information determined by a second neural network. In the context of a detection system according to this alternative, one or more processors are configured to implement at least one neural network trained to detect instances of identification markers on the transport devices in the image and to recognize marker information associated with the identification markers in the image. The detection system is configured to process the acquired image data with the at least one neural network to generate marker data (e.g., a text string or code) representing marker information (e.g., a descriptor of the transport device) present on (e.g., encoded in) the identification markers on the transport devices, and output the marker data.

[0132] Further, in the described example of transport device location identification, the location of the detected transport device 30 on the grid 15 can be determined by comparing annotation data indicating the detected transport device in the image with further annotation data corresponding to multiple virtual transport devices. In an alternative example, the location of the detected transport device 30 on the grid 15 can be determined in a two-step process. First, the multiple virtual transport devices are filtered, which includes, for example, calculating an Intersection of Union (IoU) value using the predicted / inferred data of the transport device 30 and the annotation data of all possible locations of the multiple virtual transport devices. For example, virtual transport devices with calculated IoU values ​​smaller than a predetermined threshold are filtered out. Second, all remaining virtual transport devices are sorted (e.g., in ascending order) by closest distance to the center of the camera's field of view, and the first (e.g., closest) virtual transport device is taken as the mapping. The grid location of the transport device 30 is then set to, for example, the grid location from which the annotation data of the mapped virtual transport device is created.

[0133] In examples employing storage to store data, the storage may be random access memory (RAM) such as DDR-SDRAM (double data rate synchronous dynamic random access memory). In other examples, storage 330 may include non-volatile memory such as read-only memory (ROM) or a solid-state drive (SSD) such as flash memory. Storage may in some cases include other storage media, for example, magnetic, optical, or tape media, compact discs (CDs), digital versatile discs (DVDs), or other data storage media. Storage may be removable or non-removable from the associated system.

[0134] In examples employing data processing, a processor may be employed as part of the system involved. The processor may be a general-purpose processor such as a central processing unit (CPU), microprocessor, graphics processing unit (GPU), digital signal processor (DSP), application specific integrated circuit (ASIC), field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any suitable combination thereof designed to perform the data processing functions described herein.

[0135] In examples involving neural networks, a dedicated processor may be employed as part of the system involved. The dedicated processor may be an NPU, a neural network accelerator (NNA), or other version of a hardware accelerator specialized for neural network functions. Additionally or alternatively, the neural network processing workload may be at least partially shared by one or more standard processors, e.g., a CPU or GPU.

[0136] While the term “item” has been used throughout the description, it is envisioned to include other terms, such as case, asset, unit, pallet, equipment, etc. The term “annotation data” has also been used throughout the description. However, it is envisioned that the term corresponds to predicted data or inferred data in alternative nomenclature. For example, an object detection model (e.g., comprising a neural network) may be trained using annotated images, e.g., images with annotations such as bounding boxes, that serve as ground truth for the model, e.g., predictions or inferences with a confidence of 100% or 1 when normalized. These annotations may be made by humans, for example, for purposes of training the model. Thus, object detection of the present disclosure may be interpreted as outputting predicted or inferred data (e.g., instead of “annotation data”) to indicate a prediction or inference of a transport device in an image. The predicted or inferred data may be expressed as annotations, e.g., bounding boxes and / or labels, applied to the image. The predicted or inferred data may include, for example, a confidence associated with a prediction or inference of a transport device in the image. Annotations may be applied to the image based on the generated prediction or inference data, for example. For example, the image may be updated to include a bounding box surrounding the predicted carrying device with a label indicating the confidence level of the prediction, for example, as a percentage value or a normalized value between 0 and 1.

[0137] It should also be understood that features described with respect to any one example may be used alone or in combination with other features described, and may also be used in combination with one or more features of any other of the examples, or in any combination of any other of the examples. Moreover, equivalents and modifications not described above may also be employed without departing from the scope of the appended claims. The inventions described in the claims of the present application as originally filed are set forth below. [C1] 1. A computer-implemented method for detecting a transport device in a workspace comprising a grid formed by a first set of parallel tracks extending in an X direction and a second set of parallel tracks extending in a Y direction orthogonal to the first set in a substantially horizontal plane, the grid comprising a plurality of grid spaces; wherein one or more transport devices are arranged to selectively move on said track in at least one of said X direction or said Y direction and to handle containers stacked below said track within a footprint of a single grid space; The method comprises: acquiring image data representing an image of at least a portion of the workspace; processing the image data with an object detection model trained to detect instances of transport devices on the grid; determining whether the image includes a transport device of the one or more transport devices based on the processing; in response to determining that the image includes the transport device, outputting annotation data indicative of the transport device in the image; A computer-implemented method comprising: [C2] The method of C1, wherein the method comprises generating the annotation data. [C3] The method of any one of C1 and C2, wherein the method comprises outputting an updated version of the image including the annotation data. [C4] The method of any of C1 to C3, wherein the annotation data comprises a bounding box. [C5] 5. The method of any of C1 to 4, wherein the object detection model comprises a convolutional neural network. [C6] The object detection model is further trained to detect instances of conveying devices that are misaligned with the grid, and the method comprises: processing the image data with the object detection model and determining, based on the processing, whether the image includes a conveying device that is misaligned with the grid; 6. The method of any one of C1 to C5, comprising: [C7] The method includes, in response to determining that the image includes the misaligned transport device, annotation data indicating (a prediction of) the mismatched transport device in the image; or an alert that the image contains the mismatched delivery device; 7. The method of claim 6, comprising outputting at least one of: [C8] The object detection model is further trained to detect instances of a conveying device having an activated warning signal, and the method further comprises: processing the image data with the object detection model and determining, based on the processing, whether the image includes a delivery device with the warning signal activated; 8. The method of any one of C1 to 7, comprising: [C9] The method includes, in response to determining that the image includes the delivery device with the warning signal activated, annotation data indicating a prediction of the delivery device having the warning signal activated in the image; or an alert that the image includes the delivery device with the warning signal activated; 9. The method of claim 8, further comprising: outputting at least one of: [C10] The method comprises: generating further annotation data corresponding to a plurality of virtual transport devices in a plurality of respective grid spaces in the image; determining a location of the transport device based on a comparison of the annotation data indicative of the (prediction of the) transport device in the image with the further annotation data; 10. The method of any one of C1 to 9, comprising: [C11] The method of C10, wherein the further annotation data comprises a plurality of bounding boxes corresponding to the plurality of virtual transport devices. [C12] The method of claim 10 or 11, wherein the comparing includes calculating an intersection of sets, an IoU value, using the annotation data, and wherein the determining comprises selecting the grid space corresponding to the further annotation data associated with the highest IoU value. [C13] A data processing apparatus comprising means for performing the method according to any one of C1 to 12. [C14] A computer program comprising instructions which, when executed by a computer, cause the computer to perform any of the methods set out in C1 to 12. [C15] A computer-readable data carrier storing a computer program according to C14. [C16] 1. A detection system for detecting a transport device in a workspace, the workspace comprising: a grid formed by a first set of parallel tracks extending in an X direction and a second set of parallel tracks extending in a Y direction orthogonal to the first set in a substantially horizontal plane, the grid comprising a plurality of grid spaces; wherein one or more transport devices are arranged to selectively move on said track in at least one of said X direction or said Y direction and to handle containers stacked below said track within a footprint of a single grid space; the detection system comprising: an image sensor for capturing an image of at least a portion of the workspace; an interface for obtaining a target image portion of the image representation of the workspace; an object detection model trained to detect instances of transport devices on the grid; Equipped with wherein the detection system comprises: obtaining image data representative of the image; processing the image data through the object detection model; determining whether the image includes a transport device of the one or more transport devices based on the processing; in response to determining that the image includes the transport device, outputting annotation data indicative of the transport device in the image; A detection system configured to: [C17] The detection system of C16, wherein the detection system includes a wide-angle camera equipped with the image sensor. [C18] The object detection model is further trained to detect instances of conveying devices that are misaligned with the grid, and the detection system: processing the image data with the object detection model and determining, based on the processing, whether the image includes a conveying device that is misaligned with the grid; 18. The detection system according to C16 or 17, configured to perform the following: [C19] The object detection model is further trained to detect instances of a conveying device having an activated warning signal, and the detection system: processing the image data with the object detection model and determining, based on the processing, whether the image includes a delivery device with the warning signal activated; 19. The detection system of any one of claims 16 to 18, configured to: [C20] 20. The detection system of any one of C16 to 19, wherein the object detection model comprises a convolutional neural network.

Claims

1. 1. A computer-implemented method for detecting a transport device in a workspace comprising a grid formed by a first set of parallel tracks extending in an X direction and a second set of parallel tracks extending in a Y direction orthogonal to the first set in a substantially horizontal plane, the grid comprising a plurality of grid spaces; wherein one or more transport devices are arranged to selectively move on said track in at least one of said X direction or said Y direction and to handle containers stacked below said track within a footprint of a single grid space; The method comprises: acquiring image data representing an image of at least a portion of the workspace; processing the image data with an object detection model trained to detect instances of transport devices on the grid; determining whether the image includes a transport device of the one or more transport devices based on the processing; in response to determining that the image includes the transport device, outputting annotation data indicative of the transport device in the image; A computer-implemented method comprising:

2. The method of claim 1 , wherein the method comprises generating the annotation data.

3. The method of claim 1 or 2, wherein the method comprises outputting an updated version of the image including the annotation data.

4. The method of claim 1 or 2, wherein the annotation data comprises a bounding box.

5. The method of claim 1 or 2, wherein the object detection model comprises a convolutional neural network.

6. The object detection model is further trained to detect instances of conveying devices that are misaligned with the grid, and the method comprises: processing the image data with the object detection model and determining, based on the processing, whether the image includes a conveying device that is misaligned with the grid; The method of claim 1 or 2, comprising:

7. The method includes, in response to determining that the image includes the misaligned transport device, annotation data indicating a prediction of the mismatched delivery device in the image; or an alert that the image contains the mismatched delivery device. The method of claim 6 , comprising outputting at least one of:

8. The object detection model is further trained to detect instances of a conveying device having an activated warning signal, and the method further comprises: processing the image data with the object detection model and determining, based on the processing, whether the image includes a delivery device with the warning signal activated; The method of claim 1 or 2, comprising:

9. The method includes, in response to determining that the image includes the delivery device with the warning signal activated, annotation data indicating a prediction of the delivery device having the warning signal activated in the image; or an alert that the image includes the delivery device with the warning signal activated; 9. The method of claim 8, comprising outputting at least one of:

10. The method comprises: generating further annotation data corresponding to a plurality of virtual transport devices in a plurality of respective grid spaces in the image; determining a location of the transport device based on a comparison of the annotation data indicating a prediction of the transport device in the image with the further annotation data; The method of claim 1 or 2, comprising:

11. The method of claim 10 , wherein the further annotation data comprises a plurality of bounding boxes corresponding to the plurality of virtual transport devices.

12. 11. The method of claim 10, wherein the comparing comprises calculating an intersection of unions, IoU values, using the annotation data, and wherein the determining comprises selecting the grid space corresponding to the further annotation data associated with a highest IoU value.

13. 3. A data processing apparatus comprising means for carrying out the method according to claim 1 or 2.

14. A computer program comprising instructions which, when said program is executed by a computer, cause said computer to perform the method according to claim 1 or 2.

15. A computer readable data carrier storing a computer program according to claim 14.

16. 1. A detection system for detecting a transport device in a workspace, the workspace comprising: a grid formed by a first set of parallel tracks extending in an X direction and a second set of parallel tracks extending in a Y direction orthogonal to the first set in a substantially horizontal plane, the grid comprising a plurality of grid spaces; wherein one or more transport devices are arranged to selectively move on said track in at least one of said X direction or said Y direction and to handle containers stacked below said track within a footprint of a single grid space; the detection system comprising: an image sensor for capturing an image of at least a portion of the workspace; an interface for acquiring a target image portion of the image representation of the workspace; and an object detection model trained to detect instances of transport devices on the grid. Equipped with wherein the detection system comprises: obtaining image data representative of the image; processing the image data through the object detection model; determining whether the image includes a transport device of the one or more transport devices based on the processing; in response to determining that the image includes the transport device, outputting annotation data indicative of the transport device in the image; A detection system configured to:

17. The detection system of claim 16 , wherein the detection system includes a wide-angle camera that includes the image sensor.

18. The object detection model is further trained to detect instances of conveying devices that are misaligned with the grid, and the detection system: processing the image data with the object detection model and determining, based on the processing, whether the image includes a conveying device that is misaligned with the grid; 18. The detection system of claim 16 or 17, configured to:

19. The object detection model is further trained to detect instances of a conveying device having an activated warning signal, and the detection system: processing the image data with the object detection model and determining, based on the processing, whether the image includes a delivery device with the warning signal activated; 18. The detection system of claim 16 or 17, configured to:

20. 18. The detection system of claim 16 or 17, wherein the object detection model comprises a convolutional neural network.

Citation Information

Patent Citations

  • Physical distribution system

    JP2004355419A

  • Robot control system and robot control program

    JP2011022700A

  • Position measuring system of indoor self-propelled robot

    JP2018014064A

  • Robotic system with wall-based packing mechanism and methods of operating the same

    JP2021075395A

  • Storage system and method

    WO2020170037A2