Method and system for detecting load-handling devices
A method using unique self-similar patterns and combined object detection models effectively identifies and distinguishes load-handling devices in ASRS, addressing visibility and orientation challenges, enhancing accuracy and preventing collisions.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- OCADO INNOVATION LTD
- Filing Date
- 2025-10-27
- Publication Date
- 2026-05-07
AI Technical Summary
Existing load-handling devices in automated storage and retrieval systems (ASRS) are difficult to identify accurately due to issues with visibility, orientation, and congestion, which affects the reliability of imaging systems in detecting unique identifiers.
Implementing a computer-implemented method using unique self-similar patterns on load-handling devices, processed by a combination of high- and low-level object detection models, to identify and distinguish between devices regardless of orientation and position, and detect potential derailments or malfunctions.
Enhances the accuracy and efficiency of load-handling device identification, minimizing computational expense and preventing collisions by using self-similar patterns that are easily detectable and orientation-independent, even in congested environments.
Smart Images

Figure EP2025080924_07052026_PF_FP_ABST
Abstract
Description
[0001] Method and system for detecting load-handling devices
[0002] Technical Field
[0003] The present disclosure relates generally to the field of load-handling devices for use in warehouses or fulfilment centres.
[0004] Background
[0005] Some commercial and industrial activities require systems that enable the storage and retrieval of a large number of different products. For example, WO2015 / 185628A2 (Ocado) describes an automated storage and fulfilment system (ASRS) in which stacks of storage containers are arranged within a grid storage structure. The containers are accessed from above by load-handling devices operative on rails or tracks located on the top of the grid storage structure. The load-handling devices may be those described in W02015 / 019055A1 (Ocado). In such an ASRS, items within the containers may be accessed by robotic picking stations that use a robotic manipulator or arm, such as those described in WO2017081281A1 (Ocado).
[0006] Within the storage and fulfilment system, it is important that load-handling devices can be identified. It is against this background that the present invention has been devised.
[0007] Summary
[0008] In a first aspect, there is a computer-implemented method of identifying a load-handling device within a system, the system comprising: a first set of parallel rails or tracks and a second set of parallel rails or tracks extending substantially perpendicularly to the first set of rails or tracks in a substantially horizontal plane to form a grid comprising a plurality of grid spaces; one or more load-handling devices, wherein each load-handling device is configured to move along the first and / or second set of tracks, and wherein each loadhandling device comprises a respective unique self-similar pattern, the method comprising: obtaining an image of the system; processing the image with an object detection model trained to detect an instance of a load-handling device comprising a respective unique self-similar pattern; determining, based on the processing, whether the image includes a loadhandling device; and outputting, in response to determining that the image includes a load-handling device, annotation data indicative of the load-handling device in the image. This means a load-handling device can be identified regardless of the orientation and / or position of the load-handling device within the image.
[0009] Processing the image with an object detection model may further comprise processing the image with a first object detection model trained to detect the unique self-similar pattern at a low level of granularity to detect a load-handling device, and processing the image with a second object detection model trained to detect the unique self-similar pattern at a high level of granularity to detect a specific load-handling device. This means a specific load-handling device ca be detected whilst minimising computational expense.
[0010] The first object detection model may comprise a convolutional neural network, and the second object detection model may comprise a convolutional neural network and / or an image similarity model.
[0011] The object detection model may comprise a convolutional neural network.
[0012] The method may further comprise outputting an updated image including the annotation data. The annotation data may comprise a bounding box for a load-handling device and / or a plurality of bounding boxes, where each bounding box of the plurality of bounding boxes is for a respective load-handling device. The annotation data may comprise information identifying the load-handling device. This means a specific loadhandling device may be identified in the updated image itself.
[0013] The object detection model may be further trained to detect a derailment of a loadhandling device and / or a warning signal of the load-handling device, wherein the method may further comprise determining whether the image contains a load-handling device that has derailed and / or has a warning signal. The method may further comprises upon determining that a load-handling device has derailed and / or has a warning signal, setting up an exclusion zone around the load-handling device that has derailed and / or has a warning signal. This means the system can operate without risking further collisions and / or derailments of load-handling devices.
[0014] The method may further comprise updating the annotation data to indicate the loadhandling device that has derailed and / or has a warning signal. The method further comprises outputting an updated image including the annotation data indicative of the load-handling device that has derailed and / or has a warning signal. This means a specific load-handling device fault and / or malfunction may be identified in the updated image itself.
[0015] The system may comprise an image sensor, wherein obtaining an image of the system comprises obtaining the image from the image sensor. The image sensor may be located above the grid. The system may comprise a picking station operational on the grid, the picking station comprising a robotic manipulator comprising the image senor, wherein the robotic manipulator is configured to transfer items between containers received in respective grid cells adjacent the picking station. The method may further comprise orienting the robotic manipulator of the picking station to direct the image sensor of the robotic manipulator towards the one or more load-handling devices. This means different orientations of the load-handling devices may be included in the image.
[0016] Each unique self-similar pattern may comprises a combination of one or more colours and / or one or more geometric shapes. Each unique self-similar pattern may comprise a fractal. Each unique self-similar pattern may be generated using a mathematical function, such as a recursive function. This means each load-handling device can be identified regardless of its orientation / position in the image.
[0017] Each unique self-similar pattern may be mapped to a respective load-handling device.
[0018] The load-handling device may comprise: a body or skeleton mounted on a first set of wheels being arranged to engage with the first set of parallel tracks and a second set of wheels being arranged to engage with the second set of parallel tracks; and / or a drive assembly configured to drive the first or second sets of wheels to move the load-handling device along the first or second set of parallel rails respectively; and / or a direction-change assembly configured to raise or lower the first set of wheels and / or lower or raise the second set of wheels with respect to the body or skeleton to engage and disengage the wheels with the parallel tracks; and / or a container-lifting assembly configured to raise or lower a gripping device in the vertical direction. Substantially all or at least a portion of an outer surface of the loadhandling device, and / or the body or skeleton, and / or the first and / or second sets of wheels, and / or the drive assembly, and / or the direction-change assembly, and / or the container-lifting assembly may comprise surface decoration defining the self-similar pattern. This means the self-similar pattern is integral to the design and manufacture of the load-device.
[0019] In a second aspect, there is a computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the computer- implemented method of the first aspect.
[0020] In a third aspect, there is a computer readable medium comprising the computer program of the second aspect.
[0021] In a fourth aspect, there is a data processing system comprising means for carrying out the computer implemented method of the first aspect.
[0022] In a fifth aspect, there is a system comprising: a first set of parallel rails or tracks and a second set of parallel rails or tracks extending substantially perpendicularly to the first set of rails or tracks in a substantially horizontal plane to form a grid comprising a plurality of grid spaces; one or more load-handling device, wherein each load-handling device is configured to move along the first and / or second set of tracks, and wherein each loadhandling device comprises a respective unique self-similar pattern; an image sensor; and a processor configured to carry out the method of the first aspect. Brief Description of Drawings
[0023] The invention is described with reference to the accompanying drawings, wherein:
[0024] Figure 1 shows an automated storage and retrieval system that uses load-handling devices;
[0025] Figure 2 shows a single load-handling device with container-lifting means in a lowered configuration;
[0026] Figure 3 shows a robotic picking station;
[0027] Figure 4 shows a method of detecting a load-handling device;
[0028] Figure 5 shows how the method of Figure 4 may be implemented;
[0029] Figure 6 shows a first set of patterns that may be used on the load-handling device for detection by the method of Figure 4;
[0030] Figure 7 shows a second set of patterns that may be used on the load-handling device for detection by the method of Figure 4;
[0031] Figure 8 shows image sensors that can be used in the method of Figure 4;
[0032] Figure 9 shows a load-handling device that can be identified by the method of Figure 4; and
[0033] Figure 10 shows an example processing system for implementing aspects described herein.
[0034] Detailed Description
[0035] WO2015 / 185628A (Ocado), hereby incorporated by reference, describes a known ASRS in which stacks of containers are arranged within a grid framework structure. The containers are accessed by one or more load-handling devices, otherwise known as “bots”, operative on tracks located on the top of the grid framework structure. A system of this type is illustrated schematically in Figure 1.
[0036] As shown in Figure 1 , stackable containers 10, also known as “bins” or “totes”, are stacked on top of one another to form stacks 12. The stacks 1 are arranged in a grid framework structure 14. The grid framework structure 14 is made up of a plurality of storage columns or grid columns. Each grid in the grid framework structure has at least one grid column to store a stack of containers. Each bin 10 typically holds a plurality of product items (not shown).
[0037] The grid framework structure 14 comprises a plurality of upright members 16 that support horizontal members 18, 20. A first set of parallel horizontal grid members 18 is arranged perpendicularly to a second set of parallel horizontal members 20 in a grid pattern to form a horizontal grid structure 15 supported by the upright members 16. The members 16, 18, 20 are typically manufactured from metal. The bins 10 are stacked between the members 16, 18, 20 of the grid framework structure 14, so that the grid framework structure 14 guards against horizontal movement of the stacks 12 of bins 10 and guides the vertical movement of the bins 10.
[0038] The top level of the grid framework structure 14 comprises a grid or grid structure 15, including rails 22 arranged in a grid pattern across the top of the stacks 12. The rails or tracks 22 guide a plurality of load-handling devices 30. A first set 22a of parallel rails 22 guide movement of the robotic load-handling devices 30 in a first direction (e.g. a Y- direction along track 22a) across the top of the grid framework structure 14. A second set 22b of parallel rails 22, arranged perpendicular to the first set 22a, guide movement of the load-handling devices 30 in a second direction (e.g. an X-direction along track 22b), perpendicular to the first direction. In this way, the rails 22 allow the robotic load-handling devices 30 to move laterally in two dimensions in the horizontal X-Y plane. A loadhandling device 30 can be moved into position above any of the stacks 12.
[0039] An example load-handling device 30 shown in Figure 2 is described in WO2015 / 019055 (Ocado), hereby incorporated by reference. The load-handling device 30 comprises a vehicle 32, which is arranged to travel on the rails 22 of the frame structure 14. A first set of wheels 34, consisting of a pair of wheels 34 on the front of the vehicle 32 and a pair of wheels 34 on the back of the vehicle 32, is arranged to engage with two adjacent rails of the first set 22a of rails 22. Similarly, a second set of wheels 36, consisting of a pair of wheels 36 on each side of the vehicle 32, is arranged to engage with two adjacent rails of the second set 22b of rails 22. Each set of wheels 34, 36 can be lifted and lowered, by way of a direction-change assembly, so that either the first set of wheels 34 or the second set of wheels 36 is engaged with the respective set of rails 22a, 22b at any one time. For example, when the first set of wheels 34 is engaged with the first set of rails 22a and the second set of wheels 36 is lifted clear from the rails 22, the first set of wheels 34 can be driven, by way of a drive assembly housed in the vehicle 32, to move the load-handling device 30 in the Y-direction. To achieve movement in the X-direction, the first set of wheels 34 is lifted clear of the rails 22, and the second set of wheels 36 is lowered into engagement with the second set 22b of rails 22. The drive assembly can then be used to drive the second set of wheels 36 to move the load-handling device 30 in the X direction.
[0040] The load-handling device 30 is equipped with a container-lifting device or assembly, e.g. a crane mechanism, to lift a storage container from above. The lifting device comprises a winch tether or cable 38 wound on a spool or reel and a gripper device 39. The lifting device shown in Figure 2 comprises a set of four lifting tethers 38 extending in a vertical direction. The tethers 38 are connected at or near the respective four corners of the gripper device 39, e.g. a lifting frame, for releasable connection to a storage container 10. The gripper device 39 is configured to releasably grip the top of a storage container 10 to lift it from a stack of containers in a storage system of the type shown in Figure 1 .
[0041] To remove a bin 10 from the top of a stack 12, the load-handling device 30 is first moved in the X- and Y-directions to position the gripper device 39 above the stack 12. The gripper device 39 is then lowered vertically in the Z-direction to engage with the bin 10 on the top of the stack 12. The gripper device 39 grips the bin 10, and is then pulled upwards by the cables 38, with the bin 10 attached. At the top of its vertical travel, the bin 10 is held above the rails 22 accommodated within the vehicle body (or skeleton) 32. In this way, the load-handling device 30 can be moved to a different position in the X-Y plane, carrying the bin 10 along with it, to transport the bin 10 to another location. On reaching the target location (e.g. another stack 12, an access point in the storage system, or a conveyor belt) the bin or container 10 can be lowered from the container receiving portion and released from the grabber device 39. Load-handling device 30 may also have a warning signal, such as a light, which indicates whether the load-handling device is functioning as intended. In one example, the warning signal may indicate whether the drive assembly, and / or the direction-change assembly, and / or the container-lifting assembly, and / or whether a communications system to a master controller of the ASRS, is functioning correctly. In the example of a light, green may indicate functioning correctly, whereas red may indicate a malfunction. A specific light colour pattern may indicate a specific malfunction.
[0042] With reference to Figure 3, the ASRS may further comprise a robotic picking station 50 operational on top of the storage and retrieval structure 1 , e.g. alongside the loadhandling devices 30 (not shown). In this example, the robotic picking station is mounted on the grid but could be alternatively suspended on a rail system for example for operation on top of the storage and retrieval structure. The robotic picking station 50 comprises a robotic manipulator 52 comprising a robotic arm 54 and an end effector 56 for releasably engaging a product to be manipulated, together with several designated grid cells 60, 62. The end effector 56 may be a suction device 64 connected to a vacuum source by a vacuum line within trunking 66. The robotic manipulator 52 is mounted on a plinth 58 above a single grid cell 60 and, depending on its location on the structure 1 , can be surrounded by up to eight other grid cells 62 as shown in Figure 6. In general, the robotic manipulator 52 is configured to pick an item or product from any one of the containers located in one of the designated grid cells 62 and place it in a container located in another of the designated grid cells 62. The load-handling devices collect containers from, and deliver them to, the designated grid cells 62 as necessary. In this way, the robotic picking station 50 and the load-handling devices 30 work in conjunction via a master controller to fulfil a customer order or redistribute products throughout the storage and retrieval system 1 .
[0043] An ASRS, when operating at peak throughput, can have several load-handling devices in operation. It is important to be able to automatically identify each load-handling device individually via an imaging system. One option could be to attach an identifying marker, such as a label with a number for example, to each load-handling device. Detection of the identifying marker allows the load-handling device to be identified.
[0044] However, one issue is that the imaging system may not be able to obtain full visibility of the ASRS. For example, if the imaging system uses a camera located above the ASRS, the camera may not have sufficient resolution to allow detection of the identifying marker on a load-handling device more distant from the camera. Another issue is that an area of the ASRS may be congested with load-handling devices, so it can be difficult for the imaging system to obtain a line of sight to each identifying marker on each respective load-handling device. Further, it may not be possible to detect which load-handling device a given marker is located due to the proximity of other load-handling devices. Put another way, the reliability of detection depends on the angle and orientation of the load-handling devices with respect to an image sensor.
[0045] It would be helpful to overcome the above issues when identifying a given load-handling device using a vision system.
[0046] With reference to Figure 4, a method according to an aspect is described. The method is used to detect an instance of a load-handling device operational on a system comprising a first set of parallel rails or tracks and a second set of parallel rails or tracks extending substantially perpendicularly to the first set of rails or tracks in a substantially horizontal plane to form a grid comprising a plurality of grid spaces. There may be one or more load-handling devices, wherein each load-handling device is configured to move along the first and / or second set of tracks. The system may be that described above in reference to Figures 1-2.
[0047] Each load-handling device also comprises a respective unique self-similar pattern. The self-similar pattern may be defined using surface decoration of the load handling device, such as on the outer surface of vehicle 32 in Figure 2 for example. A self-similar pattern may be thought of one that repeats itself at any level of zoom such that parts of the pattern are similar or identical to the pattern as whole. For example, the overall shape of the pattern may be the same shape, albeit at a larger scale, as individual parts of the pattern. In one example, self-similar may be thought of as a mathematical concept where each unique self-similar pattern is defined by respective coefficients of a mathematical function. Each self-similar pattern may have a unique combination of colours and / or geometric shapes. The self-similar pattern may be generated using a mathematical function, such as a recursive function. Figures 6 and 7, further described below, show example self-similar patterns. One example of a self-similar pattern is a fractal. In step 410, an image of the system is obtained from the image sensor. In step 420, the image data is processed with an object detection model trained to detect an instance of a load-handling device comprising a respective unique self-similar pattern. An instance allows a load-handling device to be detected generally and / or a specific load-handling device to be detected. An example object detection model that may be used is described in further detail below.
[0048] In step 430, as a result of the processing in step 420, whether the image includes a loadhandling device, is determined. Put another way, the object detection model can be used to detect whether a load-handling device is in the image.
[0049] In step 440, in response a determination in step 430 that the image includes a loadhandling device, annotation data indicative of the load-handling device, is output. The annotation data may be general and indicate in a binary manner whether the image contains a load-handling device. The annotation data may alternatively or additionally identify the specific load-handling device in the image. In one example may involve checking a mapping between unique self-similar patterns and respective load-handling devices in a look-up table or a database. The annotation data may alternatively or additionally indicate whether a warning signal is engaged on the load-handling device.
[0050] The method of Figure 4 may further comprise outputting an updated version of the image data including the annotation data. A bounding box may be used in the updated image to generally indicate whether a load-handling device is present. Additionally or alternatively, further bounding boxes may be used in the updated image, where each bounding box corresponds to a respective load-handling device. Additionally or alternatively, one or more bounding boxes can be used to indicate whether respective warning signals are engaged. For the warning signals, a bounding box may be colour coded and one example is green to indicate correct operation and red to indicate incorrect operation.
[0051] The method of Figure 4 has operational advantages compared to known techniques. The use of a unique self-similar pattern lends itself well to detection by an object-detection model irrespective of angle and / or orientation of one or more load-handling devices in the image. This is because each unique self-similar pattern may be mathematically delineated which assists inference in the object detection model. It also assists in the object-detection model being able to delineate between load-handling devices, and in doing so identifying each load-handling device. This overcomes the issues with known solutions where an identity marker, such as a label, may not be correctly detected and or obscured entirely. Accordingly the resolution of an imaging sensor used for the image is not as critical to the process.
[0052] Additionally and / or alternatively, the object detection model may be further trained to detect whether a load-handling device has derailed and / or fallen over and / or crashed. Upon the object detection model having detected a derailed / fallen / crashed one or more load-handling devices in the image, it may be necessary to set up an exclusion zone thereabout. The exclusion zone is set up by the master controller to ensure that loadhandling devices are prevented from entering the exclusion zone. This can prevent further derailments / crashes / falling-overs of load-handling device that would otherwise enter the now excluded zone.
[0053] Although the object detection model can be trained to detect a unique self-similar pattern of a plurality of unique self-similar patterns, two object detection models may be used instead. As shown by Figure 5, an image sensor 510 generates image data that is first processed by a high-level object detection model 520. The high-level object detection model can be trained at a low-level of granularity to detect a self-similar pattern in general. For example, the model may be trained on colours and / or geometric shapes common to a plurality of unique self-similar patterns. The high-level object detection model can be thought of being trained to detect a load-handling device within the system. In theory, the level of granulation used can be very low since load-handling devices with self-similar patterns are somewhat prominent to their surrounding environment within the system.
[0054] After a load-handling device has been detected by high-level object detection model 520, low-level object detection model 530 is then used. The low-level object detection model can be trained at a high-level of granularity to detect a unique self-similar pattern. For example, the low-level object detection model can be trained to distinguish between self-similar patterns and thus distinguish between the load-handling devices. The low-level object detection model can be thought of as being trained to detect a specific load-handling device. Put another way, the low-level object detection model can be thought of as a fine-grained image classifier. Annotation generator 550 is then used to annotate and / or update the image as described above for step 440 of Figure 4. Thus, depending on the implementation of the system which may use hundreds of loadhandling devices, it may be more computationally efficient to use the process shown in Figure 5. The models 520 and 530 of Figure 5 may be thought of as a coarse filter to identify a load-handling device in general before specific identification occurs. If the number of load-handling devices is small and / or the unique self-similar patterns are easily distinguished in terms of computational expense, a single object-detection model may be more suitable. In any case, models 520 and 530 may also be thought of as sub-models within a single object-detection model 525.
[0055] Figure 6 shows a first set of patterns 600 that may be used on the load-handling device for detection by the method of Figure 4. In Figure 6, three load-handling devices 630, 640, 650 are shown. Each load handling device has wheels 620a and 620b which engage with transverse tracks or rails 610a and 620a respectively. Each load-handling device has a respective unique pattern. In this example, load-handling device 630 has a unique self-similar pattern defined by elements 631 , 632, and 633, load-handling device 640 has a unique self-similar pattern defined by elements 641 , 642, and 643, and loadhandling device 650 has a unique self-similar pattern defined by elements 651 , 652, and 653.
[0056] In the example of Figure 6, the self-similar patterns are based on the formula:
[0057] Zn+1 = z + c
[0058] (1) where c is fixed complex number, and z represents a point on the complex plane.
[0059] In this case, the complex plane is defined by the surface on which the pattern is located. A so-called Julia set can be obtained from formula (1). Although formula (1) results in a mathematically infinite repeating self-similar pattern, it will be appreciated that infinite repetition will not be observable on a load-handling device, and instead a practical number of repetitions will be used instead, which will depend on the resolution of the surface decoration. Changing the value of c results in different patterns as shown by comparing the patterns on load-handling devices 630, 640, 650. Further, colours may be used to indicate a density of points derived from formula (1). In this example, each side of each load-handling device has the same respective self-similar pattern. Figure 7 shows a second set of patterns 700 that may be used on the load-handling device for detection by the method of Figure 4. In Figure 7, three load-handling devices 730, 740, 650 are shown. Each load handling device has wheels 720a and 720b which engage with transverse tracks or rails 710a and 720a respectively. Each load-handling device has a respective unique pattern. In this example, load-handling device 730 has a unique self-similar pattern defined by elements 731 , 732, and 733, load-handling device 740 has a unique self-similar pattern defined by elements 741 , 742, and 743, and loadhandling device 750 has a unique self-similar pattern defined by elements 751 , 752, and 753.
[0060] The self-similar patterns in Figure 7 are also based on formula 1 and illustrate how unique self-similar patterns can be derived by changing weightings or coefficients of formula 1 .
[0061] In general, it will be appreciated that Figures 6 and 7 are mere examples of the infinite number of self-similar patterns that can be generated. In general, a recursive mathematical formula can be used to generate the unique self-similar pattern. Known recursive mathematical formulas are the Mandelbrot set, Julia set, Burning Ship fractal, Nova fractal and Lyapunov fractal. In general, the number of patterns that may need to be generated can be small and in the order of 200-600 so the weighting or coefficients can be selected accordingly to give each load-handling device a significantly different unique self-similar pattern. Further, although substantially all of the outside surface (e.g. each of the top and 4 side / lateral faces / surfaces) of a load-handling devices is shown to have a self-similar pattern, it will be appreciated that only a portion may have the pattern such as a top surface only or a top surface and a top portion of at least one side surface. In general, the self-similar pattern can be located on portions of the load-handling device where it is more likely to be “seen” by the image sensor.
[0062] Figure 8 shows a system 800 that can be used with the method of Figure 4. An image sensor, such as a camera, 810 can be located above the tracks 840 of the ASRS. Image sensor 810 can provide the image data for the method of Figure 4 to detect and identify load-handling devices 850 and 860, each of which have a respective unique self-similar pattern. Additionally or alternatively, a picking station 820 of the type shown in Figure 3 and described above may have an image sensor 830 on an end-effector. Image sensor 830 can provide the image data for the method of Figure 4 to detect and identify loadhandling devices 850 and 860, each of which have a respective unique self-similar pattern. Image sensor 820 may be advantageously orientated towards a set of loadhandling devices. It will be appreciated that image data can be used from both image sensors 810 and 830 to verify the identity of a load-handling device. That is, if image data from two different image sensors independently result in the method of Figure 4 described above identifying the same load-handling device, it is more likely the correct load-handling device has been identified. It will also be appreciated that image data can be used from both image sensors 810 and 830 to verify whether a load-handling device has derailed and / or a warning signal is engaged.
[0063] Figure 9 shows a specific embodiment of a load-handling device 900 that can use the above methods. Load-handling device is similar to that further described in PCT / EP2022 / 051652, herby incorporated by reference. As described in PCT / EP2022 / 051652, load-handling device may be manufactured through additive manufacturing (also known as 3D printing). Therefore, the self-similar patterns may be printed on the outer surfaces when manufacturing one or more of the parts. A 3D printer may be coupled to a computer that generates the self-similar patterns using a mathematical formula, and maps the self-similar pattern to the load-handling device being printed.
[0064] Object detection model
[0065] The inventor has found that a suitable object detection model may be based on a pretrained model described in C. Anderson and R. Farrell, "Improving Fractal Pre-training," in 2022 IEEE / CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 2022, pp. 2412-2421 , doi: 10.1109 / WACV51458.2022.00247, which is hereby incorporated by reference. A ResNet50 CNN model is pre-trained for 90 epochs, with 1 ,000,000 training samples per epoch, and with an image resolution of 224 x 224. Each image is a fractal derived from a mathematical equation. The pretrained model was found to be useful for real-world image recognition tasks.
[0066] It was also found by the inventor that the pre-trained model can be fine-tuned using a dataset specific to the ASRS and load-handling devices described above. The dataset for fine-tuning was labelled accordingly for load-handling devices at different positions on the grid (e.g. relative to an image sensor) in different operational states including moving and stationary, and load-handling devices with respective unique self-similar patterns as described above. After the fine-tuning, the resulting object detection model was found to have a precision of about 0.9, a recall about 0.925, and an f1 score about 0.9125. Further fine-tuning of the model with a dataset labelled accordingly for load-handling devices that had derailed and / or with a warning engaged and not engaged resulted in similar classification performance metrics.
[0067] The resulting object-detection model may be used independently or act as a low-level object detection model in combination with a high level object detection model, as shown in Figure 5. The high-level object detection model may comprise a “You Only Look Once” (YOLO) object detection model, e.g. YOLOv8 or Scaled-YOLOv4, which has a CNN-based architecture. Other example high-level object detection models include neural-based approaches such as RetinatNet or R-CNN (Regions with CNN features) and non-neural approaches such as a support vector machine (SVM) to do the object classification based on determined features, e.g. Haar-like features or histogram of oriented gradients (HOG) features. The high-level object detection model may be fine-tuned with a dataset labelled accordingly indicating no load-handling devices or load-handling devices with self-similar patterns. Given the prominence of a self-similar pattern in an image, the dataset and fine-tuning required are both minimal.
[0068] The above example object detection models can be thought of as image classification models.
[0069] Independently or in addition to each or all of the image classification models described above, an image similarity model may be used such as that in A. Radford, J.W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning Transferable Visual Models From Natural Language Supervision” in arXiv:2103.00020, 2021 , which is hereby incorporated by reference. Image similarity models may operate in general by comparing embeddings of respective images to determine similarity. Therefore, load-handling devices with self-similar patterns can be compared to the self-similar patterns alone. Image similarity models have been found to deal with rotation and distortion of the image whilst allowing for accurate comparisons. It will also be appreciated that any two images may be compared by calculating a hamming distance between two hashes generated from respective images. Thus, the self-similar pattern can be detected on a load-handling device regardless of the orientation of the load-handling device. Therefore, an image similarity model can be used to verify the self-similar image detected (and thus the load-handling device identified) by the image classification model.
[0070] Figure 10 depicts a processing system 1000 for implementing aspects described. In some aspects, processing system 1000 implements logical elements from Figures 4 and 5. In this example, processing system 1000 includes one or more one or more processors 1002 configured to retrieve and execute instructions stored in one or more memories 1006, which may be volatile memory, such as a random access memory (RAM), or a non-volatile memory, such as non-volatile random access memory (NVRAM), or the like. In this example, the one or more memories 1006 include a training component 1050, an image receiving component 1051 , an object detection model component 1050, an annotation component 1053, and a self-similar pattern component 1054.
[0071] The training component 1050 may be configured to train the object detection models described above. The image receiving component 1051 may be configured to perform and / or control at least step 410 described with reference to Figure 4 and / or step 510 described with reference to Figure 5, and / or elements 810 or 820 described with reference to Figure 8. The object detection model component 1052 may be configured to perform at least step 420 and 430 described with reference to Figure 4 and / or at least step 520, 525, and 530 described with reference to Figure 5. The annotation component 1053 may be configured to perform at least step 440 described with reference to Figure
[0072] 4 and / or at least step 550 described with reference to Figure 5. The self-similar pattern component 1053 may be configured to perform at least step 440 described with reference to Figure 4 and / or at least step 540 and 550 described with reference to Figure
[0073] 5 and / or use a mathematical formula such as formula (1) above.
[0074] The one or more memories 1006 may include various additional components or data useful for performing described methods in accordance with presently described aspects. Instructions 1030 may generally implement any of components 1050-1054 for processing by the one or more processors 1002. Processing system 1000 may further include a graphics processing unit (GPU) 1008 that is operatively connected to the one or more processors 1002 and to the one or more memories 1006 to offload relevant data from the one or more processors 1002 and process data in parallel with the one or more processors 1002. Processing system 1000 may further include a video display 1016 connected by a video interface 1010, and various input / output devices such as a keyboard 1018, mouse 1020, and disk drive or solid state drive 1022 connected by an I / O interface 1012. In a known manner, the mouse 1020 may be configured to control movement of a cursor in a video display 1016, and to operate various graphical user interface (GUI) controls appearing in the video display 1016 with a mouse button. The disk drive or solid state drive 1022 may be configured to accept computer readable media 1024.
[0075] The processing system 1000 may send and receive data over a network via a network interface 1004, allowing the processing system 1000 to communicate with other suitably configured data processing systems, applications, or devices. Network interface 1004 may generally provide data access to any sort of data network, including personal area networks (PANs), local area networks (LANs), wide area networks (WANs), the Internet, and the like. Processing system 1000, which may be an example of a master controller described above, may be implemented in various ways. For example, processing system 1000 may be implemented within on-site, remote, or cloud-based processing equipment
[0076] In examples employing storage to store data, the storage may be a random-access memory (RAM) such as DDR-SDRAM (double data rate synchronous dynamic randomaccess memory). In other examples, the storage may include non-volatile memory such as Read-Only Memory (ROM) or a solid-state drive (SSD) such as Flash memory. The storage in some cases includes other storage media, e.g. magnetic, optical or tape media, a compact disc (CD), a digital versatile disc (DVD) or other data storage media. The storage may be removable or non-removable from the relevant system.
[0077] In examples employing data processing, a processor can be employed as part of the relevant system. The processor can be a general-purpose processor such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field- programmable gate array (FPGA) or another programmable logic device, a discrete gate or transistor logic, discrete hardware components, or any suitable combination thereof designed to perform the data processing functions described herein. In examples involving a neural network, a specialised processor may be employed as part of the relevant system. The specialised processor may be an NPU, a neural network accelerator (NNA) or other version of a hardware accelerator specialised for neural network functions. Additionally or alternatively, the neural network processing workload may be at least partly shared by one or more standard processors, e.g. CPU or GPU. The term “annotation data” has been used throughout the description and is envisaged to correspond with prediction data or inference data in alternative nomenclature. For example, the object detection model (e.g. comprising a neural network) may be trained using annotated images, e.g. images with annotations such as bounding boxes, which serve as a ground truth for the model, e.g. a prediction or inference with a confidence of 100% or 1 when normalised. These annotations may be made by a human for the purposes of training the model, for example. Thus, the object detection of the present disclosure can be taken to involve outputting prediction data or inference data (e.g. instead of “annotation data”) to indicate a prediction or inference of the picking device and / or load-handling device in the image. The prediction data or inference data may be represented as an annotation applied to the image, e.g. a bounding box and / or a label. The prediction data or inference data includes a confidence associated with the prediction or inference of the transport device in the image, for example. The annotation can be applied to the image based on the generated prediction data or inference data, for example. For instance, the image may be updated to include a bounding box surrounding the picking device and / or load-handling device with a label indicating the confidence level of the prediction, e.g. as a percentage value or a normalised value between 0 and 1 .
[0078] It is also to be understood that any feature described in relation to any one example may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the examples, or any combination of any other of the examples. Furthermore, equivalents and modifications not described above may also be employed without departing from the scope of the accompanying claims.
[0079] In this document, “controller” is intended to include any hardware which is suitable for controlling (e.g. providing instructions to) one or more other components. For example, a processor equipped with one or more memories and appropriate software to process data relating to a component or components and send appropriate instructions to the component(s) to enable the component(s) to perform its / their intended function(s).
[0080] Furthermore, the invention can take the form of a computer program embodied as a computer-readable medium having computer executable code for use by or in connection with a computer.
[0081] Reference to an element in the singular is not intended to mean only one unless specifically so stated, but rather “one or more.” The subsequent use of a definite article (e.g., “the” or “said”) with an element (e.g., “the processor”) is not intended to invoke a singular meaning (e.g., “only one”) on the element unless otherwise specifically stated. For example, reference to an element (e.g., “a processor,” “a controller,” “a memory,” “a transceiver,” “an antenna,” “the processor,” “the controller,” “the memory,” “the transceiver,” “the antenna,” etc.), unless otherwise specifically stated, should be understood to refer to one or more elements (e.g., “one or more processors,” “one or more controllers,” “one or more memories,” “one more transceivers,” etc.). The terms “set” and “group” are intended to include one or more elements, and may be used interchangeably with “one or more.” Where reference is made to one or more elements performing functions (e.g., steps of a method), one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function). Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions.
[0082] It will be understood that the above description is given by way of example only and that various modifications may be made by those skilled in the art. Although various embodiments have been described above with a certain degree of particularity, or with reference to one or more individual embodiments, those skilled in the art could make numerous alterations to the disclosed embodiments without departing from the scope of this invention.
Claims
Claims1 . A computer-implemented method of identifying a load-handling device within a system, the system comprising: a first set of parallel rails or tracks and a second set of parallel rails or tracks extending substantially perpendicularly to the first set of rails or tracks in a substantially horizontal plane to form a grid comprising a plurality of grid spaces; one or more load-handling devices, wherein each load-handling device is configured to move along the first and / or second set of tracks, and wherein each loadhandling device comprises a respective unique self-similar pattern, the method comprising: obtaining an image of the system; processing the image with an object detection model trained to detect an instance of a load-handling device comprising a respective unique self-similar pattern; determining, based on the processing, whether the image includes a loadhandling device; and outputting, in response to determining that the image includes a load-handling device, annotation data indicative of the load-handling device in the image.
2. The computer-implemented method of claim 1 , wherein processing the image with an object detection model further comprises: processing the image with a first object detection model trained to detect the unique self-similar pattern at a low level of granularity to detect a load-handling device; and processing the image with a second object detection model trained to detect the unique self-similar pattern at a high level of granularity to detect a specific load-handling device.
3. The computer-implemented method of claim 2, wherein the first object detection model comprises a convolutional neural network; and the second object detection model comprises a convolutional neural network and / or an image similarity model.
4. The computer-implemented method of claim 1 , wherein the object detection model comprises a convolutional neural network.
5. The computer-implemented method of claims 1-4, wherein the method further comprises outputting an updated image including the annotation data.
6. The computer-implemented method of claims 1-5, wherein the annotation data comprises a bounding box for a load-handling device and / or a plurality of bounding boxes, where each bounding box of the plurality of bounding boxes is for a respective load-handling device.
7. The computer-implemented method of claims 1-6, wherein the annotation data comprises information identifying the load-handling device.
8. The computer-implemented method of claims 1-7, wherein the object detection model is further trained to detect a derailment of a load-handling device and / or a warning signal of the load-handling device, wherein the method further comprises determining whether the image contains a load-handling device that has derailed and / or has a warning signal.
9. The computer-implemented method of claim 8, wherein the method further comprises upon determining that a load-handling device has derailed and / or has a warning signal, setting up an exclusion zone around the load-handling device that has derailed and / or has a warning signal.
10. The computer-implemented method of claim 8, wherein the method further comprises updating the annotation data to indicate the load-handling device that has derailed and / or has a warning signal.
11. The computer-implemented of claim 10, wherein the method further comprises outputting an updated image including the annotation date indicative of the load-handling device that has derailed and / or has a warning signal.
12. The computer-implemented method of claims 1-11 , wherein the system comprises an image sensor, wherein obtaining an image of the system comprises obtaining the image from the image sensor.
13. The computer-implemented method of claim 12, wherein the image sensor is located above the grid.
14. The computer-implemented method of claim 13, wherein the system comprises a picking station operational on the grid, the picking station comprising a robotic manipulator comprising the image senor, wherein the robotic manipulator is configured to transfer items between containers received in respective grid cells adjacent the picking station.
15. The computer-implemented method of claim 14, wherein the method further comprises orienting the robotic manipulator of the picking station to direct the image sensor of the robotic manipulator towards the one or more load-handling devices.
16. The computer-implemented method of claims 1-15, wherein each unique selfsimilar pattern comprises a combination of one or more colours and / or one or more geometric shapes.
17. The computer-implemented method of claims 1-16, wherein each unique selfsimilar pattern comprises a fractal.
18. The computer-implemented method of claims 1-17, wherein each unique selfsimilar pattern is generated using a mathematical function, such as a recursive function.
19. The computer-implemented method of claims 1-18, wherein each unique selfsimilar pattern is mapped to a respective load-handling device.
20. The computer-implemented method of claims 1-19, wherein the load-handling device comprises: a body or skeleton mounted on a first set of wheels being arranged to engage with the first set of parallel tracks and a second set of wheels being arranged to engage with the second set of parallel tracks; and / or a drive assembly configured to drive the first or second sets of wheels to move the load-handling device along the first or second set of parallel rails respectively; and / ora direction-change assembly configured to raise or lower the first set of wheels and / or lower or raise the second set of wheels with respect to the body or skeleton to engage and disengage the wheels with the parallel tracks; and / or a container-lifting assembly configured to raise or lower a gripping device in the vertical direction.21 . The computer-implemented method of claim 20, wherein substantially all or at least a portion of an outer surface of the load-handling device, and / or the body or skeleton, and / or the first and / or second sets of wheels, and / or the drive assembly, and / or the direction-change assembly, and / or the container-lifting assembly comprises surface decoration defining the self-similar pattern.
22. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the computer-implemented method of any preceding claim.
23. A computer readable medium comprising the computer program of claim 21 .
24. A data processing system comprising means for carrying out the computer-implemented method of claims 1-21.
25. A system comprising: a first set of parallel rails or tracks and a second set of parallel rails or tracks extending substantially perpendicularly to the first set of rails or tracks in a substantially horizontal plane to form a grid comprising a plurality of grid spaces; one or more load-handling device, wherein each load-handling device is configured to move along the first and / or second set of tracks, and wherein each loadhandling device comprises a respective unique self-similar pattern; an image sensor; and a processor configured to carry out the method of claims 1-21 .
Citation Information
Patent Citations
Apparatus for retrieving units from a storage system
WO2015019055A1
Methods, systems and apparatus for controlling movement of transporting devices
WO2015185628A2
Picking systems and methods
WO2017081281A1
Apparatus for retrieving storage containers from a storage and retrieval system
WO2023025418A1
Detecting a moving picking station on a grid
GB2625052A