Logistics autonomous vehicle with robust object detection, localization and monitoring
By introducing the vision system 400 into the automated logistics system, combining artificial neural networks and machine learning models, the navigation and detection accuracy problems caused by obstruction of stereo cameras are solved, and robust object detection and positioning in complex environments is achieved, and the operation efficiency of autonomous carriers is improved.
Patent Information
- Application Number
- CN202380081212.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-26
- Filing Date
- 2023-09-27
- Publication Date
- 2025-07-11
AI Technical Summary
In existing automated logistics systems, stereo camera pairs may be obstructed or visually impeded, resulting in degradation of image processing, affecting the accuracy of navigation and object detection and positioning of autonomous carriers, especially in the case of tightly packed box units and deformed boxes.
The vision system 400 is adopted, combined with artificial neural network ANN and machine learning model ML, to achieve robust object detection and positioning. Through binocular or monocular visual data analysis, the best detection and positioning protocol is selected to ensure that the box unit can be accurately identified and positioned when the visual system is blocked.
In a complex and limited logistics environment, the accuracy and reliability of object detection and positioning of autonomous carriers are improved, the picking failure rate is reduced, and the rapid transfer box unit is ensured.
Smart Images

Figure CN120303628A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application is a non - provisional application of and claims the benefit of U.S. Provisional Patent Application No. 63 / 377,271, filed Sep. 27, 2022, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0002] The disclosed embodiments generally relate to material handling systems and, more particularly, to the transportation for automated logistics systems. Background Art
[0003] Generally, automated logistics systems, such as automated storage and retrieval systems, use autonomous carriers that transport goods within the automated storage and retrieval system. These autonomous carriers are guided throughout the automated storage and retrieval system by positioning beacons, capacitive or inductive proximity sensors, line - following sensors, reflected beam sensors, and other narrow - focused beam - type sensors. These sensors may provide limited information for implementing the navigation of the autonomous carrier through the storage and retrieval system or limited information regarding the identification and differentiation of hazards that may be present throughout the automated storage and retrieval system.
[0004] Autonomous carriers may also be guided throughout the automated storage and retrieval system by a vision system using stereo or binocular cameras. However, in a logistics environment, due to, for example, the blocking or visual obstruction of one of the cameras in the stereo camera pair (e.g., due to payloads carried by the autonomous carrier, storage structures, etc.) and / or blurred line of sight, the stereo camera pair may be defective or not always available; or image processing may degrade due to processing repetitive image data or images that are otherwise not suitable (e.g., blurred, etc.) for guiding and positioning the autonomous carrier within the automated storage and retrieval system. Brief Description of the Drawings
[0005] The above - mentioned aspects and other features of the disclosed embodiments are illustrated in the following description in conjunction with the accompanying drawings, in which:
[0006] Figure 1A is a schematic diagram of a logistics facility incorporating aspects of the disclosed embodiments;
[0007] Figure 1B is according to aspects of the disclosed embodiments of Figure 1A the logistics facility;
[0008] Figure 2 is according to aspects of the disclosed embodiments of Figure 1A the autonomous - guided carrier of the logistics facility;
[0009] Figure 3A is according to aspects of the disclosed embodiments ofFigure 2 Schematic diagram of a part of an autonomous guided vehicle;
[0010] Figure 3B is according to an aspect of the disclosed embodiments Figure 2 Schematic diagram of a part of an autonomous guided vehicle;
[0011] Figure 3C is according to an aspect of the disclosed embodiments Figure 2 Schematic diagram of a part of an autonomous guided vehicle;
[0012] Figure 4A , Figure 4B and Figure 4C is according to an aspect of the disclosed embodiments Figure 2 Example of image data captured by the vision system of an autonomous guided vehicle;
[0013] Figure 5 is according to an aspect of the disclosed embodiments Figure 2 Schematic diagram of a part of an autonomous guided vehicle;
[0014] Figure 6 Exemplary flowchart of a method according to an aspect of the disclosed embodiments;
[0015] Figure 7A and Figure 7B Schematic diagram of a calibration fixture or jig according to an aspect of the disclosed embodiments;
[0016] Figure 8 is according to an aspect of the disclosed embodiments Figure 2 Exemplary illustration of a computer model (and parts thereof) of an autonomous guided vehicle;
[0017] Figure 9A is according to an aspect of the disclosed embodiments from Figure 2 Exemplary monocular image of a camera of the vision system of an autonomous guided vehicle;
[0018] Figure 9B is according to an aspect of the disclosed embodiments from Figure 9A Exemplary depth map generated from an exemplary monocular image;
[0019] Figure 9C is according to an aspect of the disclosed embodiments using Figure 2 Exemplary illustration generated from an aberration map of a pair of cameras of the vision system of an autonomous guided vehicle;
[0020] Figure 10 is according to an aspect of the disclosed embodiments Figure 2 Exemplary user interface of an autonomous guided vehicle;
[0021] Figure 11A is an exemplary image frame that images video stream data of an object obtained from one or more forward navigation cameras or one or more rearward navigation cameras of an autonomous guided vehicle according to aspects of the disclosed embodiments. Note that the detected image features are only bounded by bounding boxes for exemplary purposes and can be identified in the image frame in any suitable manner;
[0022] Figures 11B to 11F is an exemplary image frame that images video stream data of an object obtained from one or more bin monitoring cameras of an autonomous guided vehicle according to aspects of the disclosed embodiments. Note that the detected image features are only bounded by bounding boxes for exemplary purposes and can be identified in the image frame in any suitable manner;
[0023] Figure 12A and Figure 12B is according to aspects of the disclosed embodiments Figure 2 exemplary augmented image of the vision system of the autonomous guided vehicle;
[0024] Figure 13 is an exemplary flowchart of a method according to aspects of the disclosed embodiments;
[0025] Figure 14 is an exemplary flowchart of a method according to aspects of the disclosed embodiments;
[0026] Figure 15 is according to aspects of the disclosed embodiments Figure 1A exemplary schematic diagram of a calibration station of a logistics facility; and
[0027] Figure 16 is an exemplary flowchart of a method according to aspects of the disclosed embodiments. DETAILED DESCRIPTION
[0028] Figure 1A and 1B illustrates an exemplary automated storage and retrieval system 100 according to aspects of the disclosed embodiments. Although aspects of the disclosed embodiments will be described with reference to the figures, it should be understood that aspects of the disclosed embodiments can be embodied in various forms. Additionally, any suitable size, shape, or type of element or material can be used.
[0029] Aspects of the disclosed embodiments provide a logistics autonomous guided vehicle 110 (referred to herein as an autonomous guided vehicle) having intelligent autonomous and collaborative operations. For example, the autonomous guided vehicle 110 includes a vision system 400 (see Figure 2) having at least one camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B configured to generate video stream data imaging of an object in a logistics space such as an operating environment or space of a storage and retrieval system 100. The object is one or more of the following: at least a portion of the autonomous guided vehicle 110 frame 200, at least a portion of a payload (e.g., a bin CU), at least a portion of a transfer arm 210A, and at least a portion of a logistics item (other bin CU) or structure (e.g., a storage and retrieval system 100) in the logistics space other than the autonomous guided vehicle 110. For illustrative purposes, the vision system 400 uses at least stereo or binocular vision configured to perform detection of bin CUs and objects (such as facility structures and non-desired foreign objects / transient materials) within a logistics facility such as the automated storage and retrieval system 100 and autonomous guided vehicle positioning within the automated storage and retrieval system 100. The vision system 400 also provides cooperative vehicle operation by providing images (static images or video streams (live or recorded)) to an operator of the automated storage and retrieval system 100, where in some aspects, those images are provided as augmented images as described herein and as Figure 10 shown therein.
[0030] As will be described in more detail herein, the autonomous guided vehicle 110 includes a controller 122 programmed with one or more machine learning models ML and one or more artificial neural networks ANN, and the one or more artificial neural networks ANN access data from the vision system 400 to perform robust bin / object detection and positioning regardless of whether one or more cameras of the vision system 400 are blocked. Here, the bin / object detection is robust because the bin / object can be detected and positioned by using the artificial neural network ANN and the machine learning model ML, and even in cases where stereo vision cannot be used in a highly constrained system or operating environment, the artificial neural network ANN and the machine learning model ML provide detection and positioning effects comparable to those obtained by using the stereo vision. Highly constrained systems include but are not limited to at least the following limitations: the spacing between adjacent bins is a tight stacking spacing, the autonomous guided vehicle is configured to pick bins from the bottom (lift from below), bins of different sizes are distributed in a Gaussian distribution within a storage array SA, the bins may be deformed, and the bins may be placed on a support surface in an irregular manner, all of which affect the transfer of bin units CU between a storage shelf 555 (or other bin holding position) and the autonomous guided vehicle 110.
[0031] Another limitation of the ultra-constrained system is the transfer time of the autonomous guided vehicle 110 for transferring the bin unit between the payload platform 210B of the autonomous guided vehicle 110 and the bin holding positions (e.g., storage space, staging area, transfer station, or other bin holding positions described herein). Here, the transfer time for bin transfer is about 10 seconds or less. Thus, the vision system 400 differentiates the bin position and attitude (or the holding station position and attitude) in less than about two seconds or less than about half a second.
[0032] As mentioned above, the bins CU stored in the storage and retrieval system have a Gaussian distribution in terms of the size of the bins in the picking aisle 130A and in terms of the size of the bins in the entire storage array SA (see Figure 4A ), such that when picking and placing bins, the size of any given storage space on the storage shelf 555 varies dynamically (e.g., dynamic Gaussian bin size distribution). Thus, as described herein, the autonomous guided vehicle 110 is configured to identify the bins held in storage spaces of dynamic size (depending on the bins held therein), regardless of whether one of the stereo camera pairs implementing object detection and localization is blocked or not.
[0033] In addition, for example, as Figure 4A can be seen, the bins CU are placed in a closely coupled or closely spaced relationship on the storage shelf 555 (or other holding stations), where the distance DIST between adjacent bin units CU is half the distance between the storage shelf caps 444. The distance / width DIST between the caps 444 of the support strip 520L is about 2.5 inches. The close spacing of the bins CU can be complex (i.e., the spacing can be less than half the distance between the storage shelf caps 444), where the bins CU (e.g., deformed bins, see Figures 4A to 4C , which shows an open flap bin deformation) can exhibit deformations (e.g., such as bulging sides, open flaps, protruding sides) and / or can be skewed relative to the caps 444 on which the bins CU sit (i.e., the front of the bin may not be parallel to the front of the storage shelf 555 and the sides of the bin may not be parallel to the caps 555 of the storage shelf 555 (see Figure 4A )). Bin deformation and skewed bin placement can further reduce the spacing between adjacent bins. Thus, as described herein, the autonomous guided vehicle is configured to determine picking interference between closely spaced adjacent bins, regardless of whether one of the stereo camera pairs implementing object detection and localization is blocked or not.
[0034] Also note that the height HGT of the caps 444 is about 2 inches, where the space envelope ENV between the caps 444 is about 1.7 inches wide and about 1.2 inches high, and the fork 210AT of the transfer arm 210A of the autonomous guided vehicle 110 is inserted below the bin unit CU in this space envelope for picking bins from the storage shelf 555 and placing bins on the storage shelf (see for exampleFigure 3A , Figure 3C and Figure 4A ). The picking of the case CU by the bottom of the autonomous guided vehicle must dock with the case CU held on the storage shelf 555 at the picking / case support plane (defined by the case placement surface 444S of the cap 444, see Figure 4A ), without impact between the fork 210AT of the transfer arm 210A of the autonomous guided vehicle 110 and the cap 444 / slat 520L, without impact between the fork 210AT and the adjacent case (i.e., the unpicked case), and without impact between the picked case and the adjacent unpicked case, all of which can be implemented by placing the fork 210AT within the envelope ENV between the caps 444. Thus, as described herein, the autonomous guided vehicle is configured to detect and locate the spatial envelope ENV for inserting the fork 210AT of the transfer arm 210A below the predetermined case CU for picking the case, regardless of whether one of the stereo camera pairs for implementing object detection and location is blocked or not.
[0035] The highly constrained system described above requires the robustness of the vision system and can be regarded as defining the robustness of the vision system 400, since the vision system 400 is configured to adapt to the limitations mentioned above (even when the stereo vision provided by the vision system 400 is not available) and can provide attitude and positioning information for the case CU and / or the autonomous guided vehicle 110, with an autonomous guided vehicle picking failure rate of approximately one picking failure per approximately one million picks.
[0036] The robustness of the vision system 400 is implemented at least in part by: the controller 122 includes a control module (referred to herein as the depth guidance module DC), which includes an artificial neural network ANN and is configured to select a detection / localization protocol via the artificial neural network ANN from either (alternatively, both) a computer vision protocol (e.g., which uses binocular / stereo vision) and a machine learning protocol (e.g., which uses monocular vision data analysis that combines the use of a machine learning model ML and an artificial neural network ANN). As described herein, the controller 122 is configured such that binocular vision and monocular vision from video stream data imaging can be selectively utilized to perform object detection and localization within a predetermined reference frame (e.g., the global reference frame GREF and / or the autonomous guided vehicle reference frame BREF) based on the video stream data imaging. Each of the binocular vision object detection and localization and the monocular vision object detection and localization can be selected by the controller 122 as needed. The controller 122 is configured such that the binocular vision object detection and localization and the monocular vision object detection and localization can be interchangeably selected by the controller 122. The controller 122 has a selector 122SL (implemented using the depth guidance DC described herein), which is configured to select between the binocular vision object detection and localization and the monocular vision object detection and localization as needed based on the detection of a predetermined operating characteristic of the autonomous guided vehicle 110 (such as the video stream data imaging recorded by the controller does not support binocular vision object detection and localization). Here, aspects of the disclosed embodiments provide that the autonomous guided vehicle 110 obtains data with binocular vision (where the autonomous guided vehicle 110 traverses the storage and retrieval system 100), and when the data from the binocular vision is not suitable for performing object detection and localization as described herein, switches to monocular vision as needed (or during operation). The data obtained by the monocular vision object detection and localization enables the controller 122 to re-perform object detection and localization for any given image frame data without knowing the data from the previously obtained image frame data.
[0037] The computer vision protocol and the machine learning protocol operate simultaneously, and the depth guidance module DC determines which protocol provides the highest detection and / or localization confidence, and selects the protocol with the highest detection and / or localization confidence. Note that when the machine learning protocol is selected, using monocular vision data (processed by one or more machine learning models ML and one or more artificial neural networks ANN) has a comparable effect in place of binocular or stereo (the terms binocular and stereo are used interchangeably herein) vision data (i.e., available, unobstructed, unoccluded, and focused binocular vision or other non-defective binocular vision data of the computer vision protocol).
[0038] According to aspects of the disclosed embodiments, Figure 1A and Figure 1BThe automated storage and retrieval system 100 therein can be set, for example, in a retail distribution (logistics) center or a warehouse to fulfill orders received from retail stores to replenish goods shipped in boxes, packages, and / or parcels. The terms box, package, and parcel are used interchangeably herein and, as previously described, can be any container that can be used by a producer for shipping and can be filled with one or more product units. One or more boxes as used herein refers to box, package, or parcel units that are not stored on pallets, on handling boxes, etc. (e.g., not contained). Note that a box unit CU (which is also referred to herein as a mixed box, box, and shipping unit) can include boxes of articles / units (e.g., a box of soup cans, a box of oatmeal, etc.), or is a separate article / unit suitable for retrieval from or placement on a pallet. According to an exemplary embodiment, a shipping box or box unit (e.g., a cardboard box, a barrel, a box, a crate, a jug, a shrink-wrapped pallet or group, or any other suitable device for holding a box unit) can have a variable size and can be used to hold a box unit during shipping and can be configured such that they can be palletized for shipping. A box unit can also include a handling box, box, and / or container for one or more separate goods, and the separate goods are unpacked / removed from the original packaging (commonly referred to as unpacked goods) at an order filling station and placed into a handling box, box, and / or container (collectively referred to as a handling box) together with one or more other separate goods of a mixed or common type. Note that when, for example, an incoming bundle or pallet (e.g., from a manufacturer or supplier of box units) arrives at the storage and retrieval system for replenishing the automated storage and retrieval system 100, the contents of each pallet can be uniform (e.g., each pallet holds a predetermined number of the same items, i.e., one pallet holds soup and another pallet holds oatmeal). As can be appreciated, the boxes of such pallet loads can be substantially similar, or in other words, homogeneous boxes (e.g., similar sizes), and can have the same SKU (otherwise, as previously described, the pallet can be a "rainbow" pallet with layers formed by homogeneous boxes). When the pallet leaves the storage and retrieval system, as the box or handling box is filled to replenish an order, the pallet can accommodate any suitable number and combination of different box units (e.g., each pallet can hold a combination of different types of box units, i.e., the pallet holds a combination of canned soup, oatmeal, beverage packs, cosmetics, and household cleaners). The boxes combined onto a single pallet can have different sizes and / or different SKUs.
[0039] The automated storage and retrieval system 100 can generally be described as a storage and retrieval engine 190 coupled to a palletizer 162. Now, more specifically and still referring to Figure 1A and 1B , the storage and retrieval system 100 can be configured for installation in, for example, an existing warehouse structure or adapted to a new warehouse structure. As previously described, Figure 1A and Figure 1BThe automated storage and retrieval system 100 shown is representative and may include, for example, infeed and outfeed conveyors terminating at respective transfer stations 170, 160, lift modules 150A, 150B, a storage structure 130, and a number of autonomous guided vehicles 110. Note that the storage and retrieval engine 190 is formed at least by the storage structure 130 and the autonomous guided vehicles 110 (and in some aspects, the lift modules 150A, 150B; however, in other aspects, the lift modules 150A, 150B may form a vertical sorter in addition to the storage and retrieval engine 190, as described in U.S. Patent Application No. 17 / 091265, filed November 6, 2020, and titled "Pallet Building System with Flexible Sorting," the disclosure of which is incorporated herein by reference in its entirety). In alternative aspects, the automated storage and retrieval system 100 may also include a robot or robotic transfer station (not shown), which may provide an interface between the autonomous guided vehicles 110 and the lift modules 150A, 150B. The storage structure 130 may include a multi-level storage rack module, where each storage structure layer 130L of the storage structure 130 includes a respective picking aisle 130A and a transfer deck 130B for transferring bin units between any storage area of the storage structure 130 and the shelves of the lift modules 150A, 150B. In one aspect, the picking aisle 130A is configured to provide a guided travel of the autonomous guided vehicle 110 (such as along a track 130AR), while in other aspects, the picking aisle is configured to provide an unconstrained travel of the autonomous guided vehicle 110 (e.g., the picking aisle is open and indeterminate with respect to the guidance / travel of the autonomous guided vehicle 110). The transfer deck 130B has an open and indeterminate robotic support travel surface, and the autonomous guided vehicle 110 travels along the robotic support travel surface under the guidance and control provided by any suitable robotic steering. In one or more aspects, the transfer deck 130B has a plurality of lanes, and the autonomous guided vehicle 110 freely transitions between these lanes for accessing the picking aisle 130A and / or the lift modules 150A, 150B. As used herein, "open and indeterminate" means that the travel surface of the picking aisle and / or the transfer deck does not have mechanical constraints (such as guide rails) that define the travel of the autonomous guided vehicle 110 along any given path along the travel surface).
[0040] The picking aisle 130A and the transfer deck 130B also allow the autonomous guided vehicle 110 to place the case units CU into the picking inventory and retrieve the ordered case units CU (and define different locations where the robot performs autonomous operations, although any number of locations in the storage structure (e.g., deck, aisle, storage rack, etc.) can be one or more of the different locations). In an alternative aspect, each floor further includes a corresponding transfer station 140, which provides an indirect case transfer between the autonomous guided vehicle 110 and the lift modules 150A, 150B. The autonomous guided vehicle 110 can be configured to place case units (such as the above-mentioned retail goods) into the picking inventory in one or more storage structure layers 130L of the storage structure 130, and then selectively retrieve the ordered case units for shipping the ordered case units to, for example, a store or other suitable location. The incoming transfer station 170 and the outgoing transfer station 160 can operate with their corresponding lift modules 150A, 150B for two-way transfer of the case units CU to and from one or more storage structure layers 130L of the storage structure 130. Note that although the lift modules 150A, 150B are described as dedicated inbound lift module 150A and outbound lift module 150B, in an alternative aspect, each of the lift modules 150A, 150B can be used for both inbound and outbound transfer of case units from the storage and retrieval system 100.
[0041] As can be appreciated, the storage and retrieval system 100 can include a plurality of infeed lift modules and outfeed lift modules 150A, 150B, e.g., which can be accessed by the autonomous guided vehicle 110 of the storage and retrieval system 100 (e.g., indirectly via the transfer station 140, or directly via the bins with transfer between the lift modules 150A, 150B and the autonomous guided vehicle 110), such that one or more bin units that are unaccommodated (e.g., bin units not held in trays) or accommodated (within trays or handling bins) can be transferred from the lift modules 150A, 150B to each storage space on the corresponding level and from each storage space to any one of the lift modules 150A, 150B on the corresponding level. The autonomous guided vehicle 110 can be configured to transfer bin CUs (also referred to herein as bin units) between the storage spaces 130S (e.g., located in the picking aisles 130A or other suitable storage spaces / bin unit staging areas disposed along the transfer decks 130B) and the lift modules 150A, 150B. Generally, the lift modules 150A, 150B include at least one movable payload support that can move bin units between the infeed transfer station 160 and the outfeed transfer station 170 and the corresponding levels of the storage spaces where the bin units are stored and retrieved. The lift modules can have any suitable configuration, e.g., such as a reciprocating lift, or any other suitable configuration. The lift modules 150A, 150B include any suitable controller (such as the control server 120 or other suitable controllers coupled to the control server 120, the warehouse management system 2500, and / or the palletizer controllers 164, 164'), and can form a sorter or a picker in a manner similar to that described in U.S. Patent Application No. 16 / 444,592, filed on June 18, 2019, and titled "Vertical Sorter for Product Order Fulfillment", the disclosure of which is incorporated herein by reference in its entirety.
[0042] Automated storage and retrieval system 100 may include a control system, which for example includes one or more control servers 120, and the control servers are communicably connected via a suitable communication and control network 180 to the infeed and outfeed conveyors and transfer stations 170, 160, lift modules 150A, 150B, and autonomous guided vehicles 110. The communication and control network 180 may have any suitable architecture. For example, it may incorporate various programmable logic controllers (PLCs), such as for commanding the operation of the infeed and outfeed conveyors and transfer stations 170, 160, lift modules 150A, 150B, and other suitable systems for automation. The control server 120 may include advanced programming of a case management system (CMS) that implements a management case process system. The network 180 may further include suitable communication for implementing two-way docking with the autonomous guided vehicle 110. For example, the autonomous guided vehicle 110 may include an on-board processor / controller 122. The network 180 may include a suitable two-way communication suite that enables the autonomous guided vehicle controller 122 to request or receive commands from the control server 120 to perform a desired transport of case units (such as placing in a storage location or retrieving from a storage location), and to send desired autonomous guided vehicle 110 information and data to the control server 120, including the ephemeris, status, and other desired data of the autonomous guided vehicle 110. As Figure 1A and 1B seen, the control server 120 may further be connected to a warehouse management system 2500 for providing, for example, inventory management and customer order fulfillment information to the CMS-level program of the control server 120. As previously described, the control server 120 and / or the warehouse management system 2500 allow for at least a degree of collaborative control of the robot 110 via a user interface UI, which will be further described below. A suitable example of an automated storage and retrieval system arranged for holding and storing case units is described in U.S. Patent No. 9,096,375, issued on August 4, 2015, the disclosure of which is incorporated herein by reference in its entirety.
[0043] Now refer to Figure 1A 、 1BWith reference to FIGS. 1 and 2, the autonomous guided vehicle 110 includes a frame 200 having an integral payload support or deck 210B (also referred to herein as a payload holder). The frame 200 has a front end 200E1 and a rear end 200E2 that define a longitudinal axis LAX of the autonomous guided vehicle 110. The frame 200 can be made of any suitable material (e.g., steel, aluminum, composite materials, etc.) and includes a bin handling assembly 210 configured to handle bins / payloads transported by the autonomous guided vehicle 110. The bin handling assembly 210 includes any suitable payload deck 210B (also referred to herein as a payload compartment or payload holder) on which the payload is placed for transportation and / or any suitable transfer arm 210A (also referred to herein as a payload handler) connected to the frame. The transfer arm 210A is configured to (autonomously) transfer payloads (such as bin units CU) using a flat indeterminate placement surface disposed in the payload deck 210B to and from the payload deck 210B of the autonomous guided vehicle 110 and the storage location of the payload CU in the storage array SA (such as storage shelf 555 (see Figure 2) storage space 130S thereon, shelves of lift modules 150A, 150B, staging areas, transfer stations, and / or any other suitable storage locations), wherein the storage location 130S in the storage array SA is separate and distinct from the transfer arm 210A and the payload platform 210B. The transfer arm 210A is configured to extend laterally in the direction LAT and vertically in the direction VER to transport payloads to and from the payload platform 210B. Examples of suitable payload platforms 210B and transfer arms 210A and / or autonomous guided vehicles to which aspects of the disclosed embodiments may be applied can be found in: U.S. Patent No. 11,078,017, entitled "Automated Robot with Transfer Arm," issued on August 3, 2021; U.S. Patent No. 7,591,630, entitled "Material Handling System Using Autonomous Transfer and Transport Vehicles," issued on September 22, 2009; U.S. Patent No. 7,991,505, entitled "Material Handling System Using Autonomous Transfer and Transport Vehicles," issued on August 2, 2011; U.S. Patent No. 9,561,905, entitled "Autonomous Transport Vehicle," issued on February 7, 2017; U.S. Patent No. 9,082,112, entitled "Autonomous Transport Vehicle Charging System," issued on July 14, 2015; U.S. Patent No. 9,850,079, entitled "Storage and Retrieval System Transport Vehicle," issued on December 26, 2017; U.S. Patent No. 9,187,244, entitled "Robot Payload Alignment and Sensing," issued on November 17, 2015; U.S. Patent No. 9,499,338, entitled "Automated Robot Transfer Arm Drive System," issued on November 22, 2016; U.S. Patent No. 8,965,619, entitled "Robot with High-Speed Stability," issued on February 24, 2015; U.S. Patent No. 9,008,884, entitled "Robot Position Sensing," issued on April 14, 2015; U.S. Patent No. 8,425,173, entitled "Autonomous Transport for Storage and Retrieval Systems," issued on April 23, 2013; and U.S. Patent No. 8,696,010, entitled "Suspension System for Autonomous Transport," issued on April 15, 2014, the disclosures of which are incorporated herein by reference in their entirety.
[0044] The frame 200 includes one or more idler wheels or casters 250 disposed adjacent to the front end 200E1. Suitable examples of casters can be found in U.S. Provisional Patent Application No. 17 / 664948, filed on May 25, 2022 (), titled "Autonomous Transport Vehicle with Synthetic Carrier Dynamic Response" (having Attorney Docket No. 1127P015753-US(PAR)), and U.S. Patent Application No. 17 / 664838, filed on May 26, 2021, titled "Autonomous Transport Vehicle with Manipulation" (having Attorney Docket No. 1127P015753-US(PAR)), the disclosures of which are incorporated herein by reference in their entirety. The frame 200 also includes one or more drive wheels 260 disposed adjacent to the rear end 200E2. In other aspects, the positions of the casters 250 and the drive wheels 260 can be reversed (e.g., the drive wheels 260 are disposed at the front end 200E1 and the casters 250 are disposed at the rear end 200E2). Note that in some aspects, the autonomous guided vehicle 110 is configured to travel with the front end 200E1 leading the direction of travel or with the rear end 200E2 leading the direction of travel. In one aspect, casters 250A, 250B (which are substantially similar to the casters 250 described herein) are located at the respective front corner portions of the frame 200 at the front end 200E1, and drive wheels 260A, 260B (which are substantially similar to the drive wheels 260 described herein) are located at the respective rear corner portions of the frame 200 at the rear end 200E2 (e.g., support wheels are located at each of the four corner portions of the frame 200), such that the autonomous guided vehicle 110 can stably traverse the transfer deck 130B and the picking aisle 130A of the storage structure 130.
[0045] The autonomous guided vehicle 110 includes a drive section 261D connected to the frame 200, the drive section 261D having drive wheels 260 that support the autonomous guided vehicle 110 on a traverse / rolling surface 284, wherein the drive wheels 260 effect the traverse of the vehicle on the traverse surface 284 to move the autonomous guided vehicle 110 on the traverse surface 284 in a facility (such as a warehouse, store, etc.). The drive section 261D has at least a pair of traction drive wheels 260 (also referred to as drive wheels 260, see drive wheels 260A, 260B) across the drive section 261D. The drive wheels 260 have a fully independent suspension 280 that couples each of the at least a pair of drive wheels 260A, 260B to the frame 200 and is configured to maintain a substantially stable traction contact area between at least one of the drive wheels 260A, 260B and the rolling / travel surface 284 (also referred to as the autonomous vehicle travel surface 284) on a rolling surface transition (such as a bump, a surface transition, etc.). Suitable examples of the fully independent suspension 280 can be found in U.S. Patent Application No. 17 / 664948, filed on May 25, 2022, titled "Autonomous Transport Vehicle with Synthetic Vehicle Dynamic Response" (having Attorney Docket No. 1127P015753-US(PAR)), the disclosure of which is incorporated herein by reference in its entirety.
[0046] The autonomous guided vehicle 110 includes a physical property sensor system 270 (also referred to as an autonomous navigation operation sensor system) connected to the frame 200. The physical property sensor system 270 has electromagnetic sensors. Each of the electromagnetic sensors responds to the interaction or docking of a sensor that emits or generates an electromagnetic beam or electromagnetic field with the physical properties of (such as a storage structure or a transient object (such as a bin unit CU, debris, etc.)), wherein the electromagnetic beam or electromagnetic field is disturbed due to the interaction or docking with the physical properties. The disturbance in the electromagnetic beam is detected by the electromagnetic sensors and is used by them to effect the sensing of the physical properties, wherein the physical property sensor system 270 is configured to generate sensor data embodying at least one of information on the navigation attitude or position of the vehicle (relative to the storage and retrieval system or facility in which the autonomous guided vehicle 110 operates) and information on the payload attitude or position (relative to the storage location 130S or the payload stage 210B).
[0047] For example, and merely for purposes of illustration, the physical property sensor system 270 includes a laser sensor 271, an ultrasonic sensor 272, a barcode scanner 273, a position sensor 274, a line sensor 275, a bin sensor 276 (e.g., for sensing bin units either within the payload deck 210B carried on the vehicle 110 or on a storage shelf outside the vehicle 110), an arm proximity sensor 277, a vehicle proximity sensor 278, or one or more of any other suitable sensors for sensing the position of the vehicle 110 or the payload (e.g., bin unit CU). In some aspects, the supplementary navigation sensor system 288 may form part of the physical property sensor system 270. Suitable examples of sensors that may be included in the physical property sensor system 270 are described in U.S. Patent No. 8,425,173, titled "Autonomous Transport for Storage and Retrieval Systems," issued on April 23, 2013; U.S. Patent No. 9,008,884, titled "Robot Position Sensing," issued on April 14, 2015; and U.S. Patent No. 9,946,265, titled "Robots with High-Speed Stability," issued on April 17, 2018, the disclosures of which are incorporated herein by reference in their entirety.
[0048] The sensors of the physical property sensor system 270 may be configured to provide, for example, perception of the autonomous guided vehicle 110's environment and external objects, as well as monitoring and control of internal subsystems. For example, the sensors may provide guidance information, payload information, or any other suitable information for use in the operation of the autonomous guided vehicle 110.
[0049] The barcode scanner 273 may be mounted on the autonomous guided vehicle 110 at any suitable location. The barcode scanner 273 may be configured to provide the absolute position of the autonomous guided vehicle 110 within the storage structure 130. The barcode scanner 273 may be configured to verify aisle references and positions on the transfer deck by, for example, reading barcodes located on, for example, transfer decks, pick aisles, and transfer station floors, to verify the position of the autonomous guided vehicle 110. The barcode scanner 273 may also be configured to read barcodes located on items stored in the shelves 555.
[0050] The position sensor 274 may be mounted to the autonomous guided vehicle 110 at any suitable location. The position sensor 274 may be configured to detect reference fiducial features (or count the slats 520L of the storage shelf 555) (e.g., see Figure 5A), for determining the position of the carrier 110 relative to the shelves of, for example, a picking aisle 130A (or a staging area / transfer station located near the transfer deck 130B or the elevator 150). The controller 122 can use reference datum information to, for example, calibrate the odometer of the carrier and allow the autonomous guided carrier 110 to stop, where the support forks 210AT of the transfer arm 210A are positioned to be inserted into the space between the slats 520L (see, for example Figure 5 A). In one exemplary embodiment, the carrier 110 can include position sensors 274 on the drive end (rear end) 200E2 and the driven end (front end) 200E1 of the autonomous guided carrier 110 to allow reference datum detection, regardless of which end of the autonomous guided carrier 110 faces the direction of travel of the autonomous guided carrier 110.
[0051] The line sensor 275 can be any suitable sensor mounted to the autonomous guided carrier 110 in any suitable position, for example, for illustrative purposes only, mounted on the frame 200 and arranged to be adjacent to the drive end (rear end) 200E2 and the driven end (front end) 200E1 of the autonomous guided carrier 110. For illustrative purposes only, the line sensor 275 can be a diffuse infrared sensor. The line sensor 275 can be configured to detect a guide line 900 provided on the bottom plate of, for example, the transfer deck 130B (see Figure 1B ). The autonomous guided carrier 110 can be configured to follow these guide lines when traveling on the transfer deck 130B and define the turning ends when the carrier transitions between boarding and leaving the transfer deck 130B. The line sensor 275 can also allow the carrier 110 to detect index references for determining absolute positioning, where these index references are generated by intersecting guide lines 199 (see Figure 1B ).
[0052] The bin sensor 276 can include a bin overhang sensor and / or other suitable sensors configured to detect the position / attitude of the bin unit CU within the payload stage 210B. The bin sensor 276 can be any suitable sensor positioned on the carrier such that the field of view of the sensor spans the payload stage 210B adjacent to the top surface of the support forks 210AT (see Figure 3A and 3B ). The bin sensor 276 can be provided at the edge of the payload stage 210B (e.g., adjacent to the transport opening 1199 of the payload stage 210B) to detect any bin unit CU that at least partially extends outside the payload stage 210B.
[0053] The arm proximity sensor 277 can be mounted on the autonomous mobile vehicle 110 at any suitable location (such as, for example, on the transfer arm 210A). The arm proximity sensor 277 can be configured to sense objects around the transfer arm 210A and / or the support fork 210AT of the transfer arm 210A as the transfer arm 210A is raised / lowered and / or as the support fork 210AT is extended / retracted.
[0054] The laser sensor 271 and the ultrasonic sensor 272 can be configured to allow the autonomous mobile vehicle 110 to position itself relative to each bin unit forming the load carried by the autonomous mobile vehicle 110 before picking the bin unit from, for example, the storage rack 555 and / or the elevator 150 (or any other location suitable for retrieving the payload). The laser sensor 271 and the ultrasonic sensor 272 can also allow the vehicle to position itself relative to the empty storage location 130S for placing the bin unit in those empty storage locations 130S. The laser sensor 271 and the ultrasonic sensor 272 can also allow the autonomous mobile vehicle 110 to confirm that the storage space (or other load storage location) is empty before a payload carried by the autonomous mobile vehicle 110 is, for example, deposited into the storage space 130S. In one example, the laser sensor 271 can be mounted on the autonomous mobile vehicle 110 at a suitable location for detecting the edge of an item to be transferred to (or from) the autonomous mobile vehicle 110. The laser sensor 271 can work in conjunction with, for example, a retroreflective tape (or other suitable reflective surface, coating, or material) located at the back of the rack 555 such that the sensor can "see" the back of the storage rack 555 at all times. The reflective tape located at the back of the storage rack allows the laser sensor 1715 to be substantially unaffected by the color, reflectivity, roundness, or other suitable characteristics of the items located on the rack 555. The ultrasonic sensor 272 can be configured to measure the distance from the autonomous mobile vehicle 110 to the first item in a predetermined storage area of the rack 555 to allow the autonomous mobile vehicle 110 to determine the pick depth (e.g., the distance the support fork 210AT travels into the rack 555 to pick an item from the rack 555). One or more of the laser sensor 271 and the ultrasonic sensor 272 can allow the detection of bin orientation (e.g., the skew of the bin within the storage rack 555) by, for example, measuring the distance between the autonomous mobile vehicle 110 and the front surface of the bin unit to be picked when the autonomous mobile vehicle 110 stops adjacent to the bin unit to be picked. The bin sensor can allow verification of the placement of the bin unit on, for example, the storage rack 555 by scanning the bin unit after the bin unit is placed on the rack.
[0055] The vehicle proximity sensor 278 may also be disposed on the frame 200 for determining the position of the autonomous guided vehicle 110 in the picking aisle 130A and / or relative to the lift 150. The vehicle proximity sensor 278 is located on the autonomous guided vehicle 110 to sense a target or position determination feature disposed on the track 130AR along which the vehicle 110 travels through the picking aisle 130A (and / or at the wall of the transfer area 195 and / or at the lift 150 proximity position). The position of the target on the track 130AR is a known position to form an incremental or absolute encoder along the track 130AR. The vehicle proximity sensor 278 senses the target and provides sensor data to the controller 122 such that the controller 122 determines the position of the autonomous guided vehicle 110 along the picking aisle 130A based on the sensed target.
[0056] The sensors of the physical property sensor system 270 are communicatively coupled to the controller 122 of the autonomous guided vehicle 110. As described herein, the controller 122 is operatively connected to the drive section 261D and / or the transfer arm 210A. The controller 122 is configured to determine the vehicle attitude and position (e.g., in up to six degrees of freedom X, Y, Z, Rx, Ry, Rz) based on the information of the physical property sensor system 270 and implement the autonomous guided vehicle 110 with independent guidance to traverse the storage and retrieval facility / system 100. The controller 122 is also configured to determine the attitude and position of the payload (e.g., the bin unit CU) (either on board or outside the autonomous guided vehicle 110) based on the information of the physical property sensor system 270, implement independent bottom picking (e.g., lift the bin unit CU from below the bin unit CU) and place the payload CU to and from the storage location 130S and independently bottom pick and place the payload CU in the payload station 210B.
[0057] Reference Figure 1A 、 1B, 2, 3A, and 3B, as described above, the autonomous guided vehicle 110 includes a supplemental or auxiliary navigation sensor system 288 connected to the frame 200. The supplemental navigation sensor system 288 supplements the physical property sensor system 270. The supplemental navigation sensor system 288 is at least in part a vision system 400 having cameras configured to capture image data that informs at least one of the vehicle navigation attitude or position (relative to the storage and retrieval system structure or facility within which the vehicle 110 operates) and the payload attitude or position (relative to the storage location or payload station 210B), supplementing the information of the physical property sensor system 270. It should be noted that the term "camera" as used herein is a static imaging or video imaging device, which includes one or more of a two-dimensional camera, a two-dimensional camera having RGB (red, green, blue) pixels, a three-dimensional camera having XYZ+A resolution (where XYZ is the three-dimensional reference frame of the camera and A is one of radar echo intensity, flight time stamp, or other distance determination stamp / indicator), and an RGB / XYZ camera (which includes both RGB and three-dimensional coordinate system information), and non-limiting examples of them are provided herein.
[0058] Reference Figure 2 , 3A and 3B, the vision system 400 includes one or more of the following: bin unit monitoring cameras 410A, 410B, forward navigation cameras 420A, 420B, rearward navigation cameras 430A, 430B, one or more three-dimensional imaging systems 440A, 440B, one or more bin edge detection sensors 450A, 450B, one or more traffic monitoring cameras 460A, 460B, and one or more out-of-plane (e.g., upward or downward facing) positioning cameras 477A, 477B (note that the downward facing camera can supplement the line following sensor 275 of the physical property sensor system 270 and provide a wider field of view than the line following sensor 275 to implement the guidance / crossing of the vehicle 110. In the case where the vehicle path deviates from the guidance line 900 (i.e., the guidance line 900 is removed from the field of view of the line following sensor 275), place the guidance line 900 (see Figure 1B ) back into the field of view of the line following sensor 275). Images (static images and / or dynamic video images) from the cameras of the different vision systems 400 are requested by the controller 122 from the vision system controller 122VC as desired for the operation of any given autonomous guided vehicle 110. For example, the controller 122 obtains images from at least one or more of the forward navigation cameras 420A, 420B and the rearward navigation cameras 430A, 430B to implement the navigation of the autonomous guided vehicle 110 along the transfer deck 130B and the picking aisle 130A.
[0059] The forward navigation cameras 420A, 420B can be paired to form a stereo camera system, and the rearward navigation cameras 430A, 430B can be paired to form another stereo camera system. Refer to Figure 2 and 3A , the forward navigation cameras 420A, 420B are any suitable cameras configured to provide object detection and ranging. The forward navigation cameras 420A, 420B can be placed on opposite sides of the longitudinal centerline LAXCL of the autonomous transport vehicle 110 and separated by any suitable distance such that the forward-facing fields of view 420AF, 420BF provide stereo vision for the autonomous transport vehicle 110. The forward navigation cameras 420A, 420B are any suitable high-resolution or low-resolution video cameras (where video images including more than about 480 vertical scan lines and captured at more than about 50 frames per second are considered high-resolution), time-of-flight cameras, laser ranging cameras, or any other suitable cameras configured to provide object detection and ranging for implementing the traversal of the autonomous vehicle along the transfer deck 130B and the picking aisle 130A. The rearward navigation cameras 430A, 430B can be substantially similar to the forward navigation cameras. The forward navigation cameras 420A, 420B and the rearward navigation cameras 430A, 430B provide navigation for obstacle detection and avoidance of the autonomous guided vehicle 110 (where either end 200E1 of the autonomous guided vehicle 110 leads or follows the direction of travel) and the positioning of the autonomous transport vehicle within the storage and retrieval system 100. The positioning of the autonomous transport vehicle 110 can be implemented by one or more of the forward navigation cameras 420A, 420B and the rearward navigation cameras 430A, 430B by detecting the guide lines on the travel / rolling surface 284 and / or by detecting suitable storage structures (including but not limited to storage rack (or other) structures). The line detection and / or storage structure detection can be compared with the floor map and structure information of the vision system controller 122VC (e.g., stored in the memory or accessible). The forward navigation cameras 420A, 420B and the rearward navigation cameras 430A, 430B can also send signals to the controller 122 (including or through the vision system controller 122VC) such that when an object approaches the autonomous transport vehicle 110 (with the autonomous transport vehicle 110 stopped or in motion), the autonomous transport vehicle 110 can maneuver (e.g., on the uncertain rolling surface of the transfer deck 130B or within the picking aisle 130A (which can have a determined or undetermined rolling surface)) to avoid the approaching object (e.g., another autonomous guided vehicle, bin unit, or other transient object within the storage and retrieval system 100).
[0060] Forward navigation cameras 420A, 420B and rearward navigation cameras 430A, 430B may also be provided for a fleet of carriers 110 along pick aisle 130A or transfer deck 130B, where one carrier 110 follows another carrier 110A at a predetermined fixed distance. As an example, Figure 1B A fleet of three carriers 110 is shown, where one carrier closely follows another carrier at a predetermined fixed distance.
[0061] As another example, the controller 122 may obtain images from one or more of the three-dimensional imaging systems 440A, 440B, the case edge detection sensors 450A, 450B, and the case unit monitoring cameras 410A, 410B for the carrier 110 to perform case handling. Still referring to Figure 2 and 3A , one or more case edge detection sensors 450A, 450B are any suitable sensors configured to scan the shelves of the storage and retrieval system 100 to verify that the shelves are clear for placing the case unit CU, or to verify the case unit size and position before picking the case unit CU, such as a laser measurement sensor. Although one case edge detection sensor 450A, 450B is shown on each side of the centerline CLPB of the payload platform 210B (see Figure 3A ), more or fewer than two case edge detection sensors may be placed at any suitable location on the autonomous transport carrier 110 such that the carrier 110 can pass by the case unit CU and scan the case unit CU when the front end 200E1 leads the direction of travel of the carrier or the tail / rear end 200E2 leads the direction of travel of the carrier. Note that case handling includes picking and placing cases from the case unit holding position (such as verification of case unit positioning, verification of the case unit, and placement of the case unit within the payload platform 210B and / or at the case unit holding position (such as a storage shelf or a staging area position)).
[0062] Images from the off-plane positioning cameras 477A, 477B may be obtained by the controller 122 to perform navigation of the autonomous guided carrier 110 and / or provide data (e.g., image data) to supplement the positioning / navigation data from one or more of the forward navigation cameras 420A, 420B and the rearward navigation cameras 430A, 430B. Images from one or more traffic monitoring cameras 460A, 460B may be obtained by the controller 122 to perform the travel transition of the autonomous guided carrier 110 from the pick aisle 130A to the transfer deck 130B (e.g., entering the transfer deck 130B and merging the autonomous guided carrier 110 with other autonomous guided carriers traveling along the transfer deck 130B).
[0063] The one or more out-of-plane (e.g., facing up or down) positioned cameras 477A, 477B are disposed on the frame 200 of the autonomous transport vehicle 110 to sense / detect position reference points (e.g., position markings (such as barcodes, etc.), lines 900 (see Figure 1B ) etc.) disposed on the top plate of the storage and retrieval system or on the rolling surface 284 of the storage and retrieval system. The position reference points have known positions within the storage and retrieval system and can provide unique identification marks / patterns (e.g., processed data obtained from the positioning cameras 477A, 477B) that can be recognized by the vision system controller 122VC. Based on the detected position reference points, the vision system controller 122VC compares the detected position reference points with known position reference points (e.g., stored in the memory of the vision system controller 122VC or accessible by the vision system controller 122VC) to determine the position of the autonomous transport vehicle 110 within the storage structure 130.
[0064] The one or more traffic monitoring cameras 460A, 460B are disposed on the frame 200 such that the corresponding fields of view 460AF, 460BF face the side in the lateral direction LAT1. Although the one or more traffic monitoring cameras 460A, 460B are shown adjacent to the transfer opening 1199 of the transfer table 210B (e.g., on the picking side from which the arm 210A of the autonomous transport vehicle 110 extends), in other aspects, the traffic monitoring cameras can be disposed on the non-picking side of the frame 200 such that the fields of view of the traffic monitoring cameras face laterally in the LAT2 direction. The traffic monitoring cameras 460A, 460B provide for the autonomous merging of the autonomous transport vehicle 110 leaving, for example, the picking aisle 130A or the elevator transfer area 195 onto the transfer deck 130B (see Figure 1B ). For example, the autonomous transport vehicle 110V leaving the elevator transfer area 195 ( Figure 1B ) detects the autonomous transport vehicle 110T traveling along the transfer deck 130B. Here, the controller 122 autonomously formulates the merging (e.g., entering the transfer deck in front of or behind the autonomous transport vehicle 110T, accelerating onto the transfer deck based on the speed of the approaching vehicle 110T, etc.) onto the transfer deck based on the information (e.g., distance, speed, etc.) of the autonomous transport vehicle 110T collected by the traffic monitoring cameras 460A, 460B and transmitted to and processed by the vision system controller 122VC.
[0065] The case unit monitoring cameras 410A, 410B are any suitable high-resolution or low-resolution video cameras (where video images including more than about 480 vertical scan lines and captured at more than about 50 frames per second are considered high-resolution). The case unit monitoring cameras 410A, 410B are arranged relative to each other to form a stereo vision camera system configured to monitor the entry and exit of case units CU to and from the payload stage 210B. The case unit monitoring cameras 410A, 410B are coupled to the frame 200 in any suitable manner and focused at least on the payload stage 210B. In one or more aspects, the case unit monitoring cameras 410A, 410B are coupled to the transfer arm 210A so as to move along with the transfer arm 210A in the direction LAT (such as when picking and placing case units CU) and positioned so as to be focused on the payload stage 210B and the support forks 210AT of the transfer arm 210A.
[0066] Also refer to Figure 5 A, the case unit monitoring cameras 410A, 410B at least partially implement one or more of the verification of case unit determination, case unit positioning, case unit position verification, and case unit adjustment features (such as adjusting the vanes 471 and the pusher 470) and case transfer features (such as the forks 210AT, the puller 472, and the payload stage floor 473). For example, the case unit monitoring cameras 410A, 410B detect one or more of the case unit lengths CL, CL1, CL2, CL3, the case unit heights CH1, CH2, CH3, and the case unit yaw angle YW (e.g., relative to the extension / retraction direction LAT of the transfer arm 210A). Data from case handling sensors (such as those mentioned above) can also provide the location / position of the pusher 470, the puller 472, and the adjusting vanes 471, such as the position where the payload stage 210B is empty (e.g., not holding a case unit).
[0067] The case unit monitoring cameras 410A, 410B are also configured to implement the determination of the front face case center point FFCP relative to the reference position of the autonomous guided vehicle 110 (e.g., in the X, Y, and Z directions, where the case unit is set on a shelf or other holding area outside the vehicle 110) using the vision system controller 122VC. The reference position of the autonomous guided vehicle 110 can be defined by one or more adjustment surfaces of the payload stage 210B or the center line CLPB of the payload stage 210B. For example, the front face case center point FFCP can be determined relative to the center line CLPB of the payload stage 210B along the longitudinal axis LAX (e.g., in the Y direction) ( Figure 3A ). The front face case center point FFCP can be relative to the case unit support plane PSP of the payload stage 210B along the vertical axis VER (e.g., in the Z direction) ( Figure 3A and 3B, which is determined by one or more of the forks 210AT of the transfer arm 210A and the payload deck bottom plate 473). The front face box center point FFCP can be determined relative to the adjustment planar surface JPP of the pusher 470 Figure 3B ) along the lateral axis LAT (e.g., in the X direction). As a non-limiting example, determining the front face box center point FFCP of the bin unit CU located in the storage rack 555 (see Figure 3A and 4A ) or other bin unit holding positions provides positioning of the autonomous guided vehicle 110 relative to the bin unit CU to be picked, which reflects the position of the bin unit within the storage structure (e.g., in a manner similar to that described in U.S. Patent No. 9,242,800, titled "Bin Unit Detection in a Storage and Retrieval System," issued on January 26, 2016, the disclosure of which is incorporated herein by reference in its entirety), and / or the picking and placement accuracy relative to other bin units on the storage rack 555 (e.g., to maintain a predetermined gap size between bin units). The determination of the front face box center point FFCP also implements a comparison of the "real-world" environment in which the autonomous guided vehicle 110 operates with the virtual model 400VM of the operating environment, such that the controller 122 of the autonomous guided vehicle 110 compares what it will "see" substantially directly using the vision system 400 with what it expects to "see" based on the simulation of the storage and retrieval system structure, in a manner similar to that described in U.S. Patent Application No. 17 / 804,026, titled "Autonomous Transport Vehicle with Vision System," filed on May 25, 2022 (Attorney Docket No. 1127P016037-US(PAR)), the disclosure of which is incorporated herein by reference in its entirety. Additionally, in one aspect, as Figure 5 shown in A, the objects (bin units) and features determined by the vision system controller 122VC are mosaicked (combined, overlapped) with the virtual model 400VM, enhancing the resolution of the object pose relative to the facility or global reference frame GREF (up to six degrees of freedom resolution) (see Figure 2) As can be appreciated, registering the cameras of the vision system 400 with the global reference frame GREF allows for enhanced resolution of the attitude and / or position of the vehicle 110 relative to the global reference (facility features rendered in the virtual model 400VM) and the imaging object. More specifically, object position differences or anomalies become apparent and are recognized when aligning the object image with the virtual model 400VM (e.g., the edge spacing between the fiducial edges of the bin cell or the bin cell tilt or skew relative to the slat 520L of the virtual model 400VM), and if greater than a predetermined nominal threshold, describe a misalignment of one or more of the bins, racks, and / or the vehicle 110. The determination as to whether the error is related to the attitude / position of the bin, rack, or vehicle 110, one or more of which is determined by comparing with the attitude data from the sensor 270 and the supplementary navigation sensor system 288.
[0068] As an example of the above enhanced resolution, if a bin cell imaged by the vision system 400 located on a shelf is flipped relative to a juxtaposed bin cell (also imaged by the vision system) and the virtual model 400VM on the same shelf, the vision system 400 can determine that the bin is skewed and provide enhanced bin position information to the controller 122 for operating and positioning the transfer arm 210A to pick a bin based on the enhanced resolution of the attitude and position of the bin. As another example, if the edge of the bin is offset from the edge of the slat 520L (see Figures 4A - 4C ) by more than a predetermined threshold, the vision system 400 can generate a position error for the bin; note that if the offset is within the threshold, the supplementary information from the supplementary navigation sensor system 288 enhances the attitude / position resolution (e.g., the offset is substantially equal to the attitude / position of the bin relative to the slat 520L and the transfer arm 210A frame of the payload platform 210B of the vehicle 110). It is also noted that if only one bin is skewed / offset relative to the edge of the slat 520L, the vision system can generate a bin position error; however, if it is determined that two or more juxtaposed bins are skewed relative to the edge of the slat 520L, the vision system can generate a vehicle 110 attitude error and implement repositioning of the vehicle 110 (e.g., correct the position of the vehicle 110 based on the offset determined from the supplementary information of the supplementary navigation sensor system 288) or a maintenance message to the operator (e.g., where the vision system 400 implements the "dashboard camera" collaboration mode as described herein, which provides for remote control of the vehicle 110 by the operator, and the images (static images and / or real-time video) from the vision system are transmitted to the operator to perform remote control operations). The vehicle 110 can stop (e.g., not cross the picking aisle 130A or the transfer deck 130B) until the operator initiates remote control of the vehicle 110.
[0069] The case unit monitoring cameras 410A, 410B can also provide feedback on the case unit adjustment features and case transfer features of the autonomous guided vehicle 110, for example, before and / or after picking / placing a case unit from a storage shelf or other holding position (e.g., a location / position for verifying the adjustment features and case transfer features in order to perform picking / placing of the case unit with the transfer arm 210A without the transfer arm being obstructed). For example, as described above, the case unit monitoring cameras 410A, 410B have a field of view that covers the payload stage 210B. The vision system controller 122VC is configured to receive sensor data from the case unit monitoring cameras 410A, 410B and use any suitable image recognition algorithm stored in the memory of the vision system controller 122VC or accessible by the vision system controller 122VC to determine the positions of the pusher 470, adjustment vane 471, puller 472, fork 210AT, and / or any other features of the payload stage 210B that engage the case unit held on the payload stage 210B. The positions of the pusher 470, adjustment vane 471, puller 472, fork 210AT, and / or any other features of the payload stage 210B can be used by the controller 122 to verify the corresponding positions of the pusher 470, adjustment vane 471, puller 472, fork 210AT, and / or any other features of the payload stage 210B as determined by the motor encoder or other corresponding position sensors; and in some aspects, the positions determined by the vision system controller 122VC can be used as a redundancy in the event of encoder / position sensor failure.
[0070] The adjustment position of the case unit CU within the payload stage 210B can also be verified by the case unit monitoring cameras 410A, 410B. For example, also referring to Figure 3C , the vision system controller 122VC is configured to receive sensor data from the case unit monitoring cameras 410A, 410B and use any suitable image recognition algorithm stored in the memory of the vision system controller 122VC or accessible by the vision system controller 122VC to determine the reference / original position 470 of the case unit in the X, Y, Z directions relative to, for example, the centerline CLPB of the payload stage 210B, the adjustment plane surface JPP of the pusher 470 ( Figure 3B ), and one or more of the case unit support planes PSP ( Figure 3A and 3B ). Here, determining the position of the case unit CU within the payload stage 210B at least affects the placement accuracy relative to other case units on the storage shelf 555 (e.g., in order to maintain a predetermined gap size between the case units).
[0071] Referring to Figure 2 、 3A, 3B, and 5, one or more three-dimensional imaging systems 440A, 440B include any suitable three-dimensional imager, including but not limited to, for example, a time-of-flight camera, an imaging radar system, a light detection and ranging (LIDAR), etc. One or more three-dimensional imaging systems 440A, 440B provide enhanced positioning of the autonomous guided vehicle 110 relative to a global reference frame GREF (see Figure 2 ) of, for example, a storage and retrieval system 100. For example, one or more three-dimensional imaging systems 440A, 440B may utilize a vision system controller 122VC to implement invariants with respect to a reference position of the autonomous guided vehicle 110 and the shelves of the support bin unit CU to determine the size (e.g., height and width) of the front face (i.e., the front surface) of the bin unit CU and the front face bin center point FFCP (e.g., in the X, Y, and Z directions) (e.g., one or more three-dimensional imaging systems 440A, 440B implement bin unit CU positioning without referring to the shelves of the support bin unit CU (this positioning of the bin unit CU within the automated storage and retrieval system 100 is defined in the global reference frame GREF), and implement a determination as to whether the bin unit is supported on the shelves by determining shelf-invariant characteristics of the bin unit). Here, the determination of the front surface and the bin center point FFCP also implements a comparison of the "real world" environment in which the autonomous guided vehicle 110 operates with the virtual model 400VM, such that the controller 122 of the autonomous guided vehicle 110 compares what the vision system 400 can essentially directly "see" with what the autonomous guided vehicle 110 expects to "see" based on a simulation of the structure of the storage and retrieval system, as described in U.S. Patent Application No. 17 / 804,026, filed on May 25, 2022, and titled "Autonomous Transport Vehicle with Vision System" (Attorney Docket No. 1127P016037-US(PAR)), the disclosure of which is incorporated herein by reference in its entirety. In the case where the data from the cameras 410A, 410B is incomplete or missing, the image data obtained from one or more three-dimensional imaging systems 440A, 440B may supplement and / or enhance the image data from the cameras 410A, 410B. Here, object detection and positioning with respect to the pose of the autonomous guided vehicle 110 within the global reference frame GREF can be determined with high accuracy and confidence by one or more three-dimensional imaging systems 440A, 440B; however, in other aspects, this object detection and positioning may be implemented using one or more sensors of the physical property sensor system 270 of the autonomous guided vehicle 110 and / or wheel encoders / inertial sensors.
[0072] As Figure 5As shown, one or more three-dimensional imaging systems 440A, 440B have respective fields of view that extend substantially along the direction LAT through the payload stage 210B, such that each three-dimensional imaging system 440A, 440B is arranged to sense bin units CU adjacent to but outside the payload stage 210B (such as bin units CU arranged to extend in a row or rows along the length of a substrate staging area / transfer station along the picking aisle 130A (see Figure 5 A) or along the transfer deck 130B, which is structurally similar to the storage racks 599 and their shelves 555 arranged along the picking aisle 130A). The fields of view 440AF, 440BF of each three-dimensional imaging system 440A, 440B cover spatial volumes 440AV, 440BV that extend from the height 670 of the picking range of the autonomous guided vehicle 110 (e.g., the range / height in the direction VER ( Figure 2 ), where the arm 210A can move to pick / place bin units to shelves or stacked shelves accessible from the common rolling surface 284 on which the autonomous guided vehicle 110 is located (see Figure 2 , e.g., the transfer deck 130B or the picking aisle 130A).
[0073] The vision system 400 also implements the operation control of the autonomous transport vehicle 110 in cooperation with the operator. The vision system 400 provides data (images) and the vision system data is recorded by the vision system controller 122VC, which (a) determines information characteristics (and then provides them to the controller 122), or (b) the information is transferred to the controller 122 without being characterized (objects in the predetermined criteria) and the characterization is done by the controller 122. Regardless of (a) or (b), it is the controller 122 that determines the choice to switch to the cooperative state. After the switch, the cooperative operation is implemented by the user accessing the vision system 400 via the vision system controller 122VC and / or the controller 122 through the user interface UI (see Figure 10 ). However, in its simplest form, the vision system 400 can be regarded as providing a cooperative operation mode for the autonomous transport vehicle 110. Here, the vision system 400 supplements the autonomous navigation / operation sensor system 270 to implement the cooperative discrimination and mitigation of objects / hazards 299 (see Figure 3A , where such objects / hazards include fluids, bins, solid debris, etc.) (e.g., intruding into the travel / rolling surface 284), as described in U.S. Patent Application No. 17 / 804026, filed on May 25, 2022, and titled "Autonomous Transport Vehicle with Vision System" (Attorney Docket No. 1127P016037-US(PAR)), the disclosure of which is incorporated herein by reference in its entirety.
[0074] On the one hand, an operator can select or switch the control of the autonomous guided vehicle (e.g., via a user interface UI) from autonomous operation to cooperative operation (e.g., the operator remotely controls the operation of the autonomous transport vehicle 110 via the user interface UI). For example, the user interface UI can include a capacitive touchpad / screen, a joystick, a haptic screen, or other input devices that transmit motion direction commands (e.g., steering, accelerating, decelerating, etc.) from the user interface UI to the autonomous transport vehicle 110 to implement operator control input in the cooperative operation mode of the autonomous transport vehicle 110. For example, the vision system 400 provides a "dashboard camera" (or dashcam) that transmits video and / or still images from the autonomous transport vehicle 110 to the operator (via the user interface UI) to allow remote operation or monitoring of the area relative to the autonomous transport vehicle 110 in a manner similar to that described in U.S. Patent Application No. 17 / 804,026, filed on May 25, 2022, and titled "Autonomous Transport Vehicle with Vision System" (Attorney Docket No. 1127P016037-US(PAR)), the disclosure of which is incorporated herein by reference in its entirety.
[0075] Reference Figure 10, in the cooperative operation mode of the autonomous guided vehicle 110, frames from the image data stream 1000 are presented to the operator. In one or more aspects, the frames from the image data stream 1000 are augmented images (as described herein), where the image augmentation provides at least to the operator the identification of objects within the corresponding frames. The user interface may provide a frame segmentation object selection 1000C, where a portion of a frame from the image data stream 1000 is selected for manipulation by the operator. The frame segmentation object selection 1000C may include one or more object windows 1010, 1020, 1030, where these object windows 1010, 1020, 1030 are configured to provide to the operator the selection for implementing object selection and displaying data related to the selected object. For example, the object selector 1010 may be presented to the operator through the display of the frame segmentation selection 1000C. The object selector 1010 may include a drop-down menu (or other suitable interface) that implements the operator's selection of the objects shown in the frame segmentation selection 1000C (e.g., in this example, the forks and pusher / puller of the vehicle 110 and the bin unit CU are shown and presented in the object selector for the operator to select). With the selection of an object, a bounding box may be presented around the selected object to identify the object in the frame segmentation selection 1000C. The selection of an object may also implement the presentation of object information 1020 (e.g., in this example, the "part" object is selected and the information presented for the "part" object informs the operator that the "part" object represents or otherwise identifies a portion of the bin unit CU recognized by the vision system controller 122VC). An action list 1030 may also be presented in the user interface UI, where the action list depends on the selected object. For example, as the bin unit CU that has been partially identified, the operator may indicate through the action list 1030 whether the autonomous guided vehicle picks or does not pick the bin unit CU. In other aspects, the user interface UI may present any suitable information regarding the operation of the autonomous guided vehicle 110 to the operator implementing the cooperative operation of the autonomous guided vehicle 110.
[0076] Reference Figure 1A, as described above, the autonomous guided vehicle 110 is provided with a vision system 400 having an architecture based on camera pairs (e.g., such as camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, 477A and 477B), configured for stereo or binocular object detection and depth determination (e.g., by using the respective disparities / depth maps from the recorded video frames / images captured using the respective cameras). Object detection and depth determination provide the positioning of the autonomous guided vehicle 110 relative to objects (e.g., at least bin holding positions, e.g., such as on shelves and / or lifts, and bins to be picked); however, as mentioned herein, stereo vision from the camera pairs may not always be available (i.e., stereo vision deficit). Thus, the disclosed embodiments provide to the controller 122 of the autonomous guided vehicle 110 both a computer vision object detection and localization protocol and a machine learning detection and localization protocol that use the stereo vision of the vision system 400. The computer vision object detection and localization protocol uses the vision system controller 120VC to determine at least one computer vision parameter difference or disparity map for different features on the image frames obtained using the stereo cameras. The machine learning detection and localization protocol uses machine learning, wherein at least one machine learning model ML and at least one artificial neural network ANN perform robust object determination and localization through monocular vision (using one (i.e., non-deficient) camera in the stereo camera pair) with a detection confidence comparable to that of non-deficient binocular vision.
[0077] As described above and also referring to Figure 2 , the artificial neural network ANN is included in the depth guidance module DC. The depth guidance module DC is communicatively coupled to the vision controller 122VC. The depth guidance module DC (which may be referred to as a deep learning graphics processing unit) is located onboard the autonomous guided vehicle 110 and is a separate control module for the controller 122. Here, the shared memory SH is located onboard the autonomous guided vehicle 110. The shared memory SH is communicatively connected to the vision system 400 and the vision system controller 122VC to receive video stream data imaging of objects in the logistics space (e.g., provided substantially in real time via the dashcam operation mode of the vision system 400 as described herein and / or provided by the cache operation mode in which the video is stored in the shared memory for retrieval as needed). The shared memory SH includes a configuration to generate image frames from the video stream data imaging (e.g., such as Figure 9A and Figure 9CNon - transient image frame generation computer program code (such as those shown in
[0078] ) that transmits the image frame to the depth guidance module DC to implement the selection of one or the computer vision (detection / localization) protocol and the machine learning (detection / localization) protocol. In other aspects, the non - transient image frame generation computer program code is included in the depth guidance module DC. The non - transient image frame generation computer program code can be, for example, an OpenCV Mat object / image processing code or any other suitable code for generating image frames from video stream data imaging. Figure 1A ) or other remotely located computers / servers) and is communicatively connected to the autonomous guided vehicle 110. In the case where the depth guidance DC is located away from the autonomous guided vehicle 110, the server / client socket - based application is configured such that the media server MS is located on board the autonomous guided vehicle 110. The media server MS is communicatively connected to the vision system 400 and the vision system controller 122VC to receive video stream data imaging of objects in the logistics space from the vision system 400 (such as provided via the dashcam operation mode of the vision system 400 described herein and provided substantially in real - time, and / or provided by the cache operation mode in which video is stored in a shared memory for retrieval as needed). Here, the media server MS is communicatively connected to at least one camera (such as those of the autonomous guided vehicle 110 described herein) and records video stream data from the at least one camera (in any suitable memory). Although the media server is described as docking with the vision system controller 122VC (for example, the vision system controller 122VC is located on board the autonomous guided vehicle 110), the media server MS can also (for example, via the network 180) dock with any suitable controller located away from the autonomous guided vehicle 110 (such as the control server 120 or the warehouse management system 2500). The media server MS includes non - transient image frame generation computer program code to generate image frames from video stream data imaging (for example, such as Figure 9A and Figure 9C ) that transmits the image frame to the depth guidance module DC to implement the selection of one or the computer vision (detection / localization) protocol and the machine learning (detection / localization) protocol. The media server MS is configured to generate image frames in any suitable manner (such as using OpenCV Mat object / image processing) as described above.
[0079] Both the shared memory SH and the media server MS are configured to remove image frames that are similar to each other in the case of imaging the video stream data into image frames. When analyzing the received image frames, the removal of similar (e.g., substantially replicated) image frames reduces the image processing time of the depth director DC and reduces the data transfer traffic between the shared memory SH and the depth director DC and between the media server MS and the depth director DC. The shared memory SH and the media server MS are configured via respective non-transitory image frame generation computer program codes to use any suitable similarity metric (e.g., structural similarity index) that calculates the structural similarity of the image frames. A suitable example of the structural similarity index for duplicate image frame removal can be found, for example, in "Image Quality Assessment: From Error Visibility to Structural Similarity" by Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli in April 2004, IEEE Transactions on Image Processing, Vol. 13, No. 4, pages 600 - 612, (referred to herein as "Wang"), the disclosure of which is incorporated herein by reference in its entirety. This similarity metric is used by the respective non-transitory image frame generation computer program codes of the shared memory SH and the media server MS to remove similar images based on a predetermined similarity index threshold (such as that described in Wang). A predetermined similarity index threshold is set for each of the cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B so as to generate a similarity index and remove similar images when converting the video files from each respective camera into the image files for each respective camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B.
[0080] In other aspects, the video stream data imaging is streamed in any suitable manner (such as via dashcam operation) to a remotely located depth guidance module DC (located in the remotely located controller as described herein), where the depth guidance module DC includes, for example, an OpenCV Mat object / image processing for generating image frames.
[0081] Note that the autonomous guided vehicle can be provided with both an on-board shared memory SH having an on-board depth guidance module DC and a media server MS having a remotely located depth guidance module DC, where the media server MS and the remotely located depth guidance module DC can be used in cases where it is desired to conserve or limit the processing power on-board the autonomous guided vehicle 110 and / or the power stored on the autonomous guided vehicle 110. As described herein, the media server MS can also interface with the on-board shared memory SH having an on-board depth guidance module DC.
[0082] Still referring to Figure 1A 、Figure 2 and Figure 6 , in order to obtain video stream data imaging using the vision system 400, calibrate the stereo camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, 477A and 477B (see Figure 6 , frame 600). The stereo camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, 477A and 477B can be calibrated in any suitable manner (such as by, for example, intrinsic and extrinsic camera calibration) to implement sensing of the bin unit CU, storage structures (such as shelves, columns, etc.), and other structural features of the storage and retrieval system. Also refer to Figure 3A , Figure 3B , Figure 4A , Figure 4B and Figure 4C , known objects (such as bin units CU1, CU2, CU3 (or storage system structures) (for example, having known physical characteristics such as shape, size, etc.)) can be placed within the field of view of the cameras of the supplementary navigation sensor system 288 (or the carrier 110 can be positioned such that the known object is within the field of view of the cameras). These known objects can be imaged by the cameras from several angles / viewpoints to calibrate each camera such that the vision system controller 122VC is configured to detect the known object based on the sensor signals from the calibrated cameras.
[0083] For example, describe the calibration of the bin unit monitoring cameras 410A, 410B with respect to the bin units CU1, CU2, CU3 having known physical characteristics / parameters (note that the calibration of the other stereo camera pairs described herein can be implemented in a similar manner). Figures 4A to 4C are exemplary images captured from (for illustrative purposes) three different viewpoints from one of the bin unit monitoring cameras 410A, 410B. Here, the physical characteristics / parameters of the bin units CU1, CU2, CU3 (such as shape, length, width, height, etc.) are known to the vision system controller 122VC (for example, the physical characteristics of the different bin units CU1, CU2, CU3 are stored in the memory of the vision system controller 122VC or are accessible by the vision system controller 122VC). Based on, for example Figures 4A to 4CFor three (or more) different viewpoints of the bin units CU1, CU2, CU3 in the image, the vision system controller 122VC is provided with intrinsic and extrinsic cameras and bin unit parameters for performing calibration of the bin unit monitoring cameras 410A, 410B. It should be noted that each camera 410A, 410B is essentially calibrated to its own coordinate system (i.e., each camera obtains the depth of the object from the corresponding image sensor of the camera). In the case of using multiple cameras 410A, 410B, on the one hand, the calibration of the vision system 400 includes calibrating the cameras 410A, 410B to a common basic reference system (which can be the reference system of a single camera, or any other suitable basic reference system to which each of the cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B can be related so as to form a common basic reference system for all the cameras of the vision system) and calibrating the common basic reference system to the robot reference system. The calibration of the cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B to the common basic reference system includes identifying and applying a transformation between the respective reference systems of each of the cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B such that the respective reference systems of each of the cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B are transformed (or referenced) to the common basic reference system. For example, the reference system of camera 410A (although any camera can be used) represents the common basic reference system. The transformation (i.e., a rigid transformation in six degrees of freedom) is determined for each of the coordinate systems / reference systems of the other cameras 410B, 420A, 420B, 430A, 430B, 460A, 460B relative to the reference system of camera 410A such that the image data from the cameras 410B, 420A, 420B, 430A, 430B, 460A, 460B is associated or transformed into the reference system of camera 410A.
[0084] The calibration of the cameras includes recording (e.g., storing in a memory) by the vision system controller 122VC the viewpoints of the bin units CU1, CU2, CU3 relative to, for example, the bin unit monitoring cameras 410A, 410B. The vision system controller 122VC estimates the poses of the bin units CU1, CU2, CU3 relative to the bin unit monitoring cameras 410A, 410B, and estimates the poses of the bin units CU1, CU2, CU3 relative to each other. The pose estimates PE of the respective bin units CU1, CU2, CU3 are shown Figures 4A to 4C overlapped on the respective bin units CU1, CU2, CU3.
[0085] The carrier 110 moves such that any suitable number of viewpoints of the bin units CU1, CU2, CU3 are acquired / imaged by the bin unit monitoring cameras 410A, 410B to effect convergence of bin unit characteristics / parameters for each of the known bin units CU1, CU2, CU3 (e.g., estimated by the vision system controller 122VC). When the bin unit parameters converge, the bin unit monitoring cameras 410A, 410B are calibrated. The calibration process is repeated for another pair of bin unit monitoring cameras 410A, 410B. With both of the bin unit monitoring cameras 410A, 410B calibrated, the vision system controller 122VC is configured with three-dimensional rays for each pixel within each of the bin unit monitoring cameras 410A, 410B, and an estimate of the three-dimensional baseline segment separating the cameras, and the relative pose of the bin unit monitoring cameras 410A, 410B with respect to each other. The vision system controller 122VC is configured to use the three-dimensional rays for each pixel of each of the bin unit monitoring cameras 410A, 410B, the estimate of the three-dimensional baseline segment separating the cameras, and the relative pose of the bin unit monitoring cameras 410A, 410B with respect to each other such that the bin unit monitoring cameras 410A, 410B form a passive stereo vision sensor, such as when a common feature is visible within the fields of view 410AF, 410BF of the bin unit monitoring cameras 410A, 410B.
[0086] The common base reference frame can be transformed to the reference frame BREF of the autonomous guided vehicle 110 by transporting one or more of the bin units CU1 - CU3 to the payload platform 210B of the autonomous guided vehicle 110, where the one or more bin units CU1 - CU3 are adjusted within the payload platform 210B (e.g., using at least the adjustment vanes 471 and the pusher 470). With one or more bin units CU1 - CU3 at known positions within the payload platform 210B, the controller (knowing the dimensions of the bin units CU1 - CU3) characterizes the relationship between the image field of the common base reference frame and the reference frame BREF of the robot such that the positions of the bin units CU1 - CU3 in the vision system image are calibrated to the robot reference frame BREF. Other stereo camera pairs 420A and 420B, 430A and 430B, 477A and 477B can be calibrated in a similar manner, where the common base reference frame of each pair is transformed to the reference frame BREF of the autonomous guided vehicle 110 based on known relative camera positions and / or aberration maps or depth maps of parts of the autonomous guided vehicle (within the fields of view of the respective camera pairs) imaged and stored and retrieved from the bin units or other structures of the imaging and storage and retrieval system.
[0087] As can be appreciated, also refer to Figure 8, to transform the common basic reference frame of the corresponding camera pair to the reference frame BREF of the autonomous guided vehicle 110, the computer model 800 (such as a computer-aided design or CAD model) of the autonomous guided vehicle 110 can also be used by the controller 122 (or the vision controller 122VC) (either in combination with or instead of determining the reference frame transformation implemented by holding the bin unit in the payload compartment as described above). As Figure 8 can be seen, the controller 122 can extract the characteristic dimensions of the part of the autonomous guided vehicle 110 within the field of view of the camera pair, such as any suitable feature of the payload stage 210B depending on which pair of camera pairs is calibrated (in this example, it is a feature of the payload stage fence relative to the reference frame BREF, or any other suitable feature of the autonomous guided vehicle 110 via the autonomous guided vehicle model 800, and / or a suitable feature of the storage structure of the virtual model 400VM of the operating environment). These characteristic dimensions of the payload stage 210B are determined from the origin of the reference frame BREF of the autonomous guided vehicle 110. The controller 122 (or the vision controller 122VC) uses these known dimensions of the autonomous guided vehicle 110 together with the aberration map or depth map generated by the stereo camera pair to associate the common basic reference frame (or the reference frame of each camera) with the reference frame BREF of the autonomous guided vehicle 110.
[0088] As mentioned above, the calibration of the bin unit monitoring cameras 410A and 410B is described relative to the bin units CU1, CU2, and CU3, but it can be performed in a substantially similar manner for any suitable structure of the storage and retrieval system 100 (e.g., permanent or temporary (including calibration fixtures)).
[0089] Also refer to Figure 7A and Figure 7B , examples of calibrating multiple cameras using the calibration fixture / fixture 700 (also known as the common camera calibration reference structure) include: calibrating the camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A, 460B, 477A, 477B to the corresponding common basic reference frame (in a manner similar to that described above) and calibrating the common basic reference frame to the robot reference frame (in a manner similar to that described above); while in other aspects of camera calibration using bin units / storage structures and / or calibration fixtures, in the case of using one or more cameras, the reference frame of one camera or the reference frames of one or more cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B can be individually calibrated to the reference frame BREF of the autonomous guided vehicle 110.
[0090] For illustrative purposes only, refer toFigure 3A , Figure 3B , Figure 7A and Figure 7B , calibrating the cameras for 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, 477A and 477B to a corresponding common basic reference frame includes identifying and applying the transformation between the corresponding reference frames of each of the cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B such that the corresponding reference frame transformation of each of the cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B is transformed (referenced) to the corresponding common basic reference frame. Similarly, the common basic reference frame can be the reference frame of a single camera in the camera pair, or any other suitable basic reference frame to which each of the camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, 477A and 477B can be related in a manner substantially similar to that described above so as to form the corresponding common basic reference frame for all of the cameras in the camera pair. This calibration of the cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B can be performed using a calibration fixture 700 placed at any suitable holding location of the storage and retrieval system 100 (e.g., on a storage shelf or any other suitable location accessible to the autonomous guided vehicle 110 and within the field of view of the calibrated camera). As described herein, calibrating each of the cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B to a common basic reference frame and to a common basic reference frame describes the positional relationship of the corresponding camera reference frames of each of the cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B with each other camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B and with the reference frame BREF of the autonomous guided vehicle 110.
[0091] The calibration fixture 700 includes unique recognizable three-dimensional geometries 710 to 719 (in this example, squares, some of which are rotated relative to others), which provide an asymmetric pattern for the calibration fixture 1300 and limit the determination / transformation of the reference frame of the camera (e.g., from each camera) to a common basic reference frame and the transformation between the common basic reference frame and the autonomous guided vehicle reference frame BREF, as will be further described, in order to determine the relative pose of the calibration fixture 1300 (and thus the bin unit) relative to the transfer arm 210A of the autonomous guided vehicle 110. The calibration fixture 700 shown and described herein is exemplary, and any other suitable calibration fixture can be used in a similar manner as described herein. For exemplary purposes, each of the three-dimensional geometries 710 to 719 has a predetermined size that limits the identification of the corners or points C1 - C36 of the three-dimensional geometries 710 to 719, and is transformed to minimize the distance between the corresponding corners C1 - C36 (e.g., the distance between the corresponding corners C1 - C36 in the reference frame of camera 310C1 is minimized relative to each of the corresponding corners C1 - C36 identified in the reference frames 410AF, 410BF, 420AF, 420BF, 430AF, 430BF, 460AF, 460BF of the calibrated camera pairs of each camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, and cameras 477A, 477B have similar reference frames).
[0092] Each of the three-dimensional geometries 710 to 719 is imaged simultaneously (i.e., each of the three-dimensional geometries 710 to 719 is at a single position in the common basic reference frame during the imaging of all cameras of the camera pairs whose reference frames will be calibrated to the corresponding common basic reference frame), and is uniquely identified at this single position by each camera in the camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, 477A and 477B, such that the points / corners C1 - C36 of the three-dimensional geometries 1310 to 1319 identified in the image ( Figure 7B the exemplary image shown) are identified by the vision system 400 and determined uniquely independent of the calibration fixture orientation. The corners C1 - C36 identified in each image of the image set are compared between the images from the cameras in the camera pairs to define the transformation of each camera reference frame to the common basic reference frame (in one example, which may correspond to or be defined by the reference frame of camera 410A for calibrating the camera pair 410A, 410B).
[0093] When registering cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B to their respective common basic reference frames, by transferring the calibration fixture 700 (or a similar fixture) to the payload stage 210B of the autonomous guided vehicle 110 and / or by converting (e.g., recording) the respective common basic reference frames (or the reference frames of one or more cameras individually) to the autonomous guided vehicle 110 reference frame BREF in a manner similar to that mentioned above using the computer model 800 of the autonomous guided vehicle 110. For example, in the case where the calibration fixture 700 is transferred to the payload stage 210B, the calibration fixture 700 is adjusted within the payload stage 210B (e.g., using at least the adjustment vanes 471 and the pusher 470). In the case where the calibration fixture 700 is at a known (i.e., adjusted) position within the payload stage 210B, the controller (knowing the positions of the points / corners C1 - C36) characterizes the relationship between the image field of the common basic reference frame and the reference frame BREF of the autonomous guided vehicle such that the positions of the points / corners C1 - C36 in the vision system image are calibrated to the robot reference frame BREF. In the case of using the computer model 800, the aberration map or depth map generated by the stereo camera pair, together with the known dimensions of the payload stage features (from the computer model 800) and the known dimensions / positions of the points / corners C1 - C36, are used to provide, for example, the conversion between the respective common basic reference frames of the camera pairs 410A and 410B and the reference frame BREF of the autonomous guided vehicle 110.
[0094] Also refer to Figure 15 , calibration of the camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, 477A and 477B can be provided or otherwise performed at the calibration station 1510 of the storage structure 130. As Figure 15As can be seen, the calibration station 1510 can be disposed at or adjacent to the autonomous guided vehicle entry or exit position 1590 of the storage structure 130. The autonomous guided vehicle entry or exit position 1590 provides for the introduction and removal of the autonomous guided vehicle 110 to and from one or more storage levels 130L of the storage structure 130 in a manner that is substantially similar to that described in U.S. Patent No. 9,656,803, entitled "Storage and Retrieval System Rover Interface," issued on May 23, 2017, the disclosure of which is incorporated herein by reference in its entirety. For example, the autonomous guided vehicle entry or exit position 1590 includes a lift module 1591 such that access and egress of the autonomous guided vehicle 110 can be provided at each storage level 130L of the storage structure 130. The lift module 1591 can dock with the transfer deck 130B of one or more storage levels 130L. The interface between the lift module 1591 and the transfer deck 130B can be disposed at a predetermined location on the transfer deck 130B such that the access and egress of the autonomous guided vehicle 110 to each transfer deck 130B is substantially decoupled from the throughput of the automated storage and retrieval system 100 (e.g., the input and output of the autonomous guided vehicle 110 at each transfer deck do not affect the throughput). In one aspect, the lift module 1591 can dock with a branch or staging area 130B1 to 130Bn (e.g., an autonomous guided vehicle load platform) that is connected to or forms part of the transfer deck 130B of each storage level 130L. In other aspects, the lift module 1591 can dock substantially directly with the transfer deck 130B. It should be noted that the transfer deck 130B and / or the staging areas 130B1 to 130Bn can include any suitable barrier 1520 that substantially prevents the autonomous guided vehicle 110 from moving away from the transfer deck 130B and / or the staging areas 130B1 to 130Bn at the lift module interface. In one aspect, the barrier can be a movable barrier 1520 that can move between a deployed position for substantially preventing the autonomous guided vehicle 110 from moving away from the transfer deck 130B and / or the staging areas 130B1 to 130Bn and a retracted position for allowing the autonomous guided vehicle 110 to pass between the lift platform 1592 of the lift module 1591 and the transfer deck 130B and / or the staging areas 130B1 to 130Bn. In addition to inputting and removing the autonomous guided vehicle 110 to and from the storage structure 130, in one aspect, the lift module 1591 can also convey the rover 110 between the storage levels 130L without removing the autonomous guided vehicle 110 from the storage structure 130.
[0095] Each of the staging areas 130B1 through 130Bn includes a respective calibration station 1510 configured to enable the autonomous guided vehicle 110 to repeatedly calibrate the camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, 477A and 477B. Calibration of the camera pairs can be automated after the autonomous guided vehicle is registered (in a manner substantially similar to that described in U.S. Patent No. 9,656,803, which is hereby incorporated by reference in its entirety) to the storage structure 130 via the autonomous guided vehicle entry or exit location 1590. In other aspects, calibration of the camera pairs can be manual (such as in the case where the calibration station is located on the elevator 1592) and performed in a manner similar to that described herein with respect to the calibration station 1510 prior to inserting the autonomous guided vehicle 110 into the storage structure 130.
[0096] To calibrate the stereo camera pair, the autonomous guided vehicle is positioned (manually or automatically) at a predetermined location of the calibration station 1510. Automatically positioning the autonomous guided vehicle 110 at a predetermined location may use the vision system 400 of the autonomous guided vehicle 110 to detect any suitable feature of the calibration station 1510. For example, the calibration station 1510 includes any suitable positioning flag or location 1510S disposed on one or more surfaces 1200 of the calibration station 1510. The positioning flag 1510S is disposed on one or more surfaces within the field of view of at least one camera of the respective camera pair. The vision system controller 122VC is configured to detect the positioning flag 1510S and, using the detection of one or more of the positioning flags 1510S, roughly position the autonomous guided vehicle relative to the calibration fixture 700 (e.g., a shelf or other support on the calibration station 1510), the calibration bin unit (similar to the bin units CU1, CU2, CU3 mentioned above and stored on the shelf of the calibration station 1510), and / or other calibration references (or known objects) such as those described herein. In other aspects, in addition to or instead of the positioning flag 1510S, the calibration station 1510 may also include a bumper or physical stop against which the autonomous guided vehicle 110 abuts to position itself at a predetermined location of the calibration station 1510. For example, the bumper or physical stop may be a barrier 1520 or any other suitable static or deployable feature of the calibration station. The automatic positioning of the autonomous guided vehicle 110 within the calibration station 1510 may be implemented when the autonomous guided vehicle 110 is introduced into the storage and retrieval system 100 (such as when the autonomous guided vehicle exits the elevator 1592) and / or at any suitable time when the autonomous guided vehicle enters the calibration station 1510 from the transfer deck 130. Here, the autonomous guided vehicle 110 may be programmed with calibration instructions that perform stereo vision calibration when introduced into the storage structure 130, or the calibration instructions may be initialized at any suitable time when the autonomous guided vehicle 110 is operating (i.e., in service) within the storage structure 130.
[0097] As mentioned above, case units CU1, CU2, CU3 and / or calibration fixtures 700 may be stored on storage shelves of corresponding calibration stations 1510, wherein calibration of the camera pairs is performed at the corresponding calibration stations 1510 in the manner described above. In addition, one or more surfaces of each calibration station 1110 may include any suitable number of known objects GDT, which may be substantially similar to geometric shapes 710 to 719. The one or more surfaces may be any surface that can be viewed by the camera pair, including but not limited to the side walls 1511 of the calibration station 1510, the top plate 1512 of the calibration station 1510, the bottom plate / through surface 1515 of the calibration station 1510, and the barrier 1520 of the calibration station 1510. The object GDT comprising a corresponding surface (which may also be referred to as a visual reference or calibration object) may be a raised structure, a pore, an application (e.g., paint, sticker, etc.), each having known physical properties (such as shape, size, etc.), so that calibration of the camera pair is performed in a manner substantially similar to that described above with reference to the box units CU1-CU3 and / or the calibration fixture 700.
[0098] As can be appreciated, carrier positioning (e.g., positioning of a carrier at a predetermined position along a picking aisle 130A or along a transfer deck 130B relative to a pick / place location) performed by the physical property sensor system 270 can be enhanced using pixel-level position determination performed by the supplemental navigation sensor system 288. Here, the controller 122 is configured to position (which can be referred to as "rough positioning") the carrier 110 relative to the pick / place location by using one or more sensors of the physical property sensor system 270. The controller 122 is configured to use the supplemental (e.g., pixel-level) position information obtained from the vision system controller 122VC of the supplemental navigation sensor system 288 to adjust (which can be referred to as "fine-tuning") the carrier posture and position relative to the pick / place location, so that the positioning of the carrier 110 and the case unit CU placed by the carrier 110 to the storage location 130S can be maintained within a smaller tolerance (i.e., increased position accuracy) than the positioning of the carrier 110 or the case unit CU using the physical property sensor system 270 alone. Here, the pixel-level positioning provided by supplemental navigation sensor system 288 has a higher positioning clarity / resolution than the electromagnetic sensor resolution provided by physical property sensor system 270 .
[0099] Still refer to Figure 1A , Figure 2 and Figure 6 as well as Figures 9A to 9C In order to obtain video stream data imaging using the vision system 400, each camera in the stereo camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, 477A and 477B is calibrated for three-dimensional monocular vision (see Figure 6, frame 605). Here, for the monocular calibration of each camera in camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, 477A and 477B, aberration images or depth images and depth images associated with the three-dimensional imaging sensors are used (for stereo camera pairs 410A and 410B, three-dimensional imaging sensors 440A and / or 440B, note that at least one three-dimensional imaging system 440C, 440D can be placed at each end 200E1, 200E2 of the autonomous guided vehicle 110 relative to stereo camera pairs 420A and 420B, 430A and 430B, or the aberration maps or depth maps can be generated by the camera pairs themselves, because there is generally no obstruction between the corresponding camera pairs 420A and 420B, 430A and 430B at the ends 200E1, 200E2 of the autonomous guided vehicle 110 that would damage the depth maps generated therefrom). Figure 9A Shows a three-dimensional object (i.e., the bin unit CU) held, for example, in the payload bay 210B of the autonomous guided vehicle 110, but the three-dimensional object can be held at any suitable position within the field of view of each camera in the camera pair and the associated three-dimensional sensor. The bin unit CU is placed in a position and orientation visible in the fields of view of cameras 410A, 410B and three-dimensional imaging sensors 440A, 440B (e.g., within the payload bay 210B or any other suitable position). The bin unit CU can be observed in all images from cameras 410A, 410B and three-dimensional image sensors 440A, 440B, such that objects and points common to the images from cameras 410A, 410B and three-dimensional image sensors 440A, 440B can be incorporated for the transformation from the two-dimensional images from the corresponding monocular cameras 410A, 410B to the three-dimensional images for the corresponding monocular cameras 410A, 410B.
[0100] Figure 9C Shows an aberration map or a depth map generated by the vision system 400 (e.g., using cameras 410A, 410B). The aberration map of the images from cameras 410A, 410B is filtered, and then a clustering method is applied to the filtered aberration map to obtain an aberration map for object detection and localization. The aberration maps generated by cameras 410A, 410B are calibrated with depth maps from one or more of the three-dimensional image sensors 440A, 440B in any suitable manner.
[0101] Exemplary transformation equations for implementing depth determination from monocular images of monocular cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B are as follows:
[0102]
[0103]
[0104] Where R is a rotation transformation, r is a 3×3 transformation matrix, and t is a 3×1 transformation vector. The controller 122 and / or the vision controller 122VC use the above exemplary equations to convert coordinates within the monocular camera image into the global reference frame GREF of the storage and retrieval system 100. Depth information associated with the depth of an object detected in the forward-facing plane of a point cloud or map generated using the three-dimensional imaging sensors 440A, 440B is used to generate the coordinates of the global reference frame GREF. The above equations and depth information form calibration parameters, which are used by the controller 122 to convert monocular image coordinates into global reference frame coordinates GREF after deep learning detection analysis of the monocular image as described herein.
[0105] In the case where the stereo camera pair is calibrated and each of the individual cameras in the camera pair is calibrated as a monocular camera, the artificial neural network ANN ( Figure 6 , box 610) of the depth director DC is trained. The artificial neural network ANN is any suitable deep learning graphics processing algorithm configured for feature or representation learning and detection integration, and is trained in any suitable manner known in the art (such as using labeled or unlabeled data or in any other suitable manner). Briefly referring to Figure 12A , the detection integration performed by the depth director DC via the artificial neural network ANN generates a number of metrics, including but not limited to the degree of overlap (e.g., the overlapping range of two boxes) and the confidence of the detection, where for illustrative purposes only, the best-fit bounding box is selected to represent the detected object among the multiple bounding boxes generated for the detection of the object. In one aspect, as Figure 12A shows, the integration result for detecting an object is in the form of a bounding box with a confidence and the detected label presented as an overlap in the image frame. In Figure 12A , the adjustment vane or arm 471 is labeled and identified with a corresponding bounding box, and the confidence that the detected object is the adjustment vane 471 is 0.96 or 96%. Also in Figure 12A , the bin or box is labeled and identified with a corresponding bounding box, and the confidence that the detected object is the bin CU is 0.67 or 67%. As mentioned herein, although bounding boxes are shown in Figure 12A , the object identified in the image frame can be identified in any suitable manner (e.g., such as by highlighting the detected object or any other suitable image augmentation).
[0106] The object detection performed by the depth director DC is a multiple object detection based on deep learning in which multiple detections can be presented in each image frame. In the form of an augmented image frame (such as Figure 12A shows, note that although Figure 12AShows an image frame from a monocular video stream, but the detection results for augmented image frames (similar for image frames obtained from a stereoscopic vision video stream) are timestamped and can be saved in any suitable memory (e.g., such as shared memory SH, media server MS, or any other memory communicating with controller 122) for data recording, reporting, archiving, evaluation, troubleshooting, operational factor analysis, quality control and metrology, fault detection, or any other suitable purpose. As will be described herein, the output of the depth director DC (e.g., the detection results mentioned above with integrated logic and timestamp) is used to initiate one or more of a computer vision protocol for object detection and localization and a machine learning protocol for object detection and localization. As a brief example and with reference to 11F, the detection results of the depth director DC are shown, where the detections from cameras 410A, 410B in payload station 210B identify an object (e.g., at least one picking bin identified in the image data from each of cameras 410A, 410B), and the identified object is compared to any suitable predetermined threshold. For example, for a picking operation of autonomous mobile vehicle 110, the depth director DC compares the image data to verify that the number of identified picking bins is equal to or greater than one. The depth director DC can also verify whether there is motion (i.e., the autonomous mobile vehicle is traversing a surface at a speed that causes motion blur), in Figure 11F which case there is no blur. Based on the results of the image comparison, the depth director DC can initiate a computer vision protocol where the number of identified picking bins in the image data from camera 410A is the same as the number of identified picking bins in the image data from camera 410B, otherwise a machine learning protocol can be initiated by the depth director DC.
[0107] At least one machine learning model ML( Figure 6, frame 615) is generated, for example, by a controller 122 on-board the autonomous guided vehicle 110 (or by any suitable server / computer outside the autonomous guided vehicle 110 but in communication with the on-board controller 122), where at least one machine learning model ML represents the learned weights and parameters of an artificial neural network ANN. The machine learning model ML is configured to perform (with respect to the autonomous guided vehicle 110 reference frame BREF) the detection and localization of objects specific to the defined tasks of the automated guided vehicle 110. For example, for each of the execution / non-execution of operations such as vehicle traversal, picking bin CU, placing bin CU, collision avoidance, top panel label, or structure detection for the positioning of the autonomous guided vehicle 110, remote bin inspection, and any other suitable tasks and / or operating states of the autonomous guided vehicle 110, a machine learning model ML can be generated. The machine learning model ML is saved in any suitable memory of the autonomous guided vehicle 110 (such as the shared memory SH or the media server MS) such that the machine learning model ML is accessible to the controller 122 and its deep director DC.
[0108] The artificial neural network ANN and the machine learning model ML provide real-time object detection, where the artificial neural network ANN of the deep director DC determines ( Figure 6 , frame 620) which detection protocol (e.g., computer vision object detection and localization protocol, machine computer vision object detection and localization protocol, or both), and identifies the object, task, and / or motion state or otherwise utilizes the selected detection protocol to detect ( Figure 6 , frame 625). Refer to Figure 1B and FIGS. 3 to Figure 4B , Figures 11A to 11F , examples of the detected objects include, but are not limited to, the support forks (also called tines) 210AT of the transfer arm 210A, the pusher 470 of the transfer arm 210A, the puller 472 of the transfer arm 210A, the adjustment vane 471 of the transfer arm 210A, the storage shelf cap 444, the electrical panel of the storage and retrieval system 100 structure (see Figure 11E ), the vertical bars 445 of the storage and retrieval system 100 structure (see Figure 4B and also see Figure 11B ), other autonomous guided vehicles 110, the picking aisle 130A, partially identified bins, and suspicious bins ( Figure 11D , e.g., may be mis-placed on the shelf), etc. Refer to Figure 11F , examples of the autonomous guided vehicle tasks include, but are not limited to, picking tasks (e.g., the bin unit is positioned on the shelf in a suitable orientation for picking) and non-picking tasks (e.g., the bin unit is positioned on the shelf in an unsuitable orientation for picking), etc. Examples of the autonomous guided vehicle motion states include, but are not limited to, the movement of the autonomous guided vehicle across a traversing surface (see Figure 10 ,Figure 11B and Figure 11C ), this motion is indicated by blurring in the image frame(s) received from the shared memory SH and / or the media server MS, such as image frame 1000. Here, the coordinates, motion state, and / or task of the object are recognized and obtained from the image frame, such as through the coordinates shown in the corresponding bounding box ( Figures 11A to 11F as shown) and the recognition of the object, task, and / or motion state; however, in other aspects, any suitable combination of vision processing algorithms or instead of the bounding box can be used to recognize and obtain the coordinates of the object in any suitable manner.
[0109] When determining which detection protocol(s) to use (e.g., determining which detection protocol provides the highest level of confidence in object detection), the depth director DC assigns flags to each of the detected objects, tasks, and motion states via an artificial neural network ANN. Note that although flags are used for exemplary purposes of comparison with corresponding predetermined thresholds, any suitable threshold can be used. For example, a flag is a marker indicating the presence of a corresponding condition, object, or motion state. Here, flags are assigned to each detected condition, object, and motion state (e.g., 1 (or other integer greater than 0) indicates presence, and 0 indicates absence). These flags form metadata for the corresponding image frame and are used by the depth director DC (along with the image frame timestamp, detection marker, and one or more of the bounding box or object coordinates) to perform robust object detection and localization based on the video stream imaging data by selecting a detection / localization protocol from one or both of a computer vision protocol and a machine learning protocol. For example, if there is no motion detection in the image frame, the flag for the binocular depth map is set to 1 and the flag for visual maintenance (e.g., occlusion / defect in at least one of the cameras in the stereo camera pair) is set to 0, where the depth director DC selects a machine vision detection protocol for object detection and localization. If there is no motion and the flag for the binocular depth map is set to zero (e.g., binocular vision is blocked, one of the cameras in the camera pair is unavailable, an object is detected in one camera but not the other, etc.) and the flag for visual maintenance is set to 1 (e.g., due to the abnormal stereo vision described above), then the depth director DC selects a machine learning protocol for object detection and localization. If the same object is detected in the image frames of both cameras in the camera pair, the detected object is compared with any suitable threshold (e.g., a detection confidence threshold), and if the threshold is met, the flag for the binocular depth map is set to 1, and the number of detections for each camera in the camera pair for the common object is less than a predetermined threshold such that the visual maintenance flag is set to 1, then the depth director DC selects both a computer vision protocol and a machine learning protocol for object detection and localization.
[0110] Using the vehicle motion as an example, refer to Figure 11C(It shows an image from camera 410A, where the corresponding image frame from camera 410B is substantially similar but with the hand in the opposite direction), the artificial neural network detects the motion present in the image frames from cameras 410A, 410B, and the flags for the motion relative to cameras 410A, 410B are set to 1; however, it should be noted that Figure 11C shows that cameras 410A, 410B have unobstructed views relative to each other. The artificial neural network ANN (as described herein) is trained to recognize the unobstructed camera views and stereo images generated thereby. Here, since the camera views are unobstructed and depth maps can be generated from the stereo camera pair 410A, 410B, computer vision protocols can be used for object detection and localization. Thus, the depth director outputs a selection of computer vision protocols for object detection and localization relative to cameras 410A, 410B. As another example, refer to Figure 12A , which shows an image frame from camera 410B (note that the corresponding image frame from camera 410A may be substantially similar but with the hand in the opposite direction), where a bin or box and adjustment vanes are identified. Here, the bin obstructs the field of view of cameras 410A, 410B and may prevent the generation of stereo images from cameras 410A, 410B; however, each camera provides a corresponding monocular image. Thus, the depth director DC recognizes that the computer vision protocol may not be available and outputs a selection of machine learning protocols for object detection and localization relative to cameras 410A, 410B. As yet another example, Figure 11F shows an image frame from camera 410B (note that the corresponding image frame from camera 410A may be substantially similar but with the hand in the opposite direction), where a fork or fork-like member 210AT, a bin to be picked, and a shelf cap are identified. Here, both cameras have unobstructed views and the number of common objects detected in the image frames from each of the pair of cameras 410A, 410B is less than a predetermined threshold of, for example, 6 objects (in other aspects, this threshold may be greater or less than 6 objects), such that both computer vision and machine learning protocols are available. In the case where both protocols are available, the depth director DC may select both detection and localization protocols such that any suitable image integration algorithm may be used by the depth director DC to combine two-dimensional (i.e., monocular vision) image frames with three-dimensional (stereo vision) image frames and depth maps from three-dimensional sensors 440A, 440B.
[0111] In the case of selecting a machine learning protocol, the depth director DC selects one or more of the deep learning models ML based on the predetermined tasks of the autonomous guided vehicle 110 (such as picking or placing bins, traversing a transfer deck or a picking aisle, etc.) ( Figure 13 , block 1300). The depth director receives the image frames for each camera in the camera pair from the shared memory SH and / or the media server MS ( Figure 13, frame 1305), and use the selected deep learning model ML to perform object detection for a predetermined task within the image frame ( Figure 13 , frame 1310). Using the objects in the detected image frame, the depth director DC performs detection integration ( Figure 13 , frame 1315), where the two-dimensional image frame (with the detected objects) is integrated with the depth map from the 3D sensor (e.g., the image frame from camera 410A is integrated with the depth map from 3D sensor 440A) to form an integrated image or a monocular depth map (e.g., as shown in Figure 9B ). This monocular depth map provides depth information similar to that obtained using stereo or machine vision with a stereo camera pair (such as cameras 410A, 410B, see Figure 9C ). In cases where the images obtained from the video stream data recorded by the controller 122 from a stereo camera pair (e.g., cameras 410A, 410B) do not support binocular vision object detection and localization, the monocular depth map can be used by the controller 122. As described above, the monocular depth includes depth information similar to that obtained using binocular vision, such that any suitable depth information for implementing the operation of the autonomous guided vehicle 110 can be obtained from the monocular depth map. The information obtained from this monocular depth map includes, but is not limited to, identifying the distance of the bin unit from the autonomous guided vehicle, the gap between the bin units stored on the shelf 555, the gap between the picked / placed bin unit and the adjacent bin unit on the shelf 555, and the size of the storage space for the bin unit to be placed. The two-dimensional image frame (with the detected objects) is integrated with the depth map from the 3D sensor by searching, for example, the front-facing plane of the depth map generated using the 3D sensors 440A, 440B associated with the detected objects (e.g., 3D sensor 440A can be associated with the object detected by camera 410A, and 3D sensor 440B can be associated with the object detected by camera 410B). The depth director DC uses the transformed coordinates obtained during the stereo calibration of the camera pair (e.g., camera pair 410A, 410B) to perform this search of the front-facing plane. Locating the object in the integrated image ( Figure 13 , frame 1320), where the object coordinates detected from the integrated image are converted to the coordinates of the global reference frame GREF via the depth map information from one or more 3D sensors 440A, 440B. The localization may also include incorporating the selected points generated from the autonomous guided vehicle model 800 and the associated points from the corresponding images of the autonomous guided vehicle model obtained using the corresponding cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B. It should be understood that cameras 410A, 410B are described above for illustrative purposes, and the above can be implemented using any camera pair described herein.
[0112] As described herein, when an autonomous guided vehicle is performing a task, object detection and localization as determined by, for example, a machine learning protocol, a computer vision protocol, or both, is implemented in real time. As also described herein, detections are generated (e.g., image frames amplified or enhanced by a depth director DC as described herein, see Figures 11A to 12A ) and the detections are output to implement endpoint determination with respect to the operation of the autonomous guided vehicle ( Figure 6 , box 630). Here, the depth director DC outputs detections and indicates to the controller 122 (of which the depth director DC is a module) whether the operation / task of the autonomous guided vehicle 10 is to be performed (i.e., giving the autonomous guided vehicle 110 an instruction to complete the task) or not performed (giving the autonomous guided vehicle 110 an instruction not to complete the task).
[0113] The amplified image frames output by the depth director DC can be profiled live or in real time (such as when the autonomous guided vehicle is performing a task and profiling the amplified image frames to implement task completion or non - completion); and / or as described herein, recorded or stored in any suitable memory (e.g., such as the shared memory SH, the media server MS, the memory of the control server 120, the memory of the warehouse management system 2500, etc.), such that the recorded detections are or can be recognized by the corresponding autonomous guided vehicle 110 as corresponding to the corresponding autonomous guided vehicle 110. For example, in the case where the autonomous guided vehicle 110 cannot perform a task (such as a bin CU being stuck on a shelf or a part of the robot and not being fully transferable to the payload station 210B), an operator can profile live or stored image frames via the user interface UI to determine the reason for the uncompleted task and / or manually control the autonomous guided vehicle to remedy the incorrect transfer of the bin. Here, in one or more aspects, the vision system controller 122VC (and / or the controller 122) is configured to provide remote viewing using the vision system 400, where such remote viewing can be presented to the operator in augmented reality or in any other suitable manner (such as non - augmented). For example, the autonomous transport vehicle 110 is communicatively connected to the warehouse management system 2500 (e.g., via the control server 120) through the network 180 (or any other suitable wireless network). The warehouse management system 2500 includes one or more warehouse control center user interfaces UI. The warehouse control center user interface UI can be any suitable interface, such as, a desktop computer, a laptop computer, a tablet computer, a smart phone, a virtual reality headset, or any other suitable user interface configured to present visual and / or auditory data obtained from the autonomous transport vehicle 110. In some aspects, the vehicle 110 can include one or more microphones MCP ( Figure 2), where one or more microphones and / or remote viewing (e.g., live video images output by the depth director DC and / or augmented images of real-time dissection) can assist in preventive maintenance / troubleshooting diagnosis of system components such as carrier 110, other carriers, elevators, storage racks, etc. The warehouse control center user interface UI is configured such that a warehouse control center user requests or otherwise supplies (such as when an unidentifiable object 299 is detected and / or a suspicious object is detected (see Figure 11D )) images from the autonomous transport carrier 110 and causes the requested / supplied images to be viewed on the warehouse control center user interface UI.
[0114] The supplied and / or requested images can be a live video stream, pre-recorded (and stored in any suitable memory of the autonomous transport carrier 110 or the warehouse management system 2500) images, images or (e.g., one or more static images and / or dynamic videos), which correspond to a specified (user-selectable or preset) time interval or number of images obtained as needed or output by the depth director DC, and are substantially real-time with respect to the corresponding image request. It should be noted that the live video stream and / or image capture provided by the vision system 400, the vision system controller 122VC, and the depth director DC can provide real-time remote control operation (e.g., remote control operation) of the autonomous transport carrier 110 by the warehouse control center user through the warehouse control center user interface UI.
[0115] In some aspects, the live video is streamed (augmented or non-augmented) from the vision system 400 of the supplementary navigation sensor system 288 to the user interface UI as a regular video stream (e.g., the image is presented on the user interface without augmentation, and what the camera "sees" is what is presented), in a manner similar to that described in U.S. Patent Application No. 17 / 804026, filed on May 25, 2022, and titled "Autonomous Transport Carrier with Vision System" (with Attorney Docket No. 1127P016037-US(PAR)), the disclosure of which is incorporated herein by reference in its entirety. A virtual reality headset is used by the user to view the streamed video. Images from the front box unit monitoring camera 410A can be presented in the viewfinder of the virtual reality headset corresponding to the user's left eye, and images from the rear box unit monitoring camera 410B can be presented in the viewfinder of the virtual reality headset corresponding to the user's right eye.
[0116] The image frames output by the depth director DC can also be presented to the user through the virtual reality headset in a manner similar to that described above for the live video. Here, in addition to the detection markers and confidence levels associated with the detected objects, a machine learning model can also be generated due to the augmented features of the detected objects in the image frames of the artificial neural network ANN training. For example, referenceFigure 12B , an image frame from camera 410B is provided with detected objects identified via depth director DC. Controller 122 uses a machine learning model to correspondingly enhance or amplify the detected objects. For example, the machine learning model can be applied by controller 122 to amplify or otherwise enhance the features of the adjustable blade 471 and the bin CU detected in the image. Here, the features of the adjustable blade 471 and the bin CU (such as edges) are highlighted or otherwise amplified to clearly define the boundaries of the detected objects. Note that although Figure 12B corresponding bounding boxes of the adjustable blade 471 and the bin CU are shown, these bounding boxes can be omitted in the case of providing amplification of the detected objects. Although Figure 12B the adjustable rod 471 and the bin CU are shown amplified, note that any suitable structure of the bin, storage and retrieval system 100 and / or the autonomous guided vehicle 110 can be amplified in the image frame output by the depth director DC such that the detected and amplified objects can be determined from the output image frame.
[0117] Reference will be made to Figure 1A , Figure 1B , Figure 2 , Figures 3A to 3C and Figure 14 to describe exemplary methods (e.g., for object detection and localization of the autonomous guided vehicle 110) according to aspects of the disclosed embodiments. Here, an autonomous guided vehicle 110 ( Figure 14 , block 1400) is provided. As described herein, the autonomous guided vehicle 110 is provided with a frame 200 having a payload holder 210B, a drive section 261D coupled to the frame 200, wherein drive wheels 260 support the autonomous guided vehicle 110 on a traversing surface 284, wherein the drive wheels 260 effect the traversing of the vehicle on the traversing surface 284, moving the autonomous guided vehicle 110 on the traversing surface 284 in a facility (e.g., such as a logistics facility, an example of which is the storage and retrieval system 100). A payload handler or transfer arm 210A is coupled to the frame 200 and is configured to transfer payloads (such as bins CU) to and from the payload holder 210B and the storage location 130S of the payload in the storage array SA using a flat indeterminate placement surface disposed in the payload holder or stage 210B. As described herein, a vision system 400 is mounted to the frame 200 and has at least one camera 410A, 410B, 420A, 420B, 430A, 430B, 477A, 477B, and a controller 122 is communicatively coupled to the vision system 400.
[0118] At least one of the cameras 410A, 410B, 420A, 420B, 430A, 430B, 477A, 477B of the vision system 400 generates video stream data imaging of objects in the logistics space (such as those described herein), Figure 14 (frame 1405), where the object is at least one of the following: at least part of the frame 200, at least part of the payload (such as the bin CU), at least part of the payload handler or transfer arm 210A, and at least part of the logistics items or structures in the logistics space other than the autonomous guided vehicle 110. The controller 122 records the video stream data imaging from at least one of the cameras 410A, 410B, 420A, 420B, 430A, 430B, 477A, 477B Figure 14 (frame 1410), and alternately performs robust object detection and positioning within a predetermined reference frame (such as the autonomous guided vehicle reference frame BREF and / or the global reference frame GREF) based on the video stream data imaging via both binocular vision and monocular vision from the video stream data imaging Figure 14 (frame 1415), and the detections determined via monocular vision have a confidence comparable to the detections determined via binocular vision.
[0119] It should be understood that although the vision system 400 and the controller 122 (including one or more machine learning models ML and one or more artificial neural networks ANN) are described herein with respect to the autonomous guided vehicle 110, in other aspects the vision system 400 and one or more machine learning models ML and one or more artificial neural networks ANN can be applied to the controller of the load handling device 150LHD (FIG. 1, which can be substantially similar to the payload platform 210B of the autonomous guided vehicle 110) and the vertical lift 150 or the palletizer that feeds into the transfer station 170. Suitable examples of load handling devices for lifts that can be combined with the vision system 400 are described in U.S. Patent No. 10,947,060, entitled "Vertical Sorter for Product Order Fulfillment," issued on March 16, 2021, the disclosure of which is incorporated herein by reference in its entirety.
[0120] Reference will be made to Figure 1A 、 Figure 1B 、 Figure 2 、 Figures 3A to 3C and Figure 16 to describe an exemplary method (e.g., for object detection and positioning of the autonomous guided vehicle 110) according to aspects of the disclosed embodiments. Here, the autonomous guided vehicle 110 is provided Figure 16, frame 1600). As described herein, the autonomous guided vehicle 110 is provided with a frame 200 having a payload holder 210B, a drive section 261D coupled to the frame 200, wherein drive wheels 260 support the autonomous guided vehicle 110 on a traversing surface 284, and wherein the drive wheels 260 effect the traversal of the vehicle on the traversing surface 284 to move the autonomous guided vehicle 110 on the traversing surface 284 in a facility (e.g., such as a logistics facility, an example of which is the storage and retrieval system 100). A payload handler or transfer arm 210A is coupled to the frame 200 and is configured to transfer a payload (such as a box CU) to and from the payload holder 210B and the storage location 130S of the payload in the storage array SA using a flat, indeterminate placement surface disposed in the payload holder or stage 210B. As described herein, a vision system 400 is mounted to the frame 200 and has at least one camera 410A, 410B, 420A, 420B, 430A, 430B, 477A, 477B. A controller 122 is communicatively coupled to the vision system 400 and is coupled to at least one or more of a three-dimensional imaging system 440A, 440B and a distance sensor (such as one or more of a laser sensor 271 and an ultrasonic sensor 272) that detects the distance to an object (such as those described herein).
[0121] At least one camera 410A, 410B, 420A, 420B, 430A, 430B, 477A, 477B of the vision system 400 generates video stream data imaging of an object in the logistics space ( Figure 16 , frame 1610), wherein the object is at least one of: at least a portion of the frame 200, at least a portion of a payload (such as a box CU), at least a portion of the payload handler or transfer arm 210A, and at least a portion of a logistics item or structure in the logistics space other than the autonomous guided vehicle 110. The controller 122 records the video stream data imaging from at least one camera 410A, 410B, 420A, 420B, 430A, 430B, 477A, 477B ( Figure 16 , frame 1620), and optionally uses binocular vision and monocular vision from the video stream data imaging to perform object detection and localization within a predetermined reference frame (such as at least one of a global reference frame GREF and an autonomous guided vehicle reference frame BREF) based on the video stream data imaging ( Figure 16 , frame 1630), and each of the binocular vision object detection and localization and the monocular vision object detection and localization can be selected by the controller 122 as needed.
[0122] According to one or more aspects of the disclosed embodiments, an autonomous guided vehicle includes: a frame having a payload holder; a drive section coupled to the frame, the drive section having drive wheels that support the autonomous guided vehicle on a traversal surface, the drive wheels effecting traversal of the vehicle on the traversal surface to move the autonomous guided vehicle on the traversal surface in a facility; a payload handler coupled to the frame and configured to transfer a payload to and from the payload holder of the autonomous guided vehicle and a storage location of the payload in a storage array using a flat and indeterminate placement surface disposed in the payload holder; a vision system mounted to the frame, the vision system having at least one camera configured to generate video stream data imaging of objects in a logistics space, the objects being at least one of: at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics item or structure in the logistics space other than the autonomous guided vehicle; and a controller communicatively connected to record the video stream data imaging from the at least one camera and communicatively connected to at least one or more of a time-of-flight sensor and a distance sensor that detect the distance of an object; wherein the controller is configured to alternately perform robust object detection and localization within a predetermined reference frame based on the video stream data imaging via both binocular vision and monocular vision from the video stream data imaging, and the detection determined via monocular vision has a confidence comparable to the detection determined via binocular vision.
[0123] According to one or more aspects of the disclosed embodiments, the controller is configured to perform object detection via monocular vision using a deep machine learning model.
[0124] According to one or more aspects of the disclosed embodiments, the controller is configured to interface with a machine learning module having an object detection function, and the machine learning module determines object detection based on monocular vision and a deep machine learning model.
[0125] According to one or more aspects of the disclosed embodiments, the autonomous guided vehicle further includes a media server communicatively connected to the at least one camera and recording the video stream data, the media server interfacing with the controller, wherein the controller is disposed on board the autonomous guided vehicle or remote from the autonomous guided vehicle.
[0126] According to one or more aspects of the disclosed embodiments, the payload handler is configured to pick up the payload from the bottom of the storage location.
[0127] According to one or more aspects of the disclosed embodiments, at least one camera of the vision system includes two cameras forming a stereo vision camera pair; and the robust object detection and localization enables the payload handler to perform bottom picking of the payload from more than two closely stacked payloads held in adjacent storage locations regardless of the availability of stereo vision from the stereo vision camera pair.
[0128] According to one or more aspects of the disclosed embodiments, at least one camera of the vision system includes two cameras forming a stereo vision camera pair; and robust object detection and localization enable a payload handler to perform bottom picking of a deformable payload from more than two closely stacked payloads held in adjacent storage locations, regardless of the availability of stereo vision from the stereo vision camera pair.
[0129] According to one or more aspects of the disclosed embodiments, at least one camera of the vision system includes two cameras forming a stereo vision camera pair; and robust object detection and localization enable a payload handler to perform bottom picking of a payload from more than two payloads held in adjacent storage locations and having a dynamic Gaussian bin size distribution within a facility, regardless of the availability of stereo vision from the stereo vision camera pair.
[0130] According to one or more aspects of the disclosed embodiments, object localization performed by monocular vision has a confidence comparable to object localization determined via binocular vision.
[0131] According to one or more aspects of the disclosed embodiments, the controller is configured such that binocular vision object detection and localization and monocular vision object detection and localization are interchangeably selectable.
[0132] According to one or more aspects of the disclosed embodiments, the controller has a selector configured to select between binocular vision object detection and localization and monocular vision object detection and localization as needed based on the detection of predetermined operating characteristics of the autonomous guided vehicle.
[0133] According to one or more aspects of the disclosed embodiments, the predetermined characteristic is that the video stream data recorded by the controller does not support binocular vision object detection and localization.
[0134] According to one or more aspects of the disclosed embodiments, a method for object detection and localization of an autonomous guided vehicle is provided. The method includes: providing an autonomous guided vehicle having: a frame with a payload holder; a drive section coupled to the frame, the drive section having drive wheels that support the autonomous guided vehicle on a traversing surface, the drive wheels effecting traversal of the vehicle on the traversing surface to move the autonomous guided vehicle on the traversing surface in a facility; a payload handler coupled to the frame and configured to transfer payloads to and from the payload holder of the autonomous guided vehicle and a storage location of the payload in a storage array using a flat and indeterminate placement surface disposed in the payload holder; a vision system mounted to the frame, the vision system having at least one camera; and a controller communicatively coupled to the vision system; generating video stream data imaging of an object in a logistics space using at least one camera of the vision system, wherein the object is at least one of: at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics item or structure in the logistics space other than the autonomous guided vehicle; recording the video stream data imaging from the at least one camera using the controller; and detecting a distance of the object using at least one or more of a time-of-flight sensor and a distance sensor communicatively coupled to the controller; wherein the controller alternately performs robust object detection and localization within a predetermined reference frame based on the video stream data imaging via both binocular vision and monocular vision from the video stream data imaging, and the detection determined via monocular vision has a confidence level comparable to the detection determined via binocular vision.
[0135] According to one or more aspects of the disclosed embodiments, the controller performs object detection via monocular vision using a deep machine learning model.
[0136] According to one or more aspects of the disclosed embodiments, the controller interfaces with a machine learning module having an object detection function, and the machine learning module determines object detection based on monocular vision and a deep machine learning model.
[0137] According to one or more aspects of the disclosed embodiments, a media server is communicatively connected to at least one camera and records the video stream data, and the media server interfaces with the controller, wherein the controller is arranged to be airborne on the autonomous guided vehicle or away from the autonomous guided vehicle.
[0138] According to one or more aspects of the disclosed embodiments, the payload handler picks up the payload from the bottom of the storage location.
[0139] According to one or more aspects of the disclosed embodiments, at least one camera of a vision system includes two cameras forming a stereo vision camera pair, and the method further includes: performing bottom picking of payloads by a payload handler from more than two closely stacked payloads held in adjacent storage locations via robust object detection and localization, regardless of the availability of stereo vision from the stereo vision camera pair.
[0140] According to one or more aspects of the disclosed embodiments, at least one camera of a vision system includes two cameras forming a stereo vision camera pair, and the method further includes: performing bottom picking of deformed payloads by a payload handler from more than two closely stacked payloads held in adjacent storage locations via robust object detection and localization, regardless of the availability of stereo vision from the stereo vision camera pair.
[0141] According to one or more aspects of the disclosed embodiments, at least one camera of a vision system includes two cameras forming a stereo vision camera pair, and the method further includes: performing bottom picking of payloads by a payload handler from more than two payloads held in adjacent storage locations and having a dynamic Gaussian bin size distribution within a facility via robust object detection and localization, regardless of the availability of stereo vision from the stereo vision camera pair.
[0142] According to one or more aspects of the disclosed embodiments, object localization performed by monocular vision has a confidence level comparable to that of object localization determined via binocular vision.
[0143] According to one or more aspects of the disclosed embodiments, binocular vision object detection and localization and monocular vision object detection and localization are interchangeably selectable via a controller.
[0144] According to one or more aspects of the disclosed embodiments, the controller has a selector configured to select between binocular vision object detection and localization and monocular vision object detection and localization as needed based on the detection of predetermined operating characteristics of an autonomous guided vehicle.
[0145] According to one or more aspects of the disclosed embodiments, the predetermined characteristic is that the video stream data recorded by the controller does not support binocular vision object detection and localization.
[0146] According to one or more aspects of the disclosed embodiments, an autonomous guided vehicle includes: a frame having a payload holder; a drive section coupled to the frame, the drive section having drive wheels that support the autonomous guided vehicle on a traversing surface, the drive wheels effecting traversal of the vehicle on the traversing surface to move the autonomous guided vehicle on the traversing surface in a facility; a payload handler coupled to the frame and configured to transfer a payload to and from the payload holder of the autonomous guided vehicle and a storage location of the payload in a storage array using a flat and indeterminate placement surface disposed in the payload holder; a vision system mounted to the frame, the vision system having at least one camera configured to generate video stream data imagery of objects in a logistics space, the objects being at least one of: at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics article or structure in the logistics space other than the autonomous guided vehicle; and a controller communicatively coupled to record video stream data imagery from the at least one camera and communicatively coupled to at least one or more of a time-of-flight sensor and a distance sensor that detect distances to the objects; wherein the controller is configured to selectively perform object detection and localization within a predetermined reference frame based on video stream data imagery using binocular vision and monocular vision from the video stream data imagery, and each of the binocular vision object detection and localization and the monocular vision object detection and localization can be selected by the controller as needed.
[0147] According to one or more aspects of the disclosed embodiments, monocular vision detection and localization has a confidence level comparable to that of detection determined via binocular vision object detection and localization.
[0148] According to one or more aspects of the disclosed embodiments, the controller is configured such that the binocular vision object detection and localization and the monocular vision object detection and localization are interchangeably selectable.
[0149] According to one or more aspects of the disclosed embodiments, the controller has a selector configured to select between binocular vision object detection and localization and monocular vision object detection and localization as needed based on detection of predetermined operating characteristics of the autonomous guided vehicle.
[0150] According to one or more aspects of the disclosed embodiments, the predetermined characteristic is that the video stream data imagery recorded by the controller does not support binocular vision object detection and localization.
[0151] According to one or more aspects of the disclosed embodiments, the controller is configured to perform object detection via monocular vision using a deep machine learning model.
[0152] According to one or more aspects of the disclosed embodiments, the controller is configured to interface with a machine learning module having object detection capabilities, and the machine learning module determines object detection based on monocular vision and a deep machine learning model.
[0153] According to one or more aspects of the disclosed embodiments, the autonomous guided vehicle further includes a media server communicatively connected to at least one camera and recording video stream data, and the media server interfaces with the controller, where the controller is arranged to be onboard the autonomous guided vehicle or away from the autonomous guided vehicle.
[0154] According to one or more aspects of the disclosed embodiments, a method for object detection and positioning of an autonomous guided vehicle is provided. The method includes: providing an autonomous guided vehicle having: a frame with a payload holder; a drive section coupled to the frame, having drive wheels that support the autonomous guided vehicle on a traversing surface, and the drive wheels effect the traversing of the vehicle on the traversing surface to move the autonomous guided vehicle on the traversing surface in a facility; a payload handler coupled to the frame and configured to transfer a payload to and from the payload holder of the autonomous guided vehicle and a storage location of the payload in a storage array using a flat and indeterminate placement surface disposed in the payload holder; a vision system mounted to the frame, having at least one camera configured to generate video stream data imaging of objects in a logistics space, where the objects are at least one of: at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics item or structure in the logistics space other than the autonomous guided vehicle; and a controller communicatively connected to the vision system and communicatively connected to at least one or more of a time-of-flight sensor and a distance sensor that detect the distance of an object; recording, using the controller, video stream data imaging from the at least one camera; and selectively performing, using the controller, object detection and positioning within a predetermined reference frame based on the video stream data imaging using binocular vision and monocular vision from the video stream data imaging, where each of the binocular vision object detection and positioning and the monocular vision object detection and positioning can be selected by the controller as needed.
[0155] According to one or more aspects of the disclosed embodiments, the monocular vision detection and positioning has a confidence level comparable to that determined by binocular vision object detection and positioning.
[0156] According to one or more aspects of the disclosed embodiments, the binocular vision object detection and positioning and the monocular vision object detection and positioning are interchangeably selected by the controller.
[0157] In accordance with one or more aspects of the disclosed embodiments, the controller has a selector configured to select, as needed, between binocular vision object detection and localization and monocular vision object detection and localization based on the detection of predetermined operating characteristics of the autonomous guided vehicle.
[0158] In accordance with one or more aspects of the disclosed embodiments, the predetermined characteristic is that the video stream data recorded by the controller does not support binocular vision object detection and localization.
[0159] In accordance with one or more aspects of the disclosed embodiments, the controller performs object detection via monocular vision using a deep machine learning model.
[0160] In accordance with one or more aspects of the disclosed embodiments, the controller interfaces with a machine learning module having an object detection function, and the machine learning module determines object detection based on monocular vision and a deep machine learning model.
[0161] In accordance with one or more aspects of the disclosed embodiments, the media server is communicatively connected to at least one camera and records video stream data, and the media server interfaces with the controller, where the controller is configured to be airborne on the autonomous guided vehicle or away from the autonomous guided vehicle.
[0162] It should be understood that the foregoing description is only illustrative of aspects of the disclosed embodiments. Those skilled in the art can design various alternatives and modifications without departing from the aspects of the disclosed embodiments. Therefore, the various aspects of the disclosed embodiments are intended to cover all such alternatives, modifications, and variations that fall within the scope of any of the appended claims. Additionally, the fact that different features are recited only in mutually different dependent or independent claims does not indicate that combinations of these features cannot be used advantageously, and such combinations are still within the scope of the aspects of the disclosed embodiments.
[0162] What is claimed is:.
Claims
1. An autonomous guided vehicle, comprising: A frame having a payload retainer; A drive section coupled to the frame, having drive wheels that support the autonomous guided vehicle on a traversing surface, the drive wheels effecting traversal of the vehicle on the traversing surface to move the autonomous guided vehicle on the traversing surface in a facility; A payload handler coupled to the frame and configured to transfer a payload to and from the payload retainer of the autonomous guided vehicle and a storage location of the payload in a storage array using a flat and indeterminate placement surface disposed in the payload retainer; A vision system mounted to the frame, having at least one camera configured to generate video stream data imaging of an object in a logistics space, the object being at least one of: at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics item or structure in the logistics space other than the autonomous guided vehicle; And A controller communicatively connected to record the video stream data imaging from the at least one camera and communicatively connected to at least one or more of a time-of-flight sensor and a distance sensor that detect a distance to the object; Wherein the controller is configured to alternately perform robust object detection and localization within a predetermined reference frame based on the video stream data imaging via both binocular vision and monocular vision from the video stream data imaging, and the detection determined via monocular vision has a confidence level comparable to the detection determined via binocular vision.
2. The autonomous guided vehicle according to claim 1, wherein, The controller is configured to perform object detection via monocular vision using a deep machine learning model.
3. The autonomous guided vehicle according to claim 1, wherein, The controller is configured to interface with a machine learning module having an object detection function, the machine learning module determining object detection based on the monocular vision and the deep machine learning model.
4. The autonomous guided vehicle according to claim 1, further comprising a media server communicatively connected to the at least one camera and recording the video stream data, the media server interfacing with the controller, wherein the controller is provided on board the autonomous guided vehicle or remote from the autonomous guided vehicle.
5. The autonomous guided vehicle according to claim 1, wherein, The payload handler is configured to pick the payload from the bottom of the storage location.
6. The autonomous guided vehicle according to claim 1, wherein: At least one camera of the vision system includes two cameras forming a stereo vision camera pair; and The robust object detection and localization enables the payload handler to perform bottom picking of the payload from more than two closely stacked payloads held in adjacent storage locations, regardless of the availability of stereo vision from the stereo vision camera pair.
7. The autonomous guided vehicle according to claim 1, wherein: At least one camera of the vision system includes two cameras forming a stereo vision camera pair; and The robust object detection and localization enables the payload handler to perform bottom picking of deformed payloads from more than two closely stacked payloads held in adjacent storage positions, regardless of the availability of stereovision from the stereovision camera pair.
8. The autonomous guided vehicle according to claim 1, wherein: at least one camera of the vision system includes two cameras forming a stereovision camera pair; and the robust object detection and localization enables the payload handler to perform bottom picking of the payload from more than two payloads held in adjacent storage positions and having a dynamic Gaussian bin size distribution within the facility, regardless of the availability of stereovision from the stereovision camera pair.
9. The autonomous guided vehicle according to claim 1, wherein, The object localization performed by the monocular vision has a confidence comparable to the object localization determined via the binocular vision.
10. The autonomous guided vehicle according to claim 1, wherein, The controller is configured such that the binocular vision object detection and localization and the monocular vision object detection and localization are interchangeably selectable.
11. The autonomous guided vehicle according to claim 1, wherein, The controller has a selector configured to select between the binocular vision object detection and localization and the monocular vision object detection and localization as needed based on the detection of a predetermined operating characteristic of the autonomous guided vehicle.
12. The autonomous guided vehicle according to claim 11, wherein, The predetermined characteristic is that the video stream data recorded by the controller does not support the binocular vision object detection and localization.
13. A method for object detection and localization of an autonomous guided vehicle, the method comprising: providing an autonomous guided vehicle having: a frame having a payload holder; a drive section coupled to the frame, the drive section having drive wheels that support the autonomous guided vehicle on a traversing surface, the drive wheels effecting traversal of the vehicle on the traversing surface to move the autonomous guided vehicle on the traversing surface within the facility; a payload handler coupled to the frame and configured to transfer payloads to and from the payload holder of the autonomous guided vehicle and the storage positions of the payloads in a storage array using a flat, indeterminate placement surface disposed in the payload holder; a vision system mounted to the frame, the vision system having at least one camera; and a controller communicatively coupled to the vision system; using the at least one camera of the vision system to generate video stream data imaging of an object in a logistics space, where the object is at least one of: at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics item or structure in the logistics space other than the autonomous guided vehicle; using the controller to record the video stream data imaging from the at least one camera; and using at least one or more of a time-of-flight sensor and a distance sensor communicatively coupled to the controller to detect the distance to the object; The controller alternately performs robust object detection and localization within a predetermined reference frame based on the video stream data imaging via both binocular vision and monocular vision from the video stream data imaging, and the detection determined via monocular vision has a confidence level comparable to that of the detection determined via binocular vision.
14. The method according to claim 13, wherein The controller utilizes a deep machine learning model to perform object detection via the monocular vision.
15. The method according to claim 13, wherein, The controller interfaces with a machine learning module having an object detection function, and the machine learning module determines object detection based on monocular vision and a deep machine learning model.
16. The method according to claim 13, wherein, A media server is communicatively connected to the at least one camera and records the video stream data, and the media server interfaces with the controller, where the controller is configured to be onboard the autonomous guided vehicle or away from the autonomous guided vehicle.
17. The method according to claim 13, wherein, The payload handler picks up the payload from the bottom of the storage location.
18. The method according to claim 13, wherein, At least one camera of the vision system includes two cameras forming a stereo vision camera pair, and the method further includes: Performing bottom picking of the payload by the payload handler from more than two closely stacked payloads held in adjacent storage locations via the robust object detection and localization, regardless of the availability of stereo vision from the stereo vision camera pair.
19. The method according to claim 13, wherein At least one camera of the vision system includes two cameras forming a stereo vision camera pair, and the method further includes: Performing bottom picking of a deformed payload by the payload handler from more than two closely stacked payloads held in adjacent storage locations via the robust object detection and localization, regardless of the availability of stereo vision from the stereo vision camera pair.
20. The method according to claim 13, wherein, At least one camera of the vision system includes two cameras forming a stereo vision camera pair, and the method further includes: Performing bottom picking of the payload by the payload handler from more than two payloads held in adjacent storage locations and having a dynamic Gaussian bin size distribution within the facility via the robust object detection and localization, regardless of the availability of stereo vision from the stereo vision camera pair.
21. The method according to claim 13, wherein, The object localization performed by the monocular vision has a confidence level comparable to that of the object localization determined via binocular vision.
22. The method according to claim 13, wherein Binocular vision object detection and localization and monocular vision object detection and localization are interchangeably selectable via the controller.
23. The method according to claim 13, wherein, The controller has a selector configured to select between binocular vision object detection and localization and monocular vision object detection and localization as needed based on the detection of predetermined operating characteristics of the autonomous guided vehicle.
24. The method according to claim 23, wherein, The predetermined characteristic is that the video stream data recorded by the controller does not support the binocular vision object detection and localization.
25. An autonomous guided vehicle, comprising: A frame having a payload holder; A drive section coupled to the frame, having drive wheels that support the autonomous guided vehicle on a traversing surface, and the drive wheels perform traversing of the vehicle on the traversing surface to move the autonomous guided vehicle on the traversing surface in the facility; A payload handler coupled to the frame and configured to transfer a payload to and from a payload retainer of the autonomous guided vehicle and a storage location of the payload in a storage array using a flat, indeterminate placement surface disposed in the payload retainer; A vision system mounted to the frame having at least one camera configured to generate video stream data imagery of an object in a logistics space, the object being at least one of: at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics article or structure in the logistics space other than the autonomous guided vehicle; And A controller communicatively coupled to record the video stream data imagery from the at least one camera and communicatively coupled to at least one or more of a time-of-flight sensor and a distance sensor that detect a distance to the object; Wherein the controller is configured to selectively perform object detection and positioning within a predetermined reference frame using binocular vision and monocular vision from the video stream data imagery, and each of binocular vision object detection and positioning and monocular vision object detection and positioning can be selected by the controller as needed.
26. The autonomous guided vehicle according to claim 25, wherein, The monocular vision detection and positioning has a confidence level comparable to the detection determined via the binocular vision object detection and positioning.
27. The autonomous guided vehicle according to claim 25, wherein, The controller is configured such that the binocular vision object detection and positioning and the monocular vision object detection and positioning are interchangeably selectable.
28. The autonomous guided vehicle according to claim 25, wherein, The controller has a selector configured to select between the binocular vision object detection and positioning and the monocular vision object detection and positioning as needed based on a detection of a predetermined operating characteristic of the autonomous guided vehicle.
29. The autonomous guided vehicle according to claim 28, wherein, The predetermined characteristic is that the video stream data imagery recorded by the controller does not support the binocular vision object detection and positioning.
30. The autonomous guided vehicle according to claim 25, wherein, The controller is configured to perform object detection via the monocular vision using a deep machine learning model.
31. The autonomous guided vehicle according to claim 25, wherein, The controller is configured to interface with a machine learning module having an object detection function, and the machine learning module determines object detection based on the monocular vision and the deep machine learning model.
32. The autonomous guided vehicle according to claim 25, further comprising a media server communicatively coupled to the at least one camera and recording the video stream data, the media server interfacing with the controller, wherein the controller is disposed onboard the autonomous guided vehicle or remote from the autonomous guided vehicle.
33. A method for object detection and positioning of an autonomous guided vehicle, the method comprising: Providing an autonomous guided vehicle having: A frame having a payload retainer; A drive section coupled to the frame having drive wheels that support the autonomous guided vehicle on a traversal surface, the drive wheels effecting traversal of the vehicle on the traversal surface to move the autonomous guided vehicle on the traversal surface in a facility; A payload handler coupled to the frame and configured to transfer a payload to and from a payload retainer of the autonomous guided vehicle and a storage location of the payload in a storage array using a flat and indefinite placement surface disposed in the payload retainer; A vision system mounted to the frame having at least one camera configured to generate video stream data imagery of an object in a logistics space, the object being at least one of: at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics article or structure in the logistics space other than the autonomous guided vehicle; And A controller communicatively coupled to the vision system and communicatively coupled to at least one or more of a time-of-flight sensor and a distance sensor that detect a distance to the object; Recording video stream data imagery from the at least one camera using the controller; And Optionally implementing object detection and positioning within a predetermined reference frame using binocular vision and monocular vision from the video stream data imagery using the controller, each of the binocular vision object detection and positioning and the monocular vision object detection and positioning being selectable by the controller as needed.
34. The method according to claim 33, wherein, The monocular vision detection and positioning has a confidence comparable to the detection determined via the binocular vision object detection and positioning.
35. The method according to claim 33, wherein, The binocular vision object detection and positioning and the monocular vision object detection and positioning are interchangeably selectable by the controller.
36. The method according to claim 33, wherein, The controller has a selector configured to select between the binocular vision object detection and positioning and the monocular vision object detection and positioning as needed based on a detection of a predetermined operating characteristic of the autonomous guided vehicle.
37. The method according to claim 36, wherein, The predetermined characteristic is that the video stream data imagery recorded by the controller does not support the binocular vision object detection and positioning.
38. The method according to claim 33, wherein, The controller implements object detection via the monocular vision using a deep machine learning model.
39. The method according to claim 33, wherein, The controller interfaces with a machine learning module having an object detection function, the machine learning module determining object detection based on the monocular vision and the deep machine learning model.
40. The method according to claim 33, wherein, A media server communicatively coupled to the at least one camera and recording the video stream data, the media server interfacing with the controller, wherein the controller is configured to be onboard the autonomous guided vehicle or remote from the autonomous guided vehicle.
Citation Information
Patent Citations
Vertical sequencer for product order fulfillment
US10947060B2
Automated bot with transfer arm
US11078017B2
Pallet building system with flexible sequencing
US11305430B2
Autonomous transport vehicle with synergistic vehicle dynamic response
US12151922B2
Vertical sequencer for product order fulfillment
US20190389671A1
Cited By
Unmanned aerial vehicle binocular vision obstacle avoidance system and distance measurement and obstacle avoidance method thereof
CN121163464A