Logistics autonomous vehicles with robust object detection, localization, and monitoring
The autonomous guided vehicle uses a vision system with binocular and monocular protocols and AI models to address navigation and detection challenges in ultra-constrained environments, achieving reliable object detection and localization with low failure rates.
Patent Information
- Application Number
- JP2025518331
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-26
- Filing Date
- 2023-09-27
- Publication Date
- 2025-11-06
AI Technical Summary
Autonomous vehicles in automated logistics systems face challenges in navigation and object detection due to occlusion, obstruction, and degraded image processing, particularly in ultra-constrained environments with closely spaced and deformed cases, which affect the reliability of stereo camera-based guidance and localization.
The autonomous guided vehicle employs a vision system with binocular and monocular vision protocols, utilizing artificial neural networks and machine learning models to enhance object detection and localization, switching between protocols based on environmental constraints to maintain robust detection and localization even when stereo vision is impaired.
The system achieves high reliability in object detection and localization, reducing picking failures to approximately 1 in 1 million, despite ultra-constrained conditions such as closely spaced and deformed cases, ensuring efficient case transfer within the logistics facility.
Smart Images

Figure 2025536455000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application is a non-provisional application of and claims the benefit of U.S. Provisional Patent Application No. 63 / 377,271, filed September 27, 2022, the entire disclosure of which is incorporated herein by reference.
[0002] [Technical field] The disclosed embodiments relate generally to material handling systems, and more particularly to conveyors for automated logistics systems. [Background technology]
[0003] [Brief description of related developments] Generally, automated logistics systems, such as automated storage and retrieval systems, utilize autonomous vehicles that transport goods within the automated storage and retrieval system. These autonomous vehicles are guided throughout the automated storage and retrieval system by location beacons, capacitive or inductive proximity sensors, line following sensors, reflective beam sensors, and other narrow-focus beam sensors. These sensors may provide limited information to effect navigation of the autonomous vehicles through the storage and retrieval system or may provide limited information regarding the identification and discrimination of hazardous materials that may be present throughout the automated storage and retrieval system.
[0004] Autonomous vehicles may also be guided throughout an automated storage and retrieval system by vision systems employing stereo or binocular cameras. However, in a logistics environment, stereo camera pairs may be impaired, for example, by occlusion or obstruction and / or obscuration of one camera of the stereo camera pair (e.g., by a payload carried by the autonomous vehicle, a storage structure, etc.), or may not always be available, or image processing quality may be degraded from processing overlapping image data or imagery that is otherwise unsuitable (e.g., blurred, etc.) for guiding and locating an autonomous vehicle within the automated storage and retrieval system. Summary of the Invention
[0005] The foregoing aspects and other features of the disclosed embodiments are explained in the following description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0006] [Figure 1A] 1 is a schematic diagram of a logistics facility incorporating aspects of the disclosed embodiments; [Figure 1B] FIG. 1B is a schematic diagram of the logistics facility of FIG. 1A in accordance with aspects of the disclosed embodiment; [Figure 2] FIG. 1B is a schematic diagram of an autonomous guided vehicle of the logistics facility of FIG. 1A in accordance with aspects of the disclosed embodiment; [Figure 3A] FIG. 3 is a schematic diagram of a portion of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 3B] FIG. 3 is a schematic diagram of a portion of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 3C] FIG. 3 is a schematic diagram of a portion of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 4A] 3 is an example of image data captured by a vision system of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 4B] 3 is an example of image data captured by a vision system of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 4C] 3 is an example of image data captured by a vision system of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 5] FIG. 3 is a schematic diagram of a portion of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 6] FIG. 10 is an exemplary flow diagram of a method according to aspects of the disclosed embodiment; [Figure 7A] FIG. 1 is a schematic illustration of a calibration fixture or jig in accordance with aspects of the disclosed embodiment; [Figure 7B]FIG. 1 is a schematic illustration of a calibration fixture or jig in accordance with aspects of the disclosed embodiment; [Figure 8] FIG. 3 is an exemplary diagram of a computer model of the autonomous guided vehicle of FIG. 2 (and portions thereof) in accordance with aspects of the disclosed embodiment. [Figure 9A] 3 is an exemplary monocular image from a camera of a vision system of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 9B] 9B is an exemplary depth map generated from the exemplary monocular image of FIG. 9A in accordance with aspects of the disclosed embodiment; [Figure 9C] FIG. 3 is an exemplary illustration of disparity map generation utilizing a pair of cameras of the vision system of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment. [Figure 10] 3 is an exemplary user interface of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 11A] 1 is an exemplary image frame of video stream data imaging of an object captured from one or more forward navigation cameras or one or more rear navigation cameras of an autonomous guided vehicle, where detected image features are enclosed by bounding boxes for illustrative purposes only and may be identified within the image frame in any suitable manner, in accordance with aspects of the disclosed embodiment; [Figure 11B] 1 is an exemplary image frame of video stream data imaging of an object taken from one or more case monitoring cameras of an autonomous guided vehicle in accordance with aspects of the disclosed embodiment, where detected image features are enclosed by bounding boxes for illustrative purposes only and may be identified within the image frame in any suitable manner. [Figure 11C] 1 is an exemplary image frame of video stream data imaging of an object taken from one or more case monitoring cameras of an autonomous guided vehicle in accordance with aspects of the disclosed embodiment, where detected image features are enclosed by bounding boxes for illustrative purposes only and may be identified within the image frame in any suitable manner. [Figure 11D]1 is an exemplary image frame of video stream data imaging of an object taken from one or more case monitoring cameras of an autonomous guided vehicle in accordance with aspects of the disclosed embodiment, where detected image features are enclosed by bounding boxes for illustrative purposes only and may be identified within the image frame in any suitable manner. [Figure 11E] 1 is an exemplary image frame of video stream data imaging of an object taken from one or more case monitoring cameras of an autonomous guided vehicle in accordance with aspects of the disclosed embodiment, where detected image features are enclosed by bounding boxes for illustrative purposes only and may be identified within the image frame in any suitable manner. [Figure 11F] 1 is an exemplary image frame of video stream data imaging of an object taken from one or more case monitoring cameras of an autonomous guided vehicle in accordance with aspects of the disclosed embodiment, where detected image features are enclosed by bounding boxes for illustrative purposes only and may be identified within the image frame in any suitable manner. [Figure 12A] 3 is an exemplary augmented image of the vision system of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 12B] 3 is an exemplary augmented image of the vision system of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 13] FIG. 10 is an exemplary flow diagram of a method according to aspects of the disclosed embodiment; [Figure 14] FIG. 10 is an exemplary flow diagram of a method according to aspects of the disclosed embodiment; [Figure 15] 1B is an exemplary diagrammatical illustration of calibration station(s) of the logistics facility of FIG. 1A in accordance with aspects of the disclosed embodiment; [Figure 16] FIG. 10 is an exemplary flow diagram of a method according to aspects of the disclosed embodiment; DETAILED DESCRIPTION OF THE INVENTION
[0007] 1A and 1B illustrate an exemplary automated storage and retrieval system 100 in accordance with aspects of the disclosed embodiment. While aspects of the disclosed embodiment will be described with reference to the drawings, it should be understood that aspects of the disclosed embodiment can be embodied in many forms. Furthermore, any suitable size, shape, or type of elements or materials may be used.
[0008] Aspects of the disclosed embodiments provide a logistics autonomous guided vehicle 110 (referred to herein as an autonomous guided vehicle) with intelligent autonomy and coordinated operation. For example, the autonomous guided vehicle 110 includes a vision system 400 (see FIG. 2 ) having at least one camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B positioned to generate video stream data imaging of an object within a logistics space (such as the operating environment or space of the storage and retrieval system 100). The object is one or more of at least a portion of the frame 200 of the autonomous guided vehicle 110, at least a portion of a payload (e.g., a case CU), at least a portion of the transfer arm 210A, and at least a portion of a logistics item (another case CU) or structure (e.g., the storage and retrieval system 100) within the logistics space outside the autonomous guided vehicle 110. By way of example, vision system 400 employs at least stereoscopic or binocular vision configured to provide detection of cases CUs and objects (such as facility structures and unwanted foreign / transient objects) within a logistics facility such as automated storage and retrieval system 100, as well as localization of autonomous guided vehicles within automated storage and retrieval system 100. Vision system 400 also provides collaborative vehicle operations by providing images (still images or video streams, live or recorded) to an operator of automated storage and retrieval system 100, where these images, in some embodiments, are provided via a user interface UI as described herein and as augmented images as illustrated in FIG. 10 .
[0009] As described in more detail herein, the autonomous guided vehicle 110 includes a controller 122 that is programmed with one or more machine learning models ML and one or more artificial neural networks ANN that access data from the vision system 400 to provide robust case / object detection and localization regardless of obstructions of one or more cameras of the vision system 400. Here, the case / object detection is robust in that the case / object is detected and localized through the employment of artificial neural networks ANN and machine learning models ML that provide detection and localization effectiveness comparable to that obtained with stereo vision, even when stereo vision is not available in an ultra-constrained system or operating environment. The hyper-constrained system includes, but is not limited to, at least the following constraints: closely spaced spacing between adjacent cases, the autonomous guided vehicle being configured to under-pick (lift from below) cases, cases of various sizes being distributed in a Gaussian distribution within the storage array SA, cases may exhibit deformation, and cases may be arranged in an irregular manner on the support surface, all of which affect the transfer of case units CU between storage shelves 555 (or other case holding locations) and the autonomous guided vehicle 110.
[0010] Another constraint of the ultra-constrained system is the transfer time required for the autonomous guided vehicle 110 to transfer a case unit(s) between the payload platform 210B of the autonomous guided vehicle 110 and a case holding location (e.g., a storage space, a buffer, a transfer station, or other case holding location described herein), where the transfer time for a case transfer is about 10 seconds or less. Therefore, the vision system 400 identifies the location and pose of the case (or the location and pose of the holding station) in less than about 2 seconds, or less than about 0.5 seconds.
[0011] As described above, the cases CU stored in the storage and retrieval system have a Gaussian distribution (see FIG. 4A ) with respect to case size within picking aisle 130A and with respect to case size across storage array SA, such that the size of any given storage space on storage shelf 555 dynamically changes (e.g., a dynamic Gaussian case size distribution) as cases are picked and placed. As such, autonomous guided vehicle 110 is configured to identify cases held in dynamically sized storage spaces (depending on the case held), despite occlusion of one camera in a set of stereo cameras that provides object detection and localization, as described herein.
[0012] Further, as seen in FIG. 4A , for example, the cases CU are arranged on the storage shelf 555 (or other storage station) in a closely coupled or closely spaced relationship, where the distance DIST between adjacent case units CU is approximately half the distance between the storage shelf hats 444. The distance / width DIST between the hats 444 of the support slats 520L is approximately 2.5 inches. The close spacing of the case CUs can be complicated (i.e., the spacing can be less than half the distance between the storage shelf hats 444) in that the case CUs (e.g., deformed cases—see FIGS. 4A-4C illustrating case deformation with open flaps) may exhibit deformations (e.g., bulging sides, open flaps, convex sides, etc.) and / or may be tilted relative to the hats 444 on which they rest (i.e., the front of the case may not be parallel to the front of the storage shelf 555 and the sides of the case may not be parallel to the hats 555 of the storage shelf 555 (see FIG. 4A )). Case deformation and tilted case placement can further reduce the spacing between adjacent cases, such that the autonomous guided vehicle is configured to determine picking interference between closely spaced adjacent cases despite occlusion of one camera in a set of stereo cameras that provides object detection and localization, as described herein.
[0013] It is also noted that the height HGT of the hats 444 is approximately 2 inches, and the space envelope ENV between the hats 444, through which the tines 210AT of the transfer arm 210A of the autonomous guided vehicle 110 are inserted under the case units CU to pick / place cases to / from the storage shelves 555, has a width of approximately 1.7 inches and a height of approximately 1.2 inches (see, for example, Figures 3A, 3C, and 4A). Under-picking of a case CU by an autonomous guided vehicle requires interfacing with a case CU held on a storage shelf 555 at a picking / case support surface (defined by case seating surfaces 444S of hats 444 (see FIG. 4A )) without collision between tines 210AT of transfer arm 210A of autonomous guided vehicle 110 and hats 444 / slats 520L, without collision between tines 210AT and adjacent cases (not being picked), and without collision between the case being picked and adjacent cases not being picked, all of which is provided by the placement of tines 210AT in an envelope ENV between hats 444. As such, the autonomous guided vehicle is configured to detect and locate a spatial envelope ENV for inserting tines 210AT of transfer arm 210A under a given case CU to pick the case, regardless of occlusion of one camera in a set of stereo cameras that provide object detection and localization, as described herein.
[0014] The above-described hyper-constraint system requires robustness of the vision system, and can be considered to define the robustness of the vision system 400 because the vision system 400 is configured to accommodate the above-described constraints even when the stereoscopic vision provided by the vision system 400 is unavailable, and can provide pose and localization information for the case CU and / or the autonomous guided vehicle 110, resulting in an autonomous guided vehicle picking failure rate of approximately 1 picking failure per approximately 1 million picks.
[0015] The robustness of the vision system 400 is achieved, at least in part, when the controller 122 includes a control module (referred to herein as a deep conductor module, DC) that includes an artificial neural network (ANN) and is configured to select, via the artificial neural network (ANN), a detection / localization protocol from one (or both) of a computer vision protocol (e.g., one employing binocular / stereoscopic vision) and a machine learning protocol (e.g., one employing monocular vision in combination with monocular vision data analysis by a machine learning model ML and an artificial neural network (ANN). As described herein, the controller 122 is configured to selectively use binocular and monocular vision from the video stream data imaging to detect and localize objects within a predetermined reference frame (e.g., a global reference frame GREF and / or an autonomous guided vehicle reference frame BREF). Each of the binocular and monocular object detection and localization is selectable on demand by the controller 122. Controller 122 is configured such that binocular vision object detection and localization and monocular vision object detection and localization are interchangeably selectable by controller 122. Controller 122 has a selector 122SL (provided by deep conductor DC described herein) arranged to select between binocular vision object detection and localization and monocular vision object detection and localization on demand based on detection of a predetermined operating characteristic of autonomous guided vehicle 110 (such as video stream data imaging recorded by the controller that does not support binocular vision object detection and localization).Here, in accordance with aspects of the disclosed embodiment, the autonomous guided vehicle 110 acquires data with binocular vision (as the autonomous guided vehicle 110 travels within the storage and retrieval system 100) and switches to monocular vision on demand (or on the fly) if the data from the binocular vision is not suitable for providing object detection and localization as described herein. The data obtained from the monocular object detection and localization is such that the controller 122 can perform object detection and localization anew for any given image frame data without knowing the data from previously acquired image frame data.
[0016] The computer vision protocol and the machine learning protocol operate simultaneously, and the deep conductor module DC determines which protocol provides the highest detection and / or localization reliability and selects the protocol with the highest detection and / or localization reliability. It is noted that once the machine learning protocol is selected, the monocular vision data (processed by one or more machine learning models ML and one or more artificial neural networks ANN) can be used with equal effect in place of binocular or stereoscopic (the terms binocular and stereoscopic are used interchangeably herein) vision data (i.e., the available, unobstructed, unobscured, focused, or otherwise unimpaired binocular vision data of the computer vision protocol).
[0017] According to aspects of the disclosed embodiment, the automated storage and retrieval system 100 of FIGS. 1A and 1B may be deployed in a retail distribution (logistics) center or warehouse to fulfill orders received from retailers for replenishment items shipped in cases, packages, and / or parcels, for example. The terms case, package, and parcel are used interchangeably herein and, as previously mentioned, may refer to any container that may be used for shipping and that may be filled with one or more product units by a manufacturer. A case(s) as used herein refers to a case, package, or parcel unit that is not stored (e.g., not contained) in a tray, on a tote, or the like. It is noted that a case unit CU (also referred to herein as a mixed case, case, and shipping unit) may include a case of items / units (e.g., a case of soup cans, cereal boxes, etc.) or individual items / units adapted to be removed from or placed on a pallet. According to an exemplary embodiment, shipping cases or case units (e.g., cartons, barrels, boxes, crates, jugs, shrink-wrapped trays or groups, or any other suitable device for holding case units) may have variable sizes, may be used to hold case units during shipping, and may be configured to be palletizable for shipping. Case units may also contain totes, boxes, and / or containers of one or more individual items (generally referred to as break-pack items) that have been opened / released from their original packaging and placed in totes, boxes, and / or containers (collectively referred to as totes) with one or more other individual items of a mixed or common type at an order filling station.For example, it is noted that when incoming bundles or pallets (e.g., from a case unit manufacturer or supplier) arrive at the automated storage and retrieval system 100 for replenishment, the contents of each pallet may be uniform (e.g., each pallet holds a predetermined number of the same items, i.e., one pallet holds soup, another pallet holds cereal). As can be appreciated, the cases in such a pallet load may be generally similar, or in other words, homogenous cases (e.g., similar dimensions) and may have the same SKUs (otherwise, as previously mentioned, the pallet may be a “rainbow” pallet with layers formed of homogenous cases). Once the pallet exits the automated storage and retrieval system, with the cases or totes filled with replenishment orders, the pallet may contain any suitable number and combination of different case units (e.g., each pallet may hold different types of case units, i.e., the pallet holds a combination of canned soup, cereal, drink cartons, cosmetics, and household cleaners). The cases combined on a single pallet may have different dimensions and / or different SKUs.
[0018] Automated storage and retrieval system 100 may generally be described as a storage and retrieval engine 190 coupled to a palletizer 162. Now in more detail, and still referring to Figures 1A and 1B, automated storage and retrieval system 100 may be configured for installation in an existing warehouse structure or adapted to a new warehouse structure, for example. As previously mentioned, automated storage and retrieval system 100 shown in Figures 1A and 1B is representative and may include, for example, infeed and outfeed conveyors terminating in respective transfer stations 170, 160, lift module(s) 150A, 150B, a storage structure 130, and several autonomous guided vehicles 110. It is noted that storage and retrieval engine 190 is formed by at least storage structure 130 and autonomous guided vehicle 110 (and in some embodiments also by lift modules 150A, 150B, while in other embodiments lift modules 150A, 150B may form a vertical sequencer in addition to storage and retrieval engine 190, as described in U.S. Patent Application No. 17 / 091,265, filed November 6, 2020, and entitled "Pallet Building System with Flexible Sequencing," the entire disclosure of which is incorporated herein by reference). In alternative embodiments, automated storage and retrieval system 100 may include a robot or bot transfer station (not shown) that may provide an interface between autonomous guided vehicle 110 and lift module(s) 150A, 150B. The storage structure 130 may include multiple levels of storage rack modules, where each storage structure level 130L of the storage structure 130 includes a respective picking aisle 130A and a transfer deck 130B for transporting case units between any of the storage areas of the storage structure 130 and the shelves of the(s) lift module(s) 150A, 150B.In one embodiment, picking aisle 130A is configured to provide guided movement of autonomous guided vehicle 110 (such as along rail 130AR), while in other embodiments, picking aisle 130A is configured to provide unconstrained movement of autonomous guided vehicle 110 (e.g., picking aisle 130A is open and non-deterministic with respect to the guidance / movement of autonomous guided vehicle 110). Transfer deck 130B has an open, non-deterministic bot-supporting movement surface along which autonomous guided vehicle 110 moves under guidance and control provided by any suitable bot steering. In one or more embodiments, transfer deck 130B has multiple lanes between which autonomous guided vehicle 110 transitions freely to access picking aisle 130A and / or lift modules 150A, 150B. As used herein, "open and non-deterministic" indicates that the movement surface of the picking aisle and / or transfer deck has no mechanical constraints (such as guide rails) that limit the movement of the autonomous guided vehicle 110 to any given path along the movement surface.
[0019] Picking aisle 130A and transfer deck 130B also enable autonomous guided vehicle 110 to place case units CU in picking stock and retrieve ordered case units CU (and define various locations at which the bot performs its autonomous tasks, although any number of locations within the storage structure (e.g., deck, aisle, storage rack, etc.) could be one or more of the various locations). In alternative embodiments, each level may include a respective transfer station 140 that provides indirect case transfer between autonomous guided vehicle 110 and lift modules 150A, 150B. Autonomous guided vehicle 110 may be configured to place case units, such as the retail items described above, in picking stock in one or more storage structure levels 130L of storage structure 130, and then selectively retrieve ordered case units and ship the ordered case units, for example, to a store or other suitable location. The infeed transfer station 170 and the outfeed transfer station 160 may operate in conjunction with respective lift module(s) 150A, 150B to transfer case units CU bidirectionally to / from one or more storage structure levels 130L of the storage structure 130. While the lift modules 150A, 150B may be described as dedicated inbound lift module 150A and outbound lift module 150B, it is noted that in alternative embodiments, each of the lift modules 150A, 150B may be used for both inbound and outbound transfer of case units from the automated storage and retrieval system 100.
[0020] As can be appreciated, the automated storage and retrieval system 100 may include multiple infeed and outfeed lift modules 150A, 150B that are accessible, for example, by the autonomous guided vehicles 110 of the automated storage and retrieval system 100 (e.g., indirectly via transfer station 140 or via direct case transfer between the lift modules 150A, 150B and the autonomous guided vehicles 110) so that one or more uncontained case units (e.g., case unit(s) not held in a tray) or one or more contained case units (in a tray or tote) can be transferred from the lift modules 150A, 150B to each storage space on each level, and from each storage space on each level to any one of the lift modules 150A, 150B. Autonomous guided vehicle 110 may be configured to transfer cases CU (also referred to herein as case units) between storage space 130S (e.g., located in picking aisle 130A or other suitable storage space / case unit buffer located along transfer deck 130B) and lift modules 150A, 150B. Generally, lift modules 150A, 150B include at least one movable payload support that can move case unit(s) between infeed and outfeed transfer stations 160, 170 and respective levels of the storage space where the case unit(s) are stored and retrieved. The lift module(s) may have any suitable configuration, such as, for example, a reciprocating lift, or any other suitable configuration.The lift module(s) 150A, 150B may include any suitable controller (such as the control server 120 or other suitable controller coupled to the control server 120, the warehouse management system 2500, and / or the palletizer controllers 164, 164′) to form a sequencer or sorter in a manner similar to that described in U.S. patent application Ser. No. 16 / 444,592, filed June 18, 2019, and entitled “Vertical Sequencer for Product Order Fulfillment,” the disclosure of which is incorporated herein by reference in its entirety.
[0021] Automated storage and retrieval system 100 may include a control system comprising, for example, one or more control servers 120 communicatively connected to infeed and outfeed conveyors and transfer stations 170, 160, lift modules 150A, 150B, and autonomous guided vehicle 110 via a suitable communication and control network 180. Communication and control network 180 may have any suitable architecture, for example, incorporating various programmable logic controllers (PLCs), such as for commanding the operation of the automation of infeed and outfeed conveyors and transfer stations 170, 160, lift modules 150A, 150B, and other suitable systems. Control server 120 may include high-level programming to provide a case management system (CMS) that manages the case flow system. Network 180 may further include suitable communications to provide a bidirectional interface with autonomous guided vehicle 110. For example, autonomous guided vehicle 110 may include an on-board processor / controller 122. Network 180 may include a suitable two-way communication suite that enables autonomous guided vehicle controller 122 to request or receive commands from control server 120 to effect the desired transport of case units (e.g., placement in or removal from a storage location) and to transmit desired autonomous guided vehicle 110 information and data to control server 120, including autonomous guided vehicle 110 ephemeris, status, and other desired data. As seen in FIGS. 1A and 1B , control server 120 may further be connected to a warehouse management system 2500 for, for example, providing inventory management and customer order fulfillment information to control server 120's CMS-level programs. As previously mentioned, control server 120 and / or warehouse management system 2500 enable some degree of collaborative control of at least bots 110 via a user interface UI, as further described below.A suitable example of an automated storage and retrieval system arranged to hold and store case units is described in U.S. Pat. No. 9,096,375, issued August 4, 2015, the entire disclosure of which is incorporated herein by reference.
[0022] 1A, 1B, and 2, autonomous guided vehicle 110 includes a frame 200 with an integrated payload support or payload platform 210B (also referred to herein as a payload holder). Frame 200 has a front end 200E1 and a back end 200E2 that define a longitudinal axis LAX of autonomous guided vehicle 110. Frame 200 may be constructed of any suitable material (e.g., steel, aluminum, composite materials, etc.) and includes a case handling assembly 210 configured to handle cases / payloads to be transported by autonomous guided vehicle 110. Case handling assembly 210 includes any suitable payload platform 210B (also referred to herein as a payload bay or payload holder) where payloads are placed for transport and / or any suitable transfer arm 210A (also referred to herein as a payload handler) connected to the frame. Transfer arm 210A is configured to (autonomously) transfer a payload (such as a case unit CU) having a flat, non-deterministic seating surface that seats on payload platform 210B to / from payload platform 210B of autonomous guided vehicle 110, and to / from a storage location for payload CU within storage array SA (such as storage space 130S on storage shelf 555 (see FIG. 2), a shelf, buffer, transfer station, and / or any other suitable storage location), where storage location 130S within storage array SA is separate and distinct from transfer arm 210A and payload platform 210B. Transfer arm 210A is configured to extend in a lateral direction LAT and / or a vertical direction VER to transport payloads to / from payload platform 210B. Examples of suitable payload platforms 210B and transfer arms 210A and / or autonomous guided vehicles to which aspects of the disclosed embodiments may be applied are described in U.S. Pat. No. 1,107,801, entitled "Automated Bot with Transfer Arm," issued Aug. 3, 2021, and entitled "Materials-Handling System Using Autonomous Transfer and Transport," the entire disclosure of which is incorporated herein by reference.U.S. Patent No. 7,591,630, issued September 22, 2009, entitled "Materials-Handling System Using Autonomous Transfer and Transport Vehicles," U.S. Patent No. 7,991,505, issued August 2, 2011, entitled "Autonomous Transport Vehicle," U.S. Patent No. 9,561,905, issued February 7, 2017, entitled "Autonomous Transport Vehicle," U.S. Patent No. 9,082,112, issued July 14, 2015, entitled "Autonomous Transport Vehicle Charging System," U.S. Patent No. 9,850,079, issued December 26, 2017, entitled "Storage and Retrieval System Transport Vehicle," U.S. Patent No. 9,187,244, issued November 17, 2015, entitled "Bot Payload Alignment and Sensing," and U.S. Patent No. 9,187,244, issued November 17, 2015, entitled "Automated Bot Transfer Arm Drive" No. 9,499,338, issued November 22, 2016, entitled "Bot Having High Speed Stability System," U.S. Patent No. 8,965,619, issued February 24, 2015, entitled "Bot Position Sensing," U.S. Patent No. 9,008,884, issued April 14, 2015, entitled "Bot Position Sensing," U.S. Patent No. 8,425,173, issued April 23, 2013, entitled "Autonomous Transports for Storage and Retrieval Systems," and U.S. Patent No. 8,696,010, issued April 15, 2014, entitled "Suspension System for Autonomous Transports."
[0023] Frame 200 includes one or more idler wheels or casters 250 positioned adjacent to front end 200E1. Suitable examples of casters can be found in U.S. Patent Application No. 17 / 664,948 (Attorney Docket No. 1127P015753-US(PAR)), filed May 25, 2022, entitled "Autonomous Transport Vehicle with Synergistic Vehicle Dynamic Response," and U.S. Patent Application No. 17 / 664,838 (Attorney Docket No. 1127P015753-US(PAR)), filed May 26, 2021, entitled "Autonomous Transport Vehicle with Steering," the entire disclosures of which are incorporated herein by reference. Frame 200 also includes one or more drive wheels 260 positioned adjacent to back end 200E2. In other aspects, the positions of caster 250 and drive wheel 260 may be reversed (e.g., drive wheel 260 is located on front end 200E1 and caster 250 is located on back end 200E2). It is noted that in some aspects, autonomous guided vehicle 110 is configured to move with front end 200E1 leading the direction of movement or with back end 200E2 leading the direction of movement. In one aspect, casters 250A, 250B (substantially similar to caster 250 described herein) are positioned at respective front corners at front end 200E1 of frame 200, and drive wheels 260A, 260B (substantially similar to drive wheel 260 described herein) are positioned at respective back corners at back end 200E2 of frame 200 (e.g., support wheels are positioned at each of the four corners of frame 200), thereby allowing autonomous guided vehicle 110 to stably travel on transfer deck 130B and picking aisle 130A of storage structure 130.
[0024] Autonomous guided vehicle 110 includes a drive section 261D connected to frame 200, with drive wheels 260 that support autonomous guided vehicle 110 on running / rolling surface 284, and drive wheels 260 provide vehicle movement on running surface 284 and move autonomous guided vehicle 110 on running surface 284 within a facility (e.g., warehouse, store, etc.). Drive section 261D has at least a pair of traction drive wheels 260 (also referred to as drive wheels 260 (see drive wheels 260A, 260B)) straddling drive section 261D. The drive wheels 260 have a fully independent suspension 280 that couples each drive wheel 260A, 260B of at least one pair of drive wheels 260 to the frame 200 and is configured to maintain a substantially steady-state traction contact patch between at least one drive wheel 260A, 260B and a rolling / moving surface 284 (also referred to as an autonomous vehicle moving surface 284) over rolling surface transients (e.g., bumps, surface transitions, etc.). A suitable example of a fully independent suspension 280 can be found in U.S. Patent Application No. 17 / 664,948, entitled "Autonomous Transport Vehicle with Synergistic Vehicle Dynamic Response," filed May 25, 2022 (having Attorney Docket No. 1127P015753-US(PAR)), the entire disclosure of which was previously incorporated herein by reference.
[0025] Autonomous guided vehicle 110 includes a physical property sensor system 270 (also referred to as an autonomous navigation operation sensor system) connected to frame 200. Physical property sensor system 270 has electromagnetic sensors. Each of the electromagnetic sensors is responsive to an interaction or interface between an electromagnetic beam or field emitted or generated by the sensor and a physical property (e.g., of a storage structure or case unit CU, debris, or other ephemeral object), where the electromagnetic beam or field is perturbed by the interaction or interface with the physical property. The perturbation of the electromagnetic beam is detected by the electromagnetic sensor, resulting in the sensing of the physical property, where physical property sensor system 270 is configured to generate sensor data embodying at least one of vehicle navigation pose or location information (with respect to a storage and retrieval system or facility in which autonomous guided vehicle 110 operates) and payload pose or location information (with respect to storage location 130S or payload platform 210B).
[0026] Physical property sensor system 270 includes, by way of example only, one or more of laser sensor(s) 271, ultrasonic sensor(s) 272, barcode scanner(s) 273, position sensor(s) 274, line sensor(s) 275, a plurality of case sensors 278 (e.g., for detecting case units in payload berth 210B onboard vehicle 110 or on storage shelves offboard vehicle 110), arm proximity sensor(s) 277, vehicle proximity sensor(s) 278, or any other suitable sensors for detecting the position of vehicle 110 or payload (e.g., case unit CU). In some aspects, supplemental navigation sensor system 288 may form part of physical property sensor system 270. Suitable examples of sensors that may be included in the physical property sensor system 270 are described in U.S. Pat. No. 8,425,173, entitled "Autonomous Transport for Storage and Retrieval Systems," issued April 23, 2013; U.S. Pat. No. 9,008,884, entitled "Bot Position Sensing," issued April 14, 2015; and U.S. Pat. No. 9,946,265, entitled "Bot Having High Speed Stability," issued April 17, 2018, the entire disclosures of which are incorporated herein by reference.
[0027] The sensors of physical property sensor system 270 may be configured to provide autonomous guided vehicle 110 with, for example, awareness of its environment and external objects, as well as monitoring and control of internal subsystems. For example, the sensors may provide guidance information, payload information, or any other suitable information for use in operating autonomous guided vehicle 110.
[0028] Barcode scanner(s) 273 may be mounted in any suitable location on autonomous guided vehicle 110. Barcode scanner(s) 273 may be configured to provide an absolute position of autonomous guided vehicle 110 within storage structure 130. Barcode scanner(s) 273 may be configured to verify aisle references and locations on the transfer deck, for example, by reading barcodes located on the transfer deck, picking aisles, and transfer station floors to verify the location of autonomous guided vehicle 110. Barcode scanner(s) 273 may also be configured to read barcodes located on items stored on shelves 555.
[0029] Position sensor(s) 274 may be mounted in any suitable location on autonomous guided vehicle 110. Position sensor 274 may be configured, for example, to detect fiducial datum features (or count slats 520L of storage shelf 555) (e.g., see FIG. 5A ) to determine the location of vehicle 110 relative to shelves in picking aisle 130A (or transfer deck 130B or a buffer / transfer station located adjacent to lift 150). The fiducial datum information may be used by controller 122, for example, to correct vehicle odometry and stop autonomous guided vehicle 110 with support tines 210AT of transfer arm 210A positioned for insertion into the spaces between slats 520L (e.g., see FIG. 5A ). In one exemplary embodiment, the vehicle 110 may include position sensors 274 at the driving (rear) end 200E2 and the driven (front) end 200E1 of the autonomous guided vehicle 110 to enable reference datum detection regardless of which end of the autonomous guided vehicle 110 is facing in the direction the autonomous guided vehicle 110 is moving.
[0030] The line sensor 275 may be any suitable sensor mounted on the autonomous guided vehicle 110 in any suitable location, such as, by way of example only, on the frame 200 positioned adjacent the driving (rear) end 200E2 and the driven (front) end 200E1 of the autonomous guided vehicle 110. By way of example only, the line sensor 275 may be a diffuse infrared sensor. The line sensor 275 may be configured to detect guide lines 900 (see FIG. 1B ), for example, provided on the floor of the transfer deck 130B. The autonomous guided vehicle 110 may be configured to follow the guide lines when traveling on the transfer deck 130B and to define the end of a turn as the vehicle transitions onto or off of the transfer deck 130B. The line sensor 275 may also enable the vehicle 110 to detect an index reference to determine absolute localization, where the index reference is generated by the crossed guide lines 119 (see FIG. 1B ).
[0031] The case sensor 276 may include a case overhang sensor and / or other suitable sensor configured to detect the location / pose of the case unit CU within the payload platform 210B. The case sensor 276 may be any suitable sensor positioned on the vehicle such that the field of view(s) of the sensor(s) spans the payload platform 210B adjacent the top surface of the support tines 210AT (see FIGS. 3A and 3B). The case sensor 276 may be positioned at an edge of the payload platform 210B (e.g., adjacent the transport opening 1199 of the payload platform 210B) to detect a case unit CU that is at least partially extending outside the payload platform 210B.
[0032] Arm proximity sensor 277 may be mounted to autonomous guided vehicle 110 in any suitable location, such as, for example, on transfer arm 210A. Arm proximity sensor 277 may be configured to detect objects around transfer arm 210A and / or support tines 210AT of transfer arm 210A as transfer arm 210A is raised / lowered and / or support tines 210AT are extended / retracted.
[0033] Laser sensor 271 and ultrasonic sensor 272 may be configured to enable autonomous guided vehicle 110 to locate itself with respect to each case unit forming a load carried by autonomous guided vehicle 110 before the case unit is picked, for example, from storage shelf 555 and / or lift 150 (or any other location suitable for retrieving a payload). Laser sensor 271 and ultrasonic sensor 272 also enable the vehicle to position itself with respect to empty storage locations 130S and place case units in those empty storage locations 130S. Laser sensor 271 and ultrasonic sensor 272 also enable autonomous guided vehicle 110 to verify that a storage space (or other load placement location) is empty before a load carried by autonomous guided vehicle 110 is placed, for example, in storage space 130S. In one example, laser sensor 271 may be mounted on autonomous guided vehicle 110 at a suitable location for detecting the edge of an item being transferred to (or from) autonomous guided vehicle 110. Laser sensor 271 may operate in conjunction with, for example, retroreflective tape (or other suitable reflective surface, coating, or material) disposed, for example, on the back of shelf 555, to enable the sensor to "see" all the way to the back of storage shelf 555. The reflective tape disposed on the back of the storage shelf renders laser sensor 1715 substantially unaffected by the color, reflectivity, roundness, or other suitable characteristics of the items located on shelf 555. Ultrasonic sensor 272 may be configured to measure the distance from autonomous guided vehicle 110 to a first item within a predetermined storage area of shelf 555 to enable autonomous guided vehicle 110 to determine the picking depth (e.g., the distance support tine 210AT travels into shelf 555 to pick item(s) from shelf 555). One or more of laser sensor 271 and ultrasonic sensor 272 may enable detection of case orientation (e.g., tilt of a case within storage shelf 555) by, for example, measuring the distance between autonomous guided vehicle 110 and the front of the case unit to be picked when autonomous guided vehicle 110 stops adjacent to the case unit to be picked.The case sensor may allow verification of the placement of a case unit, for example, on a storage shelf 555, for example, by scanning the case unit after placing it on the shelf.
[0034] Vehicle proximity sensors 278 may also be disposed on frame 200 to determine the location of autonomous guided vehicle 110 within picking aisle 130A and / or relative to lift 150. Vehicle proximity sensors 278 are disposed on autonomous guided vehicle 110 to detect targets or position-determining features disposed on rail 130AR along which vehicle 110 travels through picking aisle 130A (and / or on the walls of transfer area 195 and / or on lift 150 access locations). The positions of the targets on rail 130AR are at known locations that form incremental or absolute encoders along rail 130AR. Vehicle proximity sensors 278 detect the targets and provide sensor data to controller 122, which then determines the position of autonomous guided vehicle 110 along picking aisle 130A based on the detected targets.
[0035] The sensors of physical characteristic sensing system 270 are communicatively coupled to controller 122 of autonomous guided vehicle 110. As described herein, controller 122 is operably connected to drive section 261D and / or transfer arm 210A. Controller 122 is configured to determine the vehicle's pose and location (e.g., in up to six degrees of freedom: X, Y, Z, Rx, Ry, Rz) from information from physical characteristic sensor system 270 to provide independent guidance of autonomous guided vehicle 110 through storage and retrieval facility / system 100. Controller 122 is also configured to determine the pose and location (onboard or offboard autonomous guided vehicle 110) of payloads (e.g., case units CU) from information from physical characteristic sensor system 270 to provide independent underpicking (e.g., lifting case units CU from under case units CU) and placement of payloads CU to / from storage location 130S, and independent underpicking and placement of payloads CU on payload platform 210B.
[0036] 1A, 1B, 2, 3A, and 3B, as described above, autonomous guided vehicle 110 includes a supplemental or auxiliary navigation sensor system 288 connected to frame 200. Supplemental navigation sensor system 288 supplements physical property sensor system 270. Supplemental navigation sensor system 288 is, at least in part, a vision system 400 with a camera positioned to capture image data that informs at least one of the vehicle navigation pose or location (relative to the structure or facility of the storage and retrieval system in which vehicle 110 operates) and the payload pose or location (relative to the storage location or payload platform 210B), which supplements the information of physical property sensor system 270. It is noted that the term "camera" as used herein refers to a still or video imaging device including one or more of a two-dimensional camera, a two-dimensional camera with RGB (red, green, blue) pixels, a three-dimensional camera with XYZ+A definition (where XYZ is the camera's three-dimensional reference frame and A is one of radar return intensity, time-of-flight stamp, or other distance-determining stamp / indicator), and an RGB / XYZ camera that contains both RGB and three-dimensional coordinate system information, non-limiting examples of which are provided herein.
[0037] 2, 3A, and 3B, the vision system 400 includes one or more of case unit monitoring cameras 410A, 410B, forward navigation cameras 420A, 420B, rear navigation cameras 430A, 430B, one or more three-dimensional imaging systems 440A, 440B, one or more case edge detection sensors 450A, 450B, one or more traffic monitoring cameras 460A, 460B, and one or more out-of-plane (e.g., upward-facing or downward-facing) localization cameras 477A, 477B (it is noted that the downward-facing cameras supplement the line following sensor 275 of the physical property sensor system 270 and may provide a wider field of view than the line following sensor 275, thereby guiding / navigating the vehicle 110 to bring the guideline 900 (see FIG. 1B) back into the field of view of the line following sensor 275 if the vehicle path deviates from the guideline 900 and causes the guideline 900 to move out of the field of view of the line following sensor 275). Images (still images and / or dynamic video images) from the various vision system 400 cameras are requested by controller 122 from vision system controller 122VC as needed for the task of any given autonomous guided vehicle 110. For example, images are acquired by controller 122 from at least one or more of forward and rear navigation cameras 420A, 420B, 430A, 430B to provide navigation of autonomous guided vehicle 110 along transfer deck 130B and picking aisle 130A.
[0038] The forward navigation cameras 420A, 420B may be paired to form a stereo camera system, and the rear navigation cameras 430A, 430B may be paired to form another stereo camera system. Referring to Figures 2 and 3A, the forward navigation cameras 420A, 420B are any suitable cameras configured to provide object detection and ranging. The forward navigation cameras 420A, 420B may be positioned on opposite sides of the longitudinal centerline LAXCL of the autonomous guided vehicle 110 and may be spaced apart by any suitable distance such that the forward fields of view 420AF, 420BF provide stereoscopic vision for the autonomous guided vehicle 110. Forward navigation cameras 420A, 420B may be any suitable high-resolution or low-resolution video cameras (wherein video images including greater than approximately 480 vertical scan lines and captured at greater than approximately 50 frames per second are considered high-resolution), time-of-flight cameras, laser ranging cameras, or any other suitable cameras configured to provide object detection and ranging for navigating the autonomous vehicle along transfer deck 130B and picking aisle 130A. Rear navigation cameras 430A, 430B may be generally similar to the forward navigation cameras. Forward navigation cameras 420A, 420B and rear navigation cameras 430A, 430B provide for navigation of autonomous guided vehicle 110 with obstacle detection and avoidance (either end 200E1 of autonomous guided vehicle 110 is at the beginning or end of the direction of travel), as well as for localization of the autonomous guided vehicle within storage and retrieval system 100. Localization of the autonomous guided vehicle 110 may be provided by one or more of the forward navigation cameras 420A, 420B and rear navigation cameras 430A, 430B through detection of guidelines on the moving / rolling surface 284 and / or through detection of appropriate storage structures, including but not limited to storage rack (or other) structures. The line detection and / or storage structure detection may be compared to floor map and structure information of the vision system controller 122VC (e.g., stored in memory of or accessible by the vision system controller 122VC).The forward navigation cameras 420A, 420B and the rear navigation cameras 430A, 430B may also send a signal to the controller 122 (including or via the vision system controller 122VC) when an object approaches the autonomous transport vehicle 110 (when the autonomous transport vehicle 110 is stopped or in operation) so that the autonomous transport vehicle 110 can maneuver (e.g., on a non-deterministic rolling surface of the transfer deck 130B or within the picking aisle 130A (which may have a deterministic or non-deterministic rolling surface)) to avoid the approaching object (e.g., another autonomous transport vehicle, a case unit, or other temporary object within the storage and retrieval system 100).
[0039] The forward navigation cameras 420A, 420B and rearward navigation cameras 430A, 430B may also provide a convoy of vehicles 110 along a picking aisle 130A or a transfer deck 130B, where one vehicle 110 follows another vehicle 110A at a predetermined fixed distance. As an example, FIG. 1B illustrates a convoy of three vehicles 110, where one vehicle closely follows another vehicle at a predetermined fixed distance.
[0040] As another example, the controller 122 may acquire images from one or more of the three-dimensional imaging systems 440A, 440B, the case edge detection sensors 450A, 450B, and the case unit monitoring cameras 410A, 410B to effect case handling by the vehicle 110. Still referring to Figures 2 and 3A, the one or more case edge detection sensors 450A, 450B are any suitable sensors, such as laser measurement sensors, configured to scan the shelves of the storage and retrieval system 100 to verify whether the shelves are clear for placing a case unit CU or to verify the size and position of the case unit CU before picking it. Although one case edge detection sensor 450A, 450B is illustrated on each side of the centerline CLPB (see FIG. 3A ) of the payload platform 210B, more or less than two case edge detection sensors may be positioned at any suitable location on the autonomous guided vehicle 110 so that the vehicle 110 can pass and scan the case unit CU, with the front end 200E1 leading the direction of vehicle travel or the rear end / back end 200E2 leading the direction of vehicle travel. It is noted that case handling includes picking and placing case units from a case unit holding location (such as for locating the case unit, verifying the case unit, and verifying the placement of the case unit within the payload platform 210B and / or at a case unit holding location such as a storage shelf or buffer location).
[0041] Images from the out-of-plane localization cameras 477A, 477B may be acquired by the controller 122 to provide navigation for the autonomous guided vehicle 110 and / or may provide data (e.g., image data) that supplements the localization / navigation data from one or more of the front and rear navigation cameras 420A, 420B, 430A, 430B. Images from one or more traffic monitoring cameras 460A, 460B may be acquired by the controller 122 to provide movement transitions for the autonomous guided vehicle 110 from the picking aisle 130A to the transfer deck 130B (e.g., the entry onto the transfer deck 130B and the merging of the autonomous guided vehicle 110 with other autonomous guided vehicles traveling along the transfer deck 130B).
[0042] One or more out-of-plane (e.g., upward-facing or downward-facing) positioning cameras 477A, 477B are positioned on the frame 200 of the autonomous guided vehicle 110 to sense / detect position references (e.g., position marks (e.g., barcodes), lines 900 (see FIG. 1B), etc.) located on the ceiling of the storage and retrieval system or on the rolling surface 284 of the storage and retrieval system. The position references have known positions within the storage and retrieval system and may provide unique identification marks / patterns that are recognized by the vision system controller 122VC (e.g., processed data obtained from the positioning cameras 477A, 477B). Based on the detected position references, the vision system controller 122VC compares the detected position references with known position references (e.g., stored in the memory of or accessible to the vision system controller 122VC) to determine the location of the autonomous guided vehicle 110 within the storage structure 130.
[0043] The one or more traffic monitoring cameras 460A, 460B are positioned on the frame 200 such that their respective fields of view 460AF, 460BF are oriented sideways in a lateral direction LAT1. While the one or more traffic monitoring cameras 460A, 460B are illustrated adjacent the transfer opening 1199 of the transfer bed 210B (e.g., on the picking side from which the arm 210A of the autonomous guided vehicle 110 extends), in other embodiments, the traffic monitoring cameras may be positioned on the non-picking side of the frame 200 such that their fields of view are oriented sideways in a direction LAT2. The traffic monitoring cameras 460A, 460B provide for the autonomous merging of the autonomous guided vehicle 110, for example, as it exits the picking aisle 130A or lift transfer area 195 onto the transfer deck 130B (see FIG. 1B ). For example, autonomous guided vehicle 110V leaving lift transfer area 195 (FIG. 1B) detects autonomous guided vehicle 110T traveling along transfer deck 130B. Controller 122 then autonomously strategizes its merge onto the transfer deck (e.g., entering the transfer deck ahead of or behind autonomous guided vehicle 110T, accelerating onto the transfer deck based on the speed of the approaching vehicle 110T, etc.) based on information (e.g., distance, speed, etc.) of autonomous guided vehicle 110T collected by traffic monitoring cameras 460A, 460B and communicated to vision system controller 122VC for processing.
[0044] The case unit monitoring cameras 410A, 410B are any suitable high-resolution or low-resolution video cameras (wherein video images including greater than approximately 480 vertical scan lines and captured at greater than approximately 50 frames per second are considered high-resolution). The case unit monitoring cameras 410A, 410B are positioned relative to one another to form a stereo vision camera system configured to monitor the entry and exit of case units CU into and out of the payload platform 210B. The case unit monitoring cameras 410A, 410B are coupled to the frame 200 in any suitable manner and are focused on at least the payload platform 210B. In one or more embodiments, the case unit monitoring cameras 410A, 410B are coupled to the transfer arm 210A to move with the transfer arm 210A in the direction LAT (such as when picking and placing a case unit CU), and are positioned to be focused on the payload platform 210B and support tines 210AT of the transfer arm 210A.
[0045] 5A, the case unit monitoring cameras 410A, 410B at least partially provide one or more of case unit determination, case unit location identification, case unit position verification, and verification of case unit positioning features (e.g., positioning blade 471 and pusher 470) and case transfer features (e.g., tines 210AT, puller 472, and payload platform floor 473). For example, the case unit monitoring cameras 410A, 410B detect one or more of case unit lengths CL, CL1, CL2, CL3, case unit heights CH1, CH2, CH3, and case unit yaw YW (e.g., relative to the extension / retraction direction LAT of transfer arm 210A). Data from case handling sensors (e.g., described above) can also provide the location / position of the pusher 470, puller 472, and positioning blade 471, such as when payload platform 210B is empty (e.g., not holding a case unit).
[0046] The case unit monitoring cameras 410A, 410B are also configured, using the vision system controller 122VC, to provide for the determination of a front case center point FFCP relative to a reference position of the autonomous guided vehicle 110 (e.g., in the X, Y, and Z directions with the case unit positioned on an off-board shelf or other holding area of the vehicle 110). The reference position of the autonomous guided vehicle 110 may be defined by one or more positioning planes of the payload platform 210B or the centerline CLPB of the payload platform 210B. For example, the front case center point FFCP may be determined along a longitudinal axis LAX (e.g., in the Y direction) relative to the centerline CLPB of the payload platform 210B (FIG. 3A). The front case center point FFCP may be determined along a vertical axis VER (e.g., in the Z direction) relative to the case unit support plane PSP of the payload platform 210B (FIGS. 3A and 3B—formed by one or more of the tines 210AT of the transfer arm 210A and the payload platform floor 473). The front case center point FFCP may be determined along the horizontal axis LAT (e.g., in the X direction) relative to the positioning face surface JPP of the pusher 470 ( FIG. 3B ). Determining the front case center point FFCP of a case unit CU located on a storage shelf 555 (see FIGS. 3A and 4A ) or other case unit holding location may, by way of non-limiting example, provide location identification of the autonomous guided vehicle 110 relative to the case unit CU to be picked, mapping of the case unit's location within the storage structure (e.g., in a manner similar to that described in U.S. Pat. No. 9,242,800, issued January 26, 2016, entitled “Storage and retrieval system case unit detection,” the entire disclosure of which is incorporated herein by reference), and / or accuracy of picking and placing relative to other case units on the storage shelf 555 (e.g., to maintain a predetermined gap size between case units).Determining the front case center point FFCP also results in a comparison of the “real world” environment in which the autonomous guided vehicle 110 is operating with a virtual model 400VM of that operating environment, thereby allowing the controller 122 of the autonomous guided vehicle 110 to make a substantially direct comparison of what the vision system 400 “sees” with what the autonomous guided vehicle 110 expects to “see” based on a simulation of the storage and retrieval system structure, in a manner similar to that described in U.S. patent application Ser. No. 17 / 804,026, filed May 25, 2022, entitled “Autonomous Transport Vehicle with Vision System” (having Attorney Docket No. 1127P016037-US(PAR)), the entire disclosure of which is incorporated herein by reference. 5A , the object (case unit) and its characteristics determined by the vision system controller 122VC are superimposed (combined, overlaid) with the virtual model 400VM to enhance resolution of the object's pose relative to the facility or global reference frame GREF (see FIG. 2 ) by up to six degrees of resolution freedom. As can be appreciated, alignment of the vision system 400's cameras with the global reference frame GREF enables enhanced resolution of the vehicle 110's pose and / or location relative to both the global reference (the facility's features rendered in the virtual model 400VM) and the imaged object. More specifically, any discrepancies or anomalies in object position revealed and identified upon superimposing the object image with the virtual model 400VM (e.g., edge spacing between the reference edges of the case unit or tilt or skew of the case unit relative to the rack slats 520L of the virtual model 400VM) that exceed a predetermined nominal threshold represent an incorrect pose of one or more of the case, rack, and / or vehicle 110. Determining whether the error is due to one or more poses / positions of the case, rack, or vehicle 110 is determined through a comparison with pose data from sensors 270 and supplemental navigation sensor system 288 .
[0047] As one example of the aforementioned resolution enhancement, if a case unit placed on a shelf imaged by vision system 400 is rotated compared to a case unit and virtual model 400VM placed next to it on the same shelf (also imaged by the vision system), vision system 400 may determine that the case is tilted and provide enhanced case position information to controller 122 to operate and position transfer arm 210A to pick the case based on the enhanced resolution of the case's pose and location. As another example, it is noted that if the edge of the case is offset from the edge of slat 520L (see FIGS. 4A-4C) beyond a predetermined threshold, vision system 400 may generate a position error for the case, but if the offset is within the threshold, supplemental information from supplemental navigation sensor system 288 will provide enhanced pose / location resolution (e.g., an offset approximately equal to the determined pose / location of the case relative to the frame of slat 520L and transfer arm 210A of payload platform 210B of vehicle 110). It is further noted that while the vision system may generate a case position error if only one case is tilted / offset relative to the edge of the slat 520L, if two or more side-by-side cases are determined to be tilted relative to the edge of the slat 520L, the vision system may generate a pose error for the vehicle 110 and cause a repositioning of the vehicle 110 (e.g., correcting the position of the vehicle 110 based on an offset determined from supplemental information from the supplemental navigation sensor system 288) or send a service message to the operator (e.g., where the vision system 400 provides a "dashboard camera" collaboration mode (as described herein) that provides remote control of the vehicle 110 by the operator with images (still images and / or real-time video) from the vision system being communicated to the operator to effect remote control operation). The vehicle 110 may be stopped (e.g., not traveling down the picking aisle 130A or transfer deck 130B) until the operator initiates remote control of the vehicle 110.
[0048] Case unit monitoring cameras 410A, 410B may also provide feedback regarding the location of the case unit's positioning features and case transfer features of autonomous guided vehicle 110, for example, before and / or after picking / placing the case unit from a storage shelf or other storage location (e.g., to verify the location / position of the positioning features and case transfer features to result in picking / placing the case unit by transfer arm 210A without obstruction of the transfer arm). For example, as described above, case unit monitoring cameras 410A, 410B have a field of view that encompasses payload platform 210B. The vision system controller 122VC is configured to receive sensor data from the case unit monitoring cameras 410A, 410B and, using any suitable image recognition algorithm stored in the memory of or accessible to the vision system controller 122VC, determine the position of the pusher 470, position adjustment blade 471, puller 472, tines 210AT, and / or any other features of the payload platform 210B that engage with a case unit held on the payload platform 210B. The positions of the pusher 470, position adjustment blade 471, puller 472, tines 210AT, and / or other features of the payload platform 210B may be utilized by the controller 122 to verify the respective positions of the pusher 470, position adjustment blade 471, puller 472, tines 210AT, and / or other features of the payload platform 210B determined by the motor encoders or other respective position sensors, although in some embodiments the positions determined by the vision system controller 122VC may be utilized as a redundancy measure in the event of an encoder / position sensor failure.
[0049] The adjusted position of the case unit CU within the payload platform 210B can also be verified by the case unit monitoring cameras 410A, 410B. For example, referring also to FIG. 3C , the vision system controller 122VC is configured to receive sensor data from the case unit monitoring cameras 410A, 410B and, using any suitable image recognition algorithm stored in the memory of or accessible by the vision system controller 122VC, determine the position of the case unit in the X, Y, and Z directions, for example, relative to one or more of the centerline CLPB of the payload platform 210B, the reference / home position of the positioning surface JPP ( FIG. 3B ) of the pusher 470, and the case unit support surface PSP ( FIGS. 3A and 3B ). Here, the position determination of the case unit CU within the payload platform 210B affects at least its placement accuracy relative to other case units on the storage shelf 555 (e.g., to maintain a predetermined gap size between the case units).
[0050] 2, 3A, 3B, and 5, the one or more three-dimensional imaging systems 440A, 440B may include any suitable three-dimensional imager(s), including, but not limited to, a time-of-flight camera, an imaging radar system, or a light detection and ranging (LIDAR). The one or more three-dimensional imaging systems 440A, 440B may provide, for example, improved localization of the autonomous guided vehicle 110 relative to a global reference frame GREF (see FIG. 2) of the storage and retrieval system 100. For example, one or more three-dimensional imaging systems 440A, 440B, together with the vision system controller 122VC, may provide a determination of the size (e.g., height and width) and front case center point FFCP (e.g., in the X, Y, and Z directions) of the front face (i.e., front surface) of the case unit CU relative to a reference position of the autonomous guided vehicle 110 without relying on the shelf supporting the case unit CU (e.g., one or more three-dimensional imaging systems 440A, 440B provide a position of the case unit CU (which position of the case unit CU within the automated storage and retrieval system 100 is defined in the global reference frame GREF) without reference to the shelf supporting the case unit CU, and provide a determination of whether the case unit is supported on a shelf via determining the invariant characteristics of the shelf of the case unit). Here, the determination of the front surface and case center point FFCP also results in a comparison of the "real world" environment in which the autonomous guided vehicle 110 is operating with the virtual model 400VM, thereby allowing the controller 122 of the autonomous guided vehicle 110 to make a substantially direct comparison of what the vision system 400 "sees" with what the autonomous guided vehicle 110 expects to "see" based on a simulation of the storage and retrieval system structure, as described in U.S. Patent Application No. 17 / 804,026, filed May 25, 2022, entitled "Autonomous Transport Vehicle with Vision System" (having Attorney Docket No. 1127P016037-US(PAR)), the entire disclosure of which has previously been incorporated by reference herein.Image data obtained from one or more three-dimensional imaging systems 440A, 440B may supplement and / or enhance image data from cameras 410A, 410B when the data from cameras 410A, 410B is incomplete or missing. Here, object detection and localization relative to the pose of autonomous guided vehicle 110 in the global reference frame GREF may be determined with high accuracy and reliability by one or more three-dimensional imaging systems 440A, 440B, although in other aspects object detection and localization may be provided by one or more of physical property sensor system 270 and / or wheel encoders / inertial sensors of autonomous guided vehicle 110.
[0051] As illustrated in FIG. 5, one or more three-dimensional imaging systems 440A, 440B have respective fields of view extending beyond the payload platform 210B in a substantially LAT direction such that each three-dimensional imaging system 440A, 440B is positioned to detect case units CU adjacent to but outside the payload platform 210B (such as case units CU arranged in one or more rows extending along the length of the picking aisle 130A (see FIG. 5A), or substrate buffer / transfer stations arranged along the transfer deck 130B (similar in configuration to the storage rack 599 and its shelves 555 arranged along the picking aisle 130A). The fields of view 440AF, 440BF of each three-dimensional imaging system 440A, 440B encompass spatial volumes 440AV, 440BV extending to the picking range height 670 of the autonomous guided vehicle 110 (e.g., the range / height in the direction VER (Figure 2) at which the arm 210A can move to pick / place case units on shelves or stacked shelves accessible from the common rolling surface 284 on which the autonomous guided vehicle 110 rides (e.g., on the transfer deck 130B or the picking aisle 130A (see Figure 2)).
[0052] The vision system 400 may also cooperate with the operator to provide operational control of the autonomous guided vehicle 110. The vision system 400 provides data (images) that is recorded by the vision system controller 122VC, and either (a) the vision system 400 determines the characteristics of the information (which is then provided to the controller 122), or (b) the information is sent to the controller 122 without being characterized (objects within predetermined criteria), which is characterized by the controller 122. In either case (a) or (b), it is the controller 122 that determines the selection to switch to a collaborative state. After switching, collaborative operation is effected by a user accessing the vision system 400 via the vision system controller 122VC and / or the controller 122 via the user interface UI (see FIG. 10 ). However, in its simplest form, the vision system 400 can be thought of as providing a collaborative mode of operation for the autonomous guided vehicle 110. Here, the vision system 400 complements the autonomous navigation / operation sensor system 270 to provide coordinated identification and mitigation of objects / hazards 299 (see FIG. 3A; such objects / hazards include fluids, cases, solid debris, etc.) intruding onto the driving / rolling surface 284, as described, for example, in U.S. patent application Ser. No. 17 / 804,026, filed May 25, 2022, entitled "Autonomous Transport Vehicle with Vision System" (having attorney docket number 1127P016037-US(PAR)), the entire disclosure of which was previously incorporated by reference herein.
[0053] In one aspect, an operator may select or switch control of the autonomous guided vehicle from automatic operation to collaborative operation (e.g., via a user interface UI) (e.g., the operator remotely controls the operation of the autonomous guided vehicle 110 via the user interface UI). For example, the user interface UI may include a capacitive touchpad / screen, joystick, tactile screen, or other input device that communicates kinematic direction commands (e.g., turn, accelerate, decelerate, etc.) from the user interface UI to the autonomous guided vehicle 110 to provide operator control input in a collaborative operation mode of the autonomous guided vehicle 110. For example, the vision system 400 provides a “dashboard camera” (or dash camera) that transmits video and / or still images from the autonomous transport vehicle 110 to an operator (via a user interface UI) to enable remote operation or monitoring of an area relative to the autonomous transport vehicle 110, in a manner similar to that described in U.S. Patent Application No. 17 / 804,026, filed May 25, 2022, entitled “Autonomous Transport Vehicle with Vision System” (having Attorney Docket No. 1127P016037-US(PAR)), the entire disclosure of which was previously incorporated herein by reference, providing a “dashboard camera” (or dash camera).
[0054] 10 , in a collaborative operation mode of the autonomous guided vehicle 110, frames from an image data stream 1000 are presented to an operator. In one or more aspects, the frames from the image data stream 1000 are augmented images (as described herein), where image augmentation provides the operator with at least identification of objects within each frame. The user interface provides a frame crop object selection 1000C, where a portion of a frame from the image data stream 1000 is selected for operator manipulation. The frame crop object selection 1000C may include one or more object windows 1010, 1020, 1030, where these object windows 1010, 1020, 1030 are configured to provide the operator with selections that result in the selection of an object and the display of data related to the selected object. For example, an object selector 1010 may be presented to the operator along with a display of the frame crop selection 1000C. The object selector 1010 may include a drop-down menu (or other suitable interface) that results in operator selection of an object shown in the frame crop selection 1000C (e.g., in this example, the forks and pusher / puller of the vehicle 110 and the case unit CU are illustrated and presented in the object selector for operator selection). Selection of an object may present a bounding box around the selected object to identify the object within the frame crop selection 1000C. Selection of an object also results in presentation of object information 1020 (e.g., in this example, a "partial" object is selected and the information presented for the "partial" object informs the operator that the "partial" object represents or otherwise identifies a partial identification of the case unit CU by the vision system controller 122VC). An action list 1030 may also be presented in the user interface UI, where the action list depends on the object selected.For example, once a case unit CU is partially identified, the operator may instruct the autonomous guided vehicle to pick or not pick the case unit CU via action list 1030. In other aspects, user interface UI may present any suitable information regarding the operation of autonomous guided vehicle(s) 110 to the operator that results in coordinated operation of autonomous guided vehicle(s) 110.
[0055] 1A , as described above, autonomous guided vehicle 110 is provided with vision system 400 having an architecture based on camera pairs (e.g., camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, 477A and 477B, etc.) arranged for stereoscopic or binocular object detection and depth determination (e.g., via utilization of respective disparity / depth maps from recorded video frames / images captured by each camera). Object detection and depth determination results in localization of autonomous guided vehicle 110 relative to objects (e.g., at least case storage locations, such as shelves or lifts, and the tops of cases to be picked), although, as described herein, stereoscopic vision from camera pairs is not always available (i.e., stereoscopic vision is impaired). Thus, the disclosed embodiments provide the controller 122 of the autonomous guided vehicle 110 with both a computer vision object detection and localization protocol that utilizes the stereoscopic vision of the vision system 400, and a machine learning detection and localization protocol. The computer vision object detection and localization protocol utilizes the vision system controller 120VC to determine a disparity map or disparity map of at least one computer vision parameter for various features on image frames acquired by the stereo cameras. The machine learning detection and localization protocol employs machine learning, where at least one machine learning model ML and at least one artificial neural network ANN provide robust object determination and localization via monocular vision (utilizing one (i.e., unobstructed) camera of the stereo camera pair) with detection reliability comparable to unobstructed binocular vision.
[0056] As mentioned above and also referring to FIG. 2 , the artificial neural network ANN is included in the deep conductor module DC. The deep conductor module DC is communicatively coupled to the vision controller 122VC. The deep conductor module DC (which may be referred to as a deep learning graphics processing unit) is located onboard the autonomous guided vehicle 110 and is another control module of the controller 122. Here, a shared memory SH is located onboard the autonomous guided vehicle 110. The shared memory SH is communicatively connected to the vision system 400 and the vision system controller 122VC to receive video stream data imaging of objects in the logistics space from the vision system 400 (such as provided in substantially real time via a dash camera operating mode of the vision system 400 described herein and / or provided by a cache operating mode in which video is stored in the shared memory for on-demand searching). The shared memory SH includes non-transitory image frame generation computer program code configured to generate image frames (e.g., such as those illustrated in FIGS. 9A and 9C) from video stream data imaging, and these image frames are communicated to the deep conductor module DC to effect selection of one of a computer vision (detection / localization) protocol and a machine learning (detection / localization) protocol, while in other aspects the non-transitory image frame generation computer program code is included in the deep conductor module DC. The non-transitory image frame generation computer program code may be, for example, OpenCV Mat object / image processing code or any other suitable code for generating image frames from video stream data imaging.
[0057] In other aspects, deep conductor module DC may be located remotely from autonomous guided vehicle 110 (such as with control server 120 (see FIG. 1A) or other remotely located computer / server) or communicatively connected to autonomous guided vehicle 110 by network 180 using any suitable communication protocol, such as, for example, a server / client socket-based application. When deep conductor DC is located remotely from autonomous guided vehicle 110, the server / client socket-based application is structured such that media server MS is located onboard autonomous guided vehicle 110. Media server MS is communicatively connected to vision system 400 and vision system controller 122VC to receive video stream data imaging of objects in the logistics space from vision system 400 (such as provided in substantially real time by and via a dash camera operational mode of vision system 400 described herein and / or provided by a cache operational mode in which video is stored in shared memory for on-demand searching). Here, media server MS is communicatively connected to at least one camera (such as those described herein for autonomous guided vehicle 110) and records (in any suitable memory) video stream data from the at least one camera. Although media server MS is described as interfacing with vision system controller 122VC (e.g., vision system controller 122VC is located onboard autonomous guided vehicle 110), media server MS may also interface (e.g., via network 180) with any suitable controller located remotely from autonomous guided vehicle 110 (such as control server 120 or warehouse management system 2500).The media server MS includes non-transient image frame generation computer program code that generates image frames (e.g., such as those illustrated in FIGS. 9A and 9C) from the video stream data imaging, and these image frames are communicated to the deep conductor module DC to effect selection of one of a computer vision (detection / localization) protocol and a machine learning (detection / localization) protocol. The media server MS is configured to generate the image frames in any suitable manner, such as using OpenCV Mat objects / image processing as described above.
[0058] Both the shared memory SH and the media server MS are configured to convert the video stream data imaging into image frames and remove image frames that are similar to one another. Removal of similar (e.g., substantially duplicate) image frames reduces image processing time of the deep conductor DC when analyzing received image frames and reduces data transport traffic between the shared memory SH and the deep conductor DC and between the media server MS and the deep conductor DC. The shared memory SH and the media server MS are configured, via their respective non-temporal image frame generation computer program code, to utilize any suitable similarity metric (e.g., a structural similarity index) that calculates the structural similarity of the image frames. A suitable example of a structural similarity index for the removal of duplicate image frames can be found, for example, in Z. Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,” IEEE Transaction on Image Processing, Vol. 13, No. 4, pp. 600-612, April 2004 (referred to herein as “Wang”), the entire disclosure of which is incorporated herein by reference. This similarity metric is utilized by non-transient image frame generation computer program code in each of the shared memory SH and the media server MS to remove similar images based on a predetermined similarity index threshold, such as that described in Wang. A predetermined similarity index threshold is set for each camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B to generate a similarity index and remove similar images when converting video files from each respective camera into image files for each respective camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B.
[0059] In other aspects, the video stream data imaging is streamed in any suitable manner (such as via dash camera operation) to a remotely located Deep Conductor module DC (located in a remotely located controller as described herein), where the Deep Conductor module DC includes, for example, OpenCV Mat objects / image processing to generate image frames.
[0060] It is noted that an autonomous guided vehicle may be provided with both an on-board shared memory SH with an on-board deep conductor module DC and a media server MS with a remotely located deep conductor module DC, where the media server MS and the remotely located deep conductor module DC may be utilized in situations where the on-board processing capabilities of the autonomous guided vehicle 110 and / or the power stored in the autonomous guided vehicle 110 are conserved or limited. As described herein, the media server MS may also interface with the on-board shared memory SH with the on-board deep conductor module DC.
[0061] 1A, 2, and 6, stereo pairs of cameras 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, and 477A and 477B are calibrated (see FIG. 6, block 600) to acquire video stream data imaging using vision system 400. The stereo pairs of cameras 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, and 477A and 477B may be calibrated in any suitable manner (e.g., by intrinsic and extrinsic camera calibration, etc.) to provide detection of case units CU, storage structures (shelves, columns, etc.), and other structural features of the storage and retrieval system. 3A, 3B, 4A, 4B, and 4C, known objects (such as case units CU1, CU2, CU3 (or storage system structures)) (e.g., having known physical characteristics such as shape, size, etc.) may be placed within the field of view of the cameras of the supplemental navigation sensor system 288 (or the vehicle 110 may be positioned so that the known objects are within the field of view of the cameras). These known objects may be imaged by the cameras from several angles / perspectives in order to calibrate each camera so that the vision system controller 122VC is configured to detect the known objects based on sensor signals from the calibrated cameras.
[0062] For example, calibration of case unit monitoring cameras 410A, 410B will be described with respect to case units CU1, CU2, CU3 having known physical characteristics / parameters (it is noted that calibration of the other stereo camera pairs described herein may be effected in a similar manner). For illustrative purposes, Figures 4A-4C show exemplary images captured by one of case unit monitoring cameras 410A, 410B from three different viewpoints, where the physical characteristics / parameters (e.g., shape, length, width, height, etc.) of case units CU1, CU2, CU3 are known by vision system controller 122VC (e.g., the physical characteristics of different case units CU1, CU2, CU3 are stored in memory of or accessible to vision system controller 122VC). For example, based on the three (or more) different viewpoints of case units CU1, CU2, CU3 in the images of Figures 4A-4C, the vision system controller 122VC is provided with intrinsic and extrinsic camera and case unit parameters that result in the calibration of the case unit monitoring cameras 410A, 410B. It is noted that each camera 410A, 410B is essentially calibrated to its own coordinate system (i.e., each camera perceives the depth of the object from its respective image sensor). When multiple cameras 410A, 410B are utilized, in one embodiment, calibration of the vision system 400 includes calibrating the cameras 410A, 410B to a common base reference frame (which may be the reference frame of a single camera or any other suitable base frame of reference to which each camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B can be associated to collectively form a common base reference frame for all cameras of the vision system), and calibrating the common base reference frame to the robot reference frame.Calibrating the cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B to a common base reference frame involves identifying and applying a transformation between the respective reference frames of each camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B such that the respective reference frames of each camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B are transformed to (or referenced to) the common base reference frame. For example, the reference frame for camera 410A (although any of the cameras may be used) represents the common base reference frame. A transformation (i.e., a rigid body transformation in six degrees of freedom) is determined for each of the coordinate systems / reference frames for the other cameras 410B, 420A, 420B, 430A, 430B, 460A, 460B relative to the reference frame of camera 410A so that image data from cameras 410B, 420A, 420B, 430A, 430B, 460A, 460B is correlated or transformed to the reference frame of camera 410A.
[0063] Camera calibration involves recording (e.g., storing in memory) the viewpoints of case units CU1, CU2, and CU3, for example, relative to case unit monitoring cameras 410A and 410B, from images of vision system controller 122VC. Vision system controller 122VC estimates the poses of case units CU1, CU2, and CU3 relative to case unit monitoring cameras 410A and 410B, and estimates the poses of case units CU1, CU2, and CU3 relative to each other. The pose estimates PE of each case unit CU1, CU2, and CU3 are illustrated in Figures 4A-4C as being superimposed on each case unit CU1, CU2, and CU3.
[0064] The vehicle 110 is moved so that any suitable number of viewpoints of case units CU1, CU2, CU3 are acquired / photographed by the case unit monitoring cameras 410A, 410B, resulting in convergence of case unit characteristics / parameters (e.g., estimated by the vision system controller 122VC) for each of the known case units CU1, CU2, CU3. Once the case unit parameters have converged, the case unit monitoring cameras 410A, 410B are calibrated. The calibration process is repeated for the other case unit monitoring cameras 410A, 410B. Once both case unit monitoring cameras 410A, 410B are calibrated, the vision system controller 122VC is configured with three-dimensional rays for each pixel in each of the case unit monitoring cameras 410A, 410B, as well as estimates of the line segments of the three-dimensional baseline separating the cameras and the relative poses of the case unit monitoring cameras 410A, 410B with respect to each other. The vision system controller 122VC is configured to utilize three-dimensional rays for each pixel in each of the case unit monitoring cameras 410A, 410B, estimates of the three-dimensional baseline line segments separating the cameras, and the relative poses of the case unit monitoring cameras 410A, 410B with respect to each other so that the case unit monitoring cameras 410A, 410B form a passive stereo vision sensor, such as when a common feature is visible within the fields of view 410AF, 410BF of the case unit monitoring cameras 410A, 410B.
[0065] The common base reference frame may be transformed into the reference frame BREF of autonomous guided vehicle 110 by transferring one or more of case units CU1-CU3 to payload platform 210B of autonomous guided vehicle 110, where one or more case units CU1-CU3 are positioned within payload platform 210B (e.g., using at least positioning blade 471 and pusher 470). With one or more case units CU1-CU3 in known locations within payload platform 210B, the controller (knowing the dimensions of case unit(s) CU1-CU3) characterizes the relationship between the image field of the common base reference frame and the robot reference frame BREF, such that the position of case unit(s) CU1-CU3 in the vision system's image is calibrated to the robot reference frame BREF. The other stereo camera pairs 420A and 420B, 430A and 430B, 477A and 477B may be calibrated in a similar manner, where the common base reference frame of each pair is transformed to the reference frame BREF of the autonomous guided vehicle 110 based on known relative camera positions and / or parallax or depth map imaging of portions of the autonomous guided vehicle (within the field of view of the respective camera pair) and case units or other structures of the storage and retrieval system.
[0066] 8 , a computer model 800 (such as a computer-aided drafting model or CAD model) of autonomous guided vehicle 110 may also be employed by controller 122 (or vision controller 122VC) (in combination with, or in lieu of, determining the reference frame transformation effected by holding the case unit in the payload bay as described above) to transform the common base reference frame of each camera pair to the reference frame BREF of autonomous guided vehicle 110. As seen in FIG. 8 , characteristic dimensions of any appropriate feature of payload platform 210B, etc. (which in this example is a feature of the payload platform fence relative to the reference frame BREF, or any other appropriate feature of autonomous guided vehicle 110 via autonomous guided vehicle model 800, and / or an appropriate feature of the storage structure via virtual model 400VM of the operating environment) depending on which camera pair is calibrated may be extracted by controller 122 for the portion of autonomous guided vehicle 110 within the field of view of the camera pair. These characteristic dimensions of payload platform 210B are determined from the origin of the reference frame BREF of autonomous guided vehicle 110. These known dimensions of autonomous guided vehicle 110 are utilized by controller 122 (or vision controller 122VC), along with the disparity or depth map produced by the stereo camera pair, to correlate the common base reference frame (or each camera's reference frame) to the reference frame BREF of autonomous guided vehicle 110.
[0067] As noted above, calibration of the case unit monitoring cameras 410A, 410B has been described with respect to case units CU1, CU2, CU3, but may be performed in a generally similar manner for any suitable structure of the storage and retrieval system 100 (e.g., permanent or temporary (including calibration fixtures)).
[0068] 7A and 7B, examples of multiple camera calibration utilizing a calibration jig / fixture 700 (also referred to as a common camera calibration reference structure) include calibrating camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A, 460B, 477A, 477B to their respective common base reference frames (in a manner similar to that described above) and calibrating the common base reference frame to the robot reference frame (in a manner similar to that described above), while in other aspects of camera calibration using a case unit / storage structure and / or calibration fixture, when one or more cameras are utilized, the reference frame of one camera or the reference frames of one or more cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B may be individually calibrated to the reference frame BREF of the autonomous guided vehicle 110.
[0069] For illustrative purposes only, and with reference to Figures 3A, 3B, 7A, and 7B, calibrating each of the camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, and 477A and 477B to a common base reference frame involves identifying and applying a transformation between the respective reference frames of each of the cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, and 477B such that the respective reference frames of each camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, and 477B are transformed to (or referenced to) the respective common base reference frame. Again, the common base frame of reference may be the frame of reference of a single camera of the camera pair, or any other suitable base frame of reference to which each camera pair 410A & 410B, 420A & 420B, 430A & 430B, 460A & 460B, 477A & 477B can be associated so as to collectively form a respective common base frame of reference for all cameras of the camera pair in a manner generally similar to that described above. This calibration of cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B may be performed using a calibration fixture 700 located in any suitable holding location of storage and retrieval system 100 (e.g., a storage shelf, or any other suitable location accessible to autonomous guided vehicle 110 and within the field of view of the cameras being calibrated). As described herein, each camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B is calibrated to a common base reference frame, which describes the positional relationship between the respective camera reference frame of each camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B and each other camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B, and the reference frame BREF of the autonomous guided vehicle 110.
[0070] Calibration fixture 700 includes uniquely identifiable three-dimensional geometric shapes 710-719 (in this example, squares, some rotated relative to others) that provide an asymmetric pattern to calibration fixture 1300 and constrain the determination / transformation of the camera's reference frame (e.g., from each camera) to a common base reference frame, and the transformation between the common base reference frame and the autonomous guided vehicle's reference frame BREF, as further described to determine the relative pose of calibration fixture 1300 (and thus the case unit) with respect to transfer arm 210A of autonomous guided vehicle 110. Calibration fixture 700 shown and described herein is exemplary, and any other suitable calibration fixture may be utilized in a manner similar to that described herein. For illustrative purposes, each of the three-dimensional geometric shapes 710-719 is a predetermined size that constrains the identification of corners or points C1-C36 of the three-dimensional geometric shapes 710-719, and the transformation is such that the distance between corresponding corners C1-C36 is minimized (e.g., the distance between each of the corners C1-C36 in the frame of reference of camera 310C1 is minimized for each of the respective corners C1-C36 identified in the frames of reference 410AF, 410BF, 420AF, 420BF, 430AF, 430BF, 460AF, 460BF of each camera 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B of the camera pair being calibrated - cameras 477A, 477B have similar frames of reference).
[0071] Each of the three-dimensional geometric shapes 710-719 is imaged simultaneously (i.e., each of the three-dimensional geometric shapes 710-719 is at a single location in a common base reference frame while imaged by all cameras of the camera pair whose reference frames are calibrated to the common base reference frame of each of the camera pairs), and points / corners C1-C36 of the three-dimensional geometric shapes 1310-1319 identified in the image (one exemplary image is illustrated in FIG. 7B) are identified by the vision system 400 and uniquely identified by each of the cameras in the camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, and 477A and 477B at a single location such that they are uniquely determined independently of the orientation of the calibration fixture. The angles C1-C36 identified in each image of the image set are compared between images from the cameras in the camera pair to define a transformation to a common base reference frame for each camera reference frame (which in one example may correspond to or otherwise be defined by the reference frame of camera 410A for calibration of camera pair 410A, 410B).
[0072] Once cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B are aligned to their respective common base reference frames, each common base reference frame (or individual reference frame of one or more cameras) is transformed (e.g., aligned) to the reference frame BREF of autonomous guided vehicle 110 by transferring calibration fixture 700 (or a similar fixture) to payload platform 210B of autonomous guided vehicle 110 and / or by utilizing computer model 800 of autonomous guided vehicle 110 in a manner similar to that described above. For example, when calibration fixture 700 is transferred into payload platform 210B, calibration fixture 700 is aligned within payload platform 210B (e.g., using at least alignment blade 471 and pusher 470). When calibration fixture 700 is in a known (i.e., adjusted) position within payload platform 210B, the controller (which knows the locations of points / corners C1-C36) characterizes the relationship between the image field of the common base reference frame and the autonomous guided vehicle's frame of reference BREF such that the locations of points / corners C1-C36 in the vision system image are calibrated to the robot's frame of reference BREF. If computer model 800 is utilized, the disparity or depth map generated by the stereo camera pair is utilized with the known dimensions of the payload platform fixture (from computer model 800) and the known dimensions / locations of points / corners C1-C36 to provide a transformation between the common base reference frame of each of camera pairs 410A and 410B and the autonomous guided vehicle's frame of reference BREF, for example.
[0073] 15 , calibration of camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, and 477A and 477B may be provided or otherwise performed at a calibration station 1510 of storage structure 130. As seen in FIG. 15 , calibration station 1510 may be located at or adjacent to an autonomous guided vehicle entry or exit location 1590 of storage structure 130. The presence of autonomous guided vehicle entry or exit location 1590 allows autonomous guided vehicles 110 to be navigated to and retrieved from one or more storage levels 130L of storage structure 130 in a manner generally similar to that described in U.S. Patent No. 9,656,803, issued May 23, 2017, and entitled “Storage and Retrieval System Rover Interface,” the entire disclosure of which is incorporated herein by reference. For example, the autonomous guided vehicle entry or exit location 1590 includes a lift module 1591 such that entry and exit of the autonomous guided vehicles 110 occurs at each storage level 130L of the storage structure 130. The lift module 1591 may interface with the transfer deck 130B of one or more storage levels 130L. The interface between the lift module 1591 and the transfer deck 130B may be located at a predetermined location on the transfer deck 130B such that entry and exit of the autonomous guided vehicles 110 at each transfer deck 130B is substantially decoupled from the throughput of the automated storage and retrieval system 100 (e.g., entry and exit of the autonomous guided vehicles 110 at each transfer deck does not affect throughput). In one aspect, the lift module 1591 may interface with a spur or staging area 130B1-130Bn (e.g., an autonomous guided vehicle loading platform) connected to or forming part of the transfer deck 130B for each storage level 130L. In other embodiments, the lift module 1591 may interface substantially directly with the transfer deck 130B.It is noted that the transfer deck 130B and / or the staging areas 130B1-130Bn may include any suitable barrier 1520 that substantially prevents the autonomous guided vehicle 110 from moving away from the transfer deck 130B and / or the staging areas 130B1-130Bn at the lift module interface. In one aspect, the barrier may be a movable barrier 1520 that may be movable between a deployed position to substantially prevent the autonomous guided vehicle 110 from moving away from the transfer deck 130B and / or the staging areas 130B1-130Bn and a stowed position to allow the autonomous guided vehicle 110 to transition between the lift platform 1592 of the lift module 1591 and the transfer deck 130B and / or the staging areas 130B1-130Bn. In addition to transporting the autonomous guided vehicle 110 in and out of the storage structure 130, in one embodiment, the lift module 1591 can also transport the rover 110 between storage levels 130L without removing the autonomous guided vehicle 110 from the storage structure 130.
[0074] Each staging area 130B1-130Bn includes a respective calibration station 1510 positioned so that the autonomous guided vehicle 110 may repeatedly calibrate camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, and 477A and 477B. Calibration of the camera pairs may occur automatically upon registration of the autonomous guided vehicle within the storage structure 130 (via an autonomous guided vehicle entry or exit location 1590 in a manner similar to that described in U.S. Pat. No. 9,656,803, previously incorporated by reference). In other embodiments, calibration of the camera pairs may occur manually (such as when the calibration station is located on a lift 1592) or may be performed prior to insertion of the autonomous guided vehicle 110 into the storage structure 130 in a manner similar to that described herein with respect to the calibration station 1510.
[0075] To calibrate the stereo camera pairs, the autonomous guided vehicle is positioned (manually or automatically) at a predetermined location in the calibration station 1510. Automatic positioning of the autonomous guided vehicle 110 at the predetermined location may utilize detection of any suitable features of the calibration station 1510 using the vision system 400 of the autonomous guided vehicle 110. For example, the calibration station 1510 includes any suitable location flags or positions 1510S disposed on one or more surfaces 1200 of the calibration station 1510. The location flags 1510S are disposed on one or more surfaces within the field of view of at least one camera of each camera pair. Vision system controller 122VC is configured to detect location flags 1510S, and detection of one or more of the location flags 1510S causes the autonomous guided vehicle to be roughly positioned relative to calibration fixture 700 (e.g., stored on a shelf or other support at calibration station 1510), calibration case units (similar to case units CU1, CU2, CU3 described above and stored on a shelf at calibration station 1510), and / or other calibration datums (or known objects) such as those described herein. In other aspects, in addition to or instead of location flags 1510S, calibration station 1510 may include a buffer or physical stop against which autonomous guided vehicle 110 abuts to position itself at a predetermined location at calibration station 1510. The buffer or physical stop may be, for example, a barrier 1520 or other suitable fixed or deployable feature of the calibration station. Automatic positioning of the autonomous guided vehicle 110 at the calibration station 1510 may occur when the autonomous guided vehicle 110 is introduced into the storage and retrieval system 100 (such as when the autonomous guided vehicle exits the lift 1592) and / or at any suitable time when the autonomous guided vehicle enters the calibration station 1510 from the transfer deck 130. Here, the autonomous guided vehicle 110 may be programmed with calibration instructions that result in stereoscopic calibration upon introduction into the storage structure 130, or the calibration instructions may be initialized at any suitable time while the autonomous guided vehicle 110 is operating (i.e., in service) within the storage structure 130.
[0076] As described above, case units CU1, CU2, CU3, and / or calibration fixture 700 may be stored in storage shelves at each calibration station 1510, where calibration of the camera pair is performed at each calibration station 1510 in the manner described above. Additionally, one or more surfaces of each calibration station 1510 may include any suitable number of known objects GDTs, which may generally resemble geometric shapes 710-719. The one or more surfaces may be any surface visible by the camera pair, including, but not limited to, sidewalls 1511 of the calibration station 1510, ceiling 1512 of the calibration station 1510, floor / running surface 1515 of the calibration station 1510, and barrier 1520 of the calibration station 1510. The objects GDT (which may also be referred to as visual datums or calibration objects) included with each surface may be raised structures, apertures, appliqués (e.g., paint, stickers, etc.), each having known physical characteristics such as shape, size, etc., such that calibration of the camera pair is performed in a manner generally similar to that described above with respect to case units CU1-CU3 and / or calibration fixture 700.
[0077] As can be appreciated, the vehicle localization provided by physical property sensor system 270 (e.g., positioning of the vehicle at a predetermined location along picking aisle 130A or along transfer deck 130B relative to the pick / place location) can be augmented by pixel-level position determination provided by supplemental navigation sensor system 288. Here, controller 122 is configured to, so-called, “roughly” position vehicle 110 relative to the pick / place location by utilizing one or more sensors of physical property sensor system 270. Controller 122 is configured to, so-called, “fine-tune” the vehicle's pose and location relative to the pick / place location so that the positioning of vehicle 110 and case units CU placed by vehicle 110 at storage location 130S is held to a smaller tolerance (i.e., improved position accuracy) compared to positioning of vehicle 110 or case units CU using physical property sensor system 270 alone. Here, the pixel-level positioning provided by the supplemental navigation sensor system 288 has higher positioning definition / resolution than the electromagnetic sensor resolution provided by the physical property sensor system 270 .
[0078] 1A, 2, and 6, and FIGS. 9A-9C, to acquire video stream data imaging using vision system 400, each camera in a stereo pair of cameras 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, and 477A and 477B is calibrated for three-dimensional monocular vision (see FIG. 6, block 605). Here, the monocular calibration of each camera in camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, and 477A and 477B utilizes a disparity image or depth image and a depth image associated with a three-dimensional imaging sensor (for stereo camera pair 410A and 410B, three-dimensional imaging sensors 440A and / or 440B— At least one three-dimensional imaging sensor 440C, 440D may be positioned at each end 200E1, 200E2 of the autonomous guided vehicle 110 relative to the stereoscopic camera pairs 420A and 420B, 430A and 430B, or, it is noted that disparity or depth maps may be generated from the camera pairs themselves, generally because there are no obstacles between the respective camera pairs 420A and 420B, 430A and 430B at the ends 200E1, 200E2 of the autonomous guided vehicle 110 that would corrupt the depth map generated therefrom.) While Figure 9A illustrates, for example, a three-dimensional object (i.e., case unit CU) held in payload bay 210B of the autonomous guided vehicle 110, the three-dimensional object may be held in any suitable location within the field of view of each camera in the camera pair and its associated three-dimensional sensor(s). The case unit CU is positioned and oriented (e.g., in payload bay 210B or other suitable location) such that it is visible within the fields of view of cameras 410A, 410B and three-dimensional imaging sensors 440A, 440B. The case unit CU is observable in all images from cameras 410A, 410B and three-dimensional image sensors 440A, 440B such that objects and points common to images from cameras 410A, 410B and three-dimensional image sensors 440A, 440B can be incorporated for conversion of two-dimensional images from each monocular camera 410A, 410B to three-dimensional images for each monocular camera 410A, 410B.
[0079] 9C illustrates a disparity map or depth map generated from the vision system 400 using, for example, cameras 410A, 410B. The disparity map from the cameras 410A, 410B images is filtered, and then a clustering method is applied to the filtered disparity map to arrive at a disparity map that is utilized for object detection and localization. The disparity map generated from the cameras 410A, 410B is calibrated in any suitable manner using depth maps from one or more of the three-dimensional image sensors 440A, 440B.
[0080] An exemplary transformation equation that results in determining depth from monocular images from monocular cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B is as follows:
[0081]
number
[0082]
number
[0083] where R is the rotational transformation, r is a 3x3 translation matrix, and t is a 3x1 translation vector. The above example equations are utilized by controller 122 and / or vision controller 122VC to transform coordinates within the monocular camera image into the global reference frame GREF of storage and retrieval system 100. Depth information associated with the depth of detected objects in the front-facing plane of the point cloud or map created by the three-dimensional imaging sensor(s) 440A, 440B is utilized to generate coordinates in the global reference frame GREF. The above equations and depth information form calibration parameters utilized by controller 122 to transform monocular image coordinates into global reference frame coordinates GREF following deep learning detection analysis of the monocular images, as described herein.
[0084] With the stereo camera pair calibrated and each individual camera of the camera pair calibrated as a monocular camera, the deep conductor DC's artificial neural network (ANN) is trained ( FIG. 6 , block 610). The artificial neural network (ANN) is any suitable deep learning graphics processing algorithm configured for feature or representation learning and detection integration, and is trained in any suitable manner known in the art, such as with labeled or unlabeled data, or any other suitable manner. Briefly referring to FIG. 12A , the detection integration performed by the deep conductor DC via the artificial neural network (ANN) generates several measures, including, but not limited to, intersection-over-union (IoU) (e.g., the degree of overlap between two boxes) and detection confidence, where, for illustrative purposes only, a best-fit bounding box is selected to represent the detected object from among multiple bounding boxes generated for object detection. In one aspect, the integration result for object detection is in the form of a bounding box with confidence and detected label presented as an overlay within the image frame, as illustrated in FIG. 12A . In Figure 12A, the positioning blade or arm 471 is labeled and identified with a respective bounding box, with a 0.96 or 96% confidence that the detected object is the positioning blade 471. Also in Figure 12A, the cases or boxes are labeled and identified with a respective bounding box, with a 0.67 or 67% confidence that the detected object is the case CU. As discussed herein, although bounding boxes are illustrated in Figure 12A, the identified objects may be identified within the image frame in any suitable manner (e.g., by highlighting the detected objects or any other suitable image enhancement, etc.).
[0085] The object detection performed by the deep conductor DC is a deep learning-based multi-object detection that can present multiple detections within each image frame. The detection results, in the form of augmented image frames (such as those shown in FIG. 12A , which illustrates image frames from a monocular video stream, but note that augmented image frames are similar for image frames obtained from a stereo vision video stream), can be time-stamped and stored in any suitable memory (e.g., shared memory SH, media server MS, or any other memory in communication with controller 122) for data logging, reporting, archiving, evaluation, troubleshooting, analysis of operating factors, quality control and measurement, fault detection, or any other suitable purpose. As described herein, the output of the deep conductor DC (e.g., the above-described detection results using integrated logic and timestamps) is utilized to activate one or more of a computer vision protocol for object detection and localization and a machine learning protocol for object detection and localization. As a simple example, and with reference to FIG. 11F, a detection result of the deep conductor DC is illustrated, where objects are identified by detection from cameras 410A, 410B in payload platform 210B (e.g., at least one pick case is identified in image data from each of cameras 410A, 410B), and the identified objects are compared to any suitable predetermined threshold. For example, for a picking operation of the autonomous guided vehicle 110, the deep conductor DC compares the image data to verify that the number of identified pick cases is equal to or greater than one. The deep conductor DC may also verify whether motion is present (i.e., whether the autonomous guided vehicle is traveling across the driving surface at a speed that causes motion blur), which in the case of FIG. 11F is not blurred.Based on the results of the image comparison, Deep Conductor DC may activate a computer vision protocol if the number of identified pick cases in the image data from camera 410A is the same as the number of identified pick cases in the image data from camera 410B, otherwise a machine learning protocol may be activated by Deep Conductor DC.
[0086] At least one machine learning model ML is generated by the controller 122 onboard the autonomous guided vehicle 110 (or by any suitable server / computer offboard the autonomous guided vehicle 110 but in communication with the onboard controller 122) ( FIG. 6 , block 615), where the at least one machine learning model ML represents learned weights and parameters of an artificial neural network ANN. The machine learning model(s) ML are configured to provide detection and localization (with respect to the reference frame BREF of the autonomous guided vehicle 110) of objects specific to a defined task of the autonomous guided vehicle 110. For example, a machine learning model ML may be generated for each of vehicle drivability, case CU picking, case CU placement, collision avoidance, ceiling tag or structure detection for autonomous guided vehicle 110 location, remote case inspection, and any other suitable task and / or operating state of the autonomous guided vehicle 110. The machine learning model ML is stored in any suitable memory of the autonomous guided vehicle 110, such as a shared memory SH or a media server MS, so as to be accessible to the controller 122 and its deep conductor DC.
[0087] The artificial neural network ANN and machine learning model ML provide real-time object detection, where the deep conductor DC's artificial neural network ANN determines which detection protocol to employ (e.g., a computer vision object detection and localization protocol, a machine computer vision object detection and localization protocol, or both) (FIG. 6, block 620), and using the selected detection protocol, objects, tasks, and / or motion states are identified or otherwise detected (FIG. 6, block 625). Examples of detected objects include, but are not limited to, support tines (also called forks) 210AT of transfer arm 210A, pusher 470 of transfer arm 210A, puller 472 of transfer arm 210A, positioning blade 471 of transfer arm 210A, storage shelf hat 444, electrical panel of storage and retrieval system 100 structure (see FIG. 11E), vertical bar 445 of storage and retrieval system 100 structure (see also FIGS. 4B and 11B), other autonomous guided vehicles 110, picking aisle 130A, partially identified cases, and suspicious cases (FIG. 11D, e.g., which may be misplaced on a shelf), with reference to FIGS. Examples of tasks for an autonomous guided vehicle include, but are not limited to, a picking task (e.g., a case unit is positioned on a shelf in an orientation suitable for picking) and a no-pick task (e.g., a case unit is positioned on a shelf in an orientation unsuitable for picking), with reference to Figure 11F. Examples of motion states for an autonomous guided vehicle include, but are not limited to, motion of the autonomous guided vehicle along a driving surface (see Figures 10, 11B, and 11C), where such motion is indicated by blur in image frames (e.g., image frame 1000) received from shared memory SH and / or media server MS.Here, the coordinates of the objects, motion states, and / or tasks are identified and obtained from the image frames, such as by the coordinates of their respective bounding boxes (as illustrated in Figures 11A-11F) based on the identification of the objects, tasks, and / or motion states, although in other aspects the coordinates of the objects may be identified and obtained in any suitable manner using any suitable visual processing algorithm, in combination with or instead of the bounding boxes.
[0088] In determining which detection protocol(s) to employ (e.g., which detection protocol determination provides the highest level of confidence in object detection), the deep conductor DC assigns a flag to each detected object, task, and motion state via an artificial neural network (ANN). The flags are utilized for illustrative purposes of comparison with respective predetermined thresholds, but it is noted that any suitable thresholds may be utilized. The flags are, for example, markers indicating the presence of a respective condition, object, or motion state. Here, each detected condition, object, and motion state is assigned a flag (e.g., 1 (or other integer greater than 0) if present, and 0 if absent). These flags form metadata for each image frame and are utilized by the deep conductor DC (along with one or more of the image frame's timestamp, detection label, and bounding box or object coordinates) to provide robust object detection and localization by selecting a detection / localization protocol from one or both of computer vision protocols and machine learning protocols from the video stream imaging data. For example, if no motion is detected in an image frame, the flag for the binocular depth map is set to 1 and the flag for vision preservation (obstruction / failure of at least one camera in the stereo camera pair) is set to 0, and the deep conductor DC then selects a machine vision detection protocol for object detection and localization. If there is no motion and the flag for the binocular depth map is set to 0 (e.g., binocular vision is obstructed, one camera in the camera pair is unavailable, and an object is detected by one camera but not the other), and the flag for vision preservation is set to 1 (e.g., as a result of the aforementioned anomalous stereo vision), the deep conductor DC then selects a machine learning protocol for object detection and localization.If the same object is detected in the image frames of both cameras in the pair, the detected object is compared with any appropriate threshold (e.g., a detection confidence threshold), and if the threshold is met, a flag for the binocular depth map is set to 1, and since the number of detections for each camera in the pair due to the common object is below a predetermined threshold, a vision maintenance flag is set to 1, and the deep conductor DC selects both a computer vision protocol and a machine learning protocol for object detection and localization.
[0089] Using vehicle motion as an example, referring to FIG. 11C (which illustrates an image from camera 410A, where the corresponding image frame from camera 410B is nearly similar but mirrored), the artificial neural network detects the presence of motion in the image frames from cameras 410A, 410B, and the motion flags for cameras 410A, 410B are set to 1. However, it is noted that FIG. 11C illustrates that cameras 410A, 410B have an unobstructed view of each other. The artificial neural network ANN is trained (as described herein) to recognize the unobstructed camera views and the resulting stereo imagery. Now, because the camera views are unobstructed and a depth map can be generated from the stereo camera pair 410A, 410B, computer vision protocols can be used for object detection and localization. In this manner, the deep conductor outputs a selection of a computer vision protocol for object detection and localization for cameras 410A, 410B. As another example, referring to FIG. 12A , an image frame from camera 410B (note that the corresponding image frame from camera 410A is substantially the same but mirrored) is illustrated, identifying the case or box and the positioning blade. Here, the box may obstruct the field of view of cameras 410A, 410B, preventing the generation of a stereoscopic image from cameras 410A, 410B, although each camera provides its own monocular image. In this manner, deep conductor DC recognizes that a computer vision protocol is unavailable and outputs the selection of a machine learning protocol for object detection and location for cameras 410A, 410B. As yet another example, FIG. 11F is an example image frame from camera 410B (note that the corresponding image frame from camera 410A is substantially the same but mirrored), identifying the tine or fork 210AT, the case to be picked, and the shelf hat.Here, both cameras have an unobstructed field of view, and the number of common objects detected in the image frames from each camera 410A, 410B in the camera pair is less than a predetermined threshold of, for example, six objects (in other embodiments, the threshold may be more or less than six), so that both computer vision and machine learning protocols are available. If both protocols are available, the deep conductor DC may select both detection and localization protocols, so that any suitable image fusion algorithm may be utilized by the deep conductor DC to combine the two-dimensional (monocular) image frames with the three-dimensional (stereoscopic) image frames and depth maps from the three-dimensional sensors 440A, 440B.
[0090] If a machine learning protocol is selected, the deep conductor DC selects one or more of the deep learning models ML based on a predetermined task (e.g., picking or placing cases, navigating a transfer deck or picking aisle, etc.) for the autonomous guided vehicle 110 ( FIG. 13 , block 1300). The deep conductor receives image frames for each camera in the camera pair from the shared memory SH and / or media server MS ( FIG. 13 , block 1305) and utilizes the selected deep learning model ML to provide object detection within the image frames for the predetermined task ( FIG. 13 , block 1310). Once an object is detected within the image frame, the deep conductor DC performs detection integration ( FIG. 13 , block 1315), where the two-dimensional image frame (with the detected object) is integrated with a depth map from a three-dimensional sensor (e.g., the image frame from camera 410A is integrated with the depth map from three-dimensional sensor 440A) to form an integrated image or monocular depth map (e.g., such as that illustrated in FIG. 9B). The monocular depth map provides depth information similar to that obtained by stereoscopic vision or machine vision acquired with a stereo camera pair (e.g., cameras 410A, 410B—see FIG. 9C ). The monocular depth map may be utilized by controller 122 where images acquired from video stream data from a stereo camera pair, e.g., cameras 410A, 410B, recorded by controller 122, do not support object detection and localization through binocular vision. As described above, the monocular depth includes depth information similar to that obtained with binocular vision, such that any suitable depth information affecting the operation of autonomous guided vehicle 110 can be obtained from the monocular depth map. Such information obtained from the monocular depth map includes, but is not limited to, identifying the distance of case units from the autonomous guided vehicle, the clearance between case units stored on shelf 555, the clearance between adjacent case units on shelf 555 and the case unit to be picked / placed, and the size of the storage space in which the case units are placed.The two-dimensional image frame (with the detected object) is integrated with the depth map from the three-dimensional sensor 440A, 440B associated with the detected object, for example, by searching the front-facing plane of the depth map generated using the three-dimensional sensor 440A, 440B associated with the detected object (e.g., three-dimensional sensor 440A is associated with the object detected by camera 410A, and three-dimensional sensor 440B is associated with the object detected by camera 410B). This front-facing plane search is effected by the deep conductor DC using transformed coordinates obtained during stereo calibration of the camera pair (e.g., camera pair 410A, 410B). The object in the integrated image is localized ( FIG. 13 , block 1320), where the coordinates of the detected object from the integrated image are transformed into coordinates of the global reference frame GREF via depth map information from one or more three-dimensional sensors 440A, 440B. Localization may also include incorporating selected points generated from autonomous guided vehicle model 800 with associated points from corresponding images of the autonomous guided vehicle model taken by corresponding cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B. Cameras 410A, 410B are described above for illustrative purposes, and it should be understood that the same may be provided by any of the camera pairs described herein.
[0091] As described herein, object detection and localization determined by machine learning protocols, computer vision protocols, or both, occurs in real time as the autonomous guided vehicle performs a task. Also as described herein, detections (e.g., image frames augmented or enhanced by deep conductor DC as described herein—see FIGS. 11A-12A ) are generated and output to provide an endpoint decision regarding autonomous guided vehicle operation ( FIG. 6 , block 630). Here, deep conductor DC outputs the detections and indicates to controller 122 (deep conductor DC is a module of controller 122) whether the autonomous guided vehicle 110 operation / task is permitted (i.e., instructions are given to the autonomous guided vehicle 110 to complete the task) or not permitted (i.e., instructions are not given to the autonomous guided vehicle 110 to complete the task).
[0092] The augmented image frames output by deep conductor DC may be analyzed live or in real time (while the autonomous guided vehicle is performing a task, with analysis of the augmented image frames affecting task completion or incompletion) and / or recorded or stored as described herein in any suitable memory (e.g., shared memory SH, media server MS, memory of control server 120, memory of warehouse management system 2500, etc.) so that the recorded detection results are identifiable by or corresponding to each autonomous guided vehicle 110. For example, if an autonomous guided vehicle 110 is unable to validate a task (e.g., if a case CU is stuck on a shelf or part of the robot and cannot be fully transferred to payload platform 210B), an operator may analyze the live or stored image frames via user interface UI to determine the cause of the incomplete task and / or manually control the autonomous guided vehicle to correct the case transfer malfunction. Here, the vision system controller 122VC (and / or the controller 122) is configured, in one or more embodiments, to provide remote viewing via the vision system 400, where such remote viewing may be presented to the operator in augmented reality or any other suitable manner (e.g., non-augmented manner). For example, the autonomous guided vehicle 110 is communicatively connected to the warehouse management system 2500 (e.g., via the control server 120) over the network 180 (or any other suitable wireless network). The warehouse management system 2500 includes one or more warehouse control center user interfaces UI. The warehouse control center user interfaces UI may be any suitable interface, such as a desktop computer, a laptop computer, a tablet, a smartphone, a virtual reality headset, or any other suitable user interface configured to present visual and / or auditory data obtained from the autonomous guided vehicle 110.In some embodiments, vehicle 110 may include one or more microphones MCP (FIG. 2), where one or more microphones and / or remote viewing (e.g., of live video images and / or analyzed real-time augmented images output by Deep Conductor DC) may assist in preventative maintenance / troubleshooting diagnosis for components of the storage and retrieval system, such as vehicle 110, other vehicles, lifts, storage shelves, etc. The warehouse control center user interface UI is configured to allow a user of the warehouse control center to request images from the autonomous guided vehicle 110 or otherwise be provided images to the user (such as upon detection of an unidentifiable object 299 and / or detection of a suspicious object (see FIG. 11D)), and for the requested / provided images to be viewed on the warehouse control center user interface UI.
[0093] The supplied and / or requested images may be a live video stream, pre-recorded images (and stored in any suitable memory of the autonomous guided vehicle 110 or warehouse management system 2500), or images (e.g., one or more still images and / or dynamic video images) corresponding to a specified (user-selectable or preset) time interval or number of images taken on demand or output by Deep Conductor DC in substantially real time in response to each image request. It is noted that the live video stream and / or image capture provided by the vision system 400, vision system controller 122VC, and Deep Conductor DC may result in real-time remote control operation (e.g., teleoperation) of the autonomous guided vehicle 110 by a user at a warehouse control center via a user interface UI at the warehouse control center.
[0094] In some aspects, the live video is streamed (augmented or unaugmented) from the vision system 400 of the supplemental navigation sensor system 288 to the user interface UI as a conventional video stream (e.g., images are presented in the user interface without augmentation, presenting what the camera “sees”), in a manner similar to that described in U.S. Patent Application No. 17 / 804,026, filed May 25, 2022, entitled “Autonomous Transport Vehicle with Vision System” (having Attorney Docket No. 1127P016037-US(PAR)), the entire disclosure of which was previously incorporated herein by reference. A virtual reality headset may be utilized by the user to view the streamed video, with an image from the front case unit monitoring camera 410A presented in the viewfinder of the virtual reality headset corresponding to the user’s left eye, and an image from the rear case unit monitoring camera 410B presented in the viewfinder of the virtual reality headset corresponding to the user’s right eye.
[0095] The image frames output by the deep conductor DC may also be presented to a user via a virtual reality headset in a manner similar to that described above for live video. Here, in addition to the detection label and confidence associated with the detected object, a machine learning model may be generated as a result of training an artificial neural network (ANN) that enhances features of the detected object in the image frame. For example, referring to FIG. 12B , image frames from camera 410B are provided via the deep conductor DC with the detected object identified. Controller 122 utilizes the machine learning model to enhance or enhance the detected object accordingly. For example, the machine learning model may be applied by controller 122 to enhance or otherwise enhance features of the positioning blade 471 and case CU detected in the image. Here, features (such as edges) of the positioning blade 471 and case CU are highlighted or otherwise enhanced to explicitly define the boundaries of the detected object. While FIG. 12B illustrates bounding boxes for each of the positioning blade 471 and case CU, it is noted that such bounding boxes may be omitted if enhancements of the detected object are provided. It is noted that while FIG. 12B illustrates an expanded view of the position adjustment bar 471 and case CU, any suitable structure of the case, storage and retrieval system 100, and / or autonomous guided vehicle 110 may be expanded within the image frame output by the deep conductor DC, such that the detected and expanded object(s) may be determined from the output image frame.
[0096] 1A, 1B, 2, 3A-3C, and 14, an exemplary method (e.g., of object detection and location for an autonomous guided vehicle 110) is described in accordance with aspects of the disclosed embodiment. Here, an autonomous guided vehicle 110 is provided (FIG. 14, block 1400). As described herein, the autonomous guided vehicle 110 is provided with a frame 200 having a payload holding portion 210B and a drive section 261D connected to the frame 200, the drive section 261D including drive wheels 260 that support the autonomous guided vehicle 110 on a running surface 284, where the drive wheels 260 provide vehicle movement on the running surface 284 and move the autonomous guided vehicle 110 on the running surface 284 in a facility (e.g., a logistics facility, an example of which is the storage and retrieval system 100). Payload handler or transfer arm 210A is coupled to frame 200 and configured to transfer payloads (such as cases CU) with flat, non-deterministic seating surfaces that are seated on payload holders or payload platforms 210B to / from payload holders 210B and payload storage locations 130S within storage array SA. As described herein, vision system 400 is mounted to frame 200 and has at least one camera 410A, 410B, 420A, 420B, 430A, 430B, 477A, 477B, and controller 122 is communicatively coupled to vision system 400.
[0097] At least one camera 410A, 410B, 420A, 420B, 430A, 430B, 477A, 477B of the vision system 400 generates video stream data imaging of an object (such as those described herein) within the logistics space (FIG. 14, block 1405), where the object is at least one of at least a portion of the frame 200, at least a portion of the payload (such as the case CU), at least a portion of the payload handler or transfer arm 210A, and at least a portion of a logistics item or structure within the logistics space outside the autonomous guided vehicle 110. The controller 122 records video stream data imaging from at least one camera 410A, 410B, 420A, 420B, 430A, 430B, 477A, 477B (FIG. 14, block 1410) and provides robust object detection and localization within a predetermined reference frame (such as the autonomous guided vehicle's reference frame BREF and / or the global reference frame GREF) from the video stream data imaging alternatively via both binocular and monocular vision (FIG. 14, block 1415), wherein detections determined via monocular vision have equivalent reliability to detections determined via binocular vision.
[0098] While the vision system 400 and the controller 122 (including one or more machine learning models ML and one or more artificial neural networks ANN) are described herein with respect to an autonomous guided vehicle 110, it should be understood that in other aspects, the vision system 400 and one or more machine learning models ML and one or more artificial neural networks ANN may be applied to a load handling device 150LHD ( FIG. 1 , which may be generally similar to the payload platform 210B of the autonomous guided vehicle 110) and a controller of a vertical lift 150 or a pallet builder of an in-feed transfer station 170. A suitable example of a load handling device of a lift into which the vision system 400 may be incorporated is described in U.S. Pat. No. 10,947,060, entitled “Vertical Sequencer for Product Order Fulfilment,” issued March 16, 2021, the entire disclosure of which is incorporated herein by reference.
[0099] 1A, 1B, 2, 3A-3C, and 16, an exemplary method (e.g., of object detection and location for an autonomous guided vehicle 110) is described in accordance with aspects of the disclosed embodiment. Here, an autonomous guided vehicle 110 is provided (FIG. 16, block 1600). As described herein, the autonomous guided vehicle 110 is provided with a frame 200 having a payload holding portion 210B and a drive section 261D connected to the frame 200, the drive section 261D including drive wheels 260 that support the autonomous guided vehicle 110 on a running surface 284, where the drive wheels 260 provide vehicle movement on the running surface 284 and move the autonomous guided vehicle 110 on the running surface 284 in a facility (e.g., a logistics facility, an example of which is the storage and retrieval system 100). Payload handler or transfer arm 210A is coupled to frame 200 and configured to transfer payloads (such as cases CU) with flat, non-deterministic seating surfaces that rest on payload holders or payload platforms 210B to / from payload holders 210B and payload storage locations 130S within storage array SA. As described herein, vision system 400 is mounted to frame 200 and has at least one camera 410A, 410B, 420A, 420B, 430A, 430B, 477A, 477B. Controller 122 is communicatively coupled to vision system 400 and to three-dimensional imaging systems 440A, 440B and at least one or more distance sensors (such as one or more of laser sensor 271 and ultrasonic sensor 272) that detect the distance of an object (such as those described herein).
[0100] At least one camera 410A, 410B, 420A, 420B, 430A, 430B, 477A, 477B of the vision system 400 generates video stream data imaging of an object within the logistics space (FIG. 16, block 1610), where the object is at least one of at least a portion of the frame 200, at least a portion of the payload (such as the case CU), at least a portion of the payload handler or transfer arm 210A, and at least a portion of a logistics item or structure within the logistics space outside the autonomous guided vehicle 110. The controller 122 records video stream data imaging from at least one camera 410A, 410B, 420A, 420B, 430A, 430B, 477A, 477B (FIG. 16, block 1620) and selectively uses binocular and monocular vision from the video stream data imaging to provide object detection and localization within a predetermined reference frame (such as one or more of a global reference frame GREF and an autonomous guided vehicle reference frame BREF) from the video stream data imaging (FIG. 16, block 1630), each of the object detection and localization using binocular vision and the object detection and localization using monocular vision being selectable on demand by the controller 122.
[0101] In accordance with one or more aspects of the disclosed embodiment, an autonomous guided vehicle includes a frame having a payload holding portion, a drive section coupled to the frame, the drive section including drive wheels that support the autonomous guided vehicle on a driving surface, the drive wheels providing vehicle navigation on the driving surface and moving the autonomous guided vehicle over the driving surface at a facility, a payload handler coupled to the frame, the payload handler configured to transport payloads having a flat, non-deterministic seating surface seated on the payload holding portion to / from the payload holding portion of the autonomous guided vehicle and a storage location for the payload within a storage array, and a vision system mounted on the frame, the vision system including at least one camera positioned to generate video stream data imaging of objects within a logistics space. the object is at least one of at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics item or structure in a logistics space outside the autonomous guided vehicle; and a controller communicatively connected to record video stream data imaging from the at least one camera and communicatively connected to at least one or more of a time-of-flight sensor and a distance sensor that detects the distance of the object, the controller being configured to provide robust object detection and localization within a predetermined frame of reference from the video stream data imaging alternatively via both binocular and monocular vision from the video stream data imaging, wherein detection determined via monocular vision has equivalent reliability to detection determined via binocular vision.
[0102] In accordance with one or more aspects of the disclosed embodiment, the controller is configured to provide object detection via monocular vision using a deep machine learning model.
[0103] In accordance with one or more aspects of the disclosed embodiment, the controller is configured to interface with a machine learning module having object detection capabilities that determines object detection from monocular vision and a deep machine learning model.
[0104] In accordance with one or more aspects of the disclosed embodiment, the autonomous guided vehicle further comprises a media server communicatively connected to the at least one camera and that records the video stream data, the media server interfacing with a controller, the controller being located onboard the autonomous guided vehicle or remotely located from the autonomous guided vehicle.
[0105] In accordance with one or more aspects of the disclosed embodiment the payload handler is configured to under-pick the payload from the storage location.
[0106] According to one or more aspects of the disclosed embodiments, at least one camera of the vision system includes two cameras forming a stereo vision camera pair, and the robust object detection and localization results in under-picking by a payload handler of a payload from more than two closely packed payloads held in adjacent storage locations regardless of the availability of stereo vision from the stereo vision camera pair.
[0107] According to one or more aspects of the disclosed embodiments, at least one camera of the vision system includes two cameras forming a stereo vision camera pair, and the robust object detection and localization results in under-picking by the payload handler of a deformed payload from more than two closely packed payloads held in adjacent storage locations, regardless of the availability of stereo vision from the stereo vision camera pair.
[0108] According to one or more aspects of the disclosed embodiments, at least one camera of the vision system includes two cameras forming a stereo vision camera pair, and the robust object detection and localization results in underpicking by a payload handler of payloads from more than two payloads held in adjacent storage locations and having a dynamic Gaussian case size distribution within the facility, regardless of the availability of stereo vision from the stereo vision camera pair.
[0109] In accordance with one or more aspects of the disclosed embodiment, object localization provided by monocular vision is as reliable as object localization determined via binocular vision.
[0110] In accordance with one or more aspects of the disclosed embodiment, the controller is configured to enable interchangeable selection between binocular object detection and location and monocular object detection and location.
[0111] In accordance with one or more aspects of the disclosed embodiment, the controller has a selector arranged to select between binocular vision object detection and localization and monocular vision object detection and localization on demand based on detection of a predetermined operating characteristic of the autonomous guided vehicle.
[0112] In accordance with one or more aspects of the disclosed embodiment, the predetermined characteristic is video stream data recorded by the controller that does not support object detection and localization through binocular vision.
[0113] According to one or more aspects of the disclosed embodiments, a method for object detection and localization for an autonomous guided vehicle is provided, the method including the steps of providing an autonomous guided vehicle with a frame having a payload holding portion, a drive section coupled to the frame, the drive section including drive wheels for supporting the autonomous guided vehicle on a driving surface, the drive wheels providing vehicle movement on the driving surface and moving the autonomous guided vehicle on the driving surface at a facility, a payload handler coupled to the frame, the payload handler configured to transport payloads having a flat, non-deterministic seating surface seated on the payload holding portion to / from the payload holding portion of the autonomous guided vehicle and a payload storage location in a storage array, a vision system mounted on the frame, the vision system having at least one camera, and a controller communicatively coupled to the vision system; generating video stream data imaging of an object, wherein the object is at least one of at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics item or structure in a logistics space outside the autonomous guided vehicle; recording the video stream data imaging from the at least one camera using a controller; and detecting the distance of the object using at least one or more of a time-of-flight sensor and a distance sensor communicatively coupled to the controller, wherein the controller provides robust object detection and localization within a predetermined reference frame from the video stream data imaging alternatively via both binocular and monocular vision, wherein detection determined via monocular vision has equivalent reliability to detection determined via binocular vision.
[0114] In accordance with one or more aspects of the disclosed embodiment, the controller provides object detection via monocular vision using deep machine learning models.
[0115] In accordance with one or more aspects of the disclosed embodiment, the controller interfaces with a machine learning module having object detection capabilities that determines object detection from monocular vision and deep machine learning models.
[0116] In accordance with one or more aspects of the disclosed embodiments, a media server is communicatively connected to at least one camera to record video stream data, the media server interfacing with a controller, the controller being located onboard the autonomous guided vehicle or remotely located from the autonomous guided vehicle.
[0117] In accordance with one or more aspects of the disclosed embodiment, a payload handler underpicks a payload from a storage location.
[0118] According to one or more aspects of the disclosed embodiment, at least one camera of the vision system includes two cameras forming a stereo vision camera pair, and the method further includes providing, via robust object detection and localization, underpicking by the payload handler of a payload from more than two closely spaced payloads held in adjacent storage locations, regardless of the availability of stereo vision from the stereo vision camera pair.
[0119] According to one or more aspects of the disclosed embodiment, at least one camera of the vision system includes two cameras forming a stereo vision camera pair, and the method further includes providing under-picking by the payload handler of the deformed payload from more than two closely packed payloads held in adjacent storage locations, regardless of the availability of stereo vision from the stereo vision camera pair, via robust object detection and localization.
[0120] In accordance with one or more aspects of the disclosed embodiment, at least one camera of the vision system includes two cameras forming a stereo vision camera pair, and the method further includes, via robust object detection and localization, providing underpicking by the payload handler of payloads from more than two payloads held in adjacent storage locations and having a dynamic Gaussian case size distribution within the facility, regardless of the availability of stereo vision from the stereo vision camera pair.
[0121] In accordance with one or more aspects of the disclosed embodiment, object localization provided by monocular vision is as reliable as object localization determined via binocular vision.
[0122] In accordance with one or more aspects of the disclosed embodiment, binocular and monocular object detection and location are interchangeably selectable via a controller.
[0123] In accordance with one or more aspects of the disclosed embodiment, the controller has a selector arranged to select between binocular vision object detection and localization and monocular vision object detection and localization on demand based on detection of a predetermined operating characteristic of the autonomous guided vehicle.
[0124] In accordance with one or more aspects of the disclosed embodiment, the predetermined characteristic is video stream data recorded by the controller that does not support object detection and localization through binocular vision.
[0125] In accordance with one or more aspects of the disclosed embodiment, an autonomous guided vehicle includes a frame having a payload holding portion, a drive section coupled to the frame, the drive section having drive wheels that support the autonomous guided vehicle on a driving surface, the drive wheels effecting vehicle travel on the driving surface and moving the autonomous guided vehicle on the driving surface at a facility, a payload handler coupled to the frame, the payload handler configured to transport payloads having a flat, non-deterministic seating surface seated on the payload holding portion to / from the payload holding portion of the autonomous guided vehicle and a storage location for the payload in a storage array, and a vision system mounted on the frame, the vision system having at least one camera positioned to generate video stream data imaging of an object within a logistics space, the object being captured by a camera positioned to capture video stream data imaging of the object within the frame. the autonomous guided vehicle includes a vision system at least partially comprising at least one of at least a portion of a payload, at least a portion of a payload handler, and at least a portion of a logistics item or structure within a logistics space outside the autonomous guided vehicle; and a controller communicatively connected to record video stream data imaging from the at least one camera and communicatively connected to at least one or more of a time-of-flight sensor and a distance sensor that detect the distance of an object, the controller being configured such that object detection and localization within a predetermined reference frame is derived from the video stream data imaging selectively using binocular vision and monocular vision from the video stream data imaging, each of the object detection and localization using binocular vision and the object detection and localization using monocular vision being selectable on demand by the controller.
[0126] In accordance with one or more aspects of the disclosed embodiment, monocular detection and localization is as reliable as detection determined via binocular object detection and localization.
[0127] In accordance with one or more aspects of the disclosed embodiment, the controller is configured to enable interchangeable selection between binocular object detection and location and monocular object detection and location.
[0128] In accordance with one or more aspects of the disclosed embodiment, the controller has a selector arranged to select between binocular vision object detection and localization and monocular vision object detection and localization on demand based on detection of a predetermined operating characteristic of the autonomous guided vehicle.
[0129] In accordance with one or more aspects of the disclosed embodiment, the predetermined characteristic is video stream data imaging recorded by the controller that does not support object detection and localization through binocular vision.
[0130] In accordance with one or more aspects of the disclosed embodiment, the controller is configured to provide object detection via monocular vision using a deep machine learning model.
[0131] In accordance with one or more aspects of the disclosed embodiment, the controller is configured to interface with a machine learning module having object detection capabilities that determines object detection from monocular vision and a deep machine learning model.
[0132] In accordance with one or more aspects of the disclosed embodiment, the autonomous guided vehicle further comprises a media server communicatively connected to the at least one camera and that records the video stream data, the media server interfacing with a controller, the controller being located onboard the autonomous guided vehicle or remotely located from the autonomous guided vehicle.
[0133] In accordance with one or more aspects of the disclosed embodiments, there is provided a method for object detection and localization for an autonomous guided vehicle, the method including: an autonomous guided vehicle comprising: a frame having a payload holding portion; a drive section coupled to the frame, the drive section having drive wheels for supporting the autonomous guided vehicle on a driving surface, the drive wheels providing vehicle navigation on the driving surface and moving the autonomous guided vehicle on the driving surface at a facility; a payload handler coupled to the frame, the payload handler configured to transport a payload having a flat, non-deterministic seating surface seated on the payload holding portion to / from the payload holding portion of the autonomous guided vehicle and a storage location for the payload in a storage array; and a vision system mounted on the frame, the vision system having at least one camera positioned to generate video stream data imaging of an object in a logistics space, the object being captured by at least a portion of the frame, at least one of the payload and the payload. The method includes providing a vision system, at least one of a payload handler, a payload handler, and a logistics item or structure within a logistics space outside the autonomous guided vehicle, and a controller communicatively connected to the vision system and to at least one or more of a time-of-flight sensor and a distance sensor that detect the distance of an object; using the controller to record video stream data imaging from the at least one camera; and using the controller to selectively use binocular vision and monocular vision from the video stream data imaging to provide object detection and localization within a predetermined reference frame from the video stream data imaging, wherein each of the binocular vision object detection and localization and the monocular vision object detection and localization is selectable on demand by the controller.
[0134] In accordance with one or more aspects of the disclosed embodiment, monocular detection and localization is as reliable as detection determined via binocular object detection and localization.
[0135] In accordance with one or more aspects of the disclosed embodiment, binocular and monocular object detection and location are interchangeably selectable by a controller.
[0136] In accordance with one or more aspects of the disclosed embodiment, the controller has a selector arranged to select between binocular vision object detection and localization and monocular vision object detection and localization on demand based on detection of a predetermined operating characteristic of the autonomous guided vehicle.
[0137] In accordance with one or more aspects of the disclosed embodiment, the predetermined characteristic is video stream data imaging recorded by the controller that does not support object detection and localization through binocular vision.
[0138] In accordance with one or more aspects of the disclosed embodiment, the controller provides object detection via monocular vision using deep machine learning models.
[0139] In accordance with one or more aspects of the disclosed embodiment, the controller interfaces with a machine learning module having object detection capabilities that determines object detection from monocular vision and deep machine learning models.
[0140] In accordance with one or more aspects of the disclosed embodiments, a media server is communicatively connected to at least one camera for recording video stream data, the media server interfacing with a controller, the controller being located onboard the autonomous guided vehicle or remotely located from the autonomous guided vehicle.
[0141] It should be understood that the foregoing description is merely illustrative of aspects of the disclosed embodiments. Various substitutions and modifications may be contemplated by those skilled in the art without departing from the aspects of the disclosed embodiments. Accordingly, aspects of the disclosed embodiments are intended to embrace all such substitutions, modifications, and variations that fall within the scope of any claims appended hereto. Furthermore, the mere fact that different features are recited in mutually different dependent or independent claims does not indicate that a combination of these features cannot be used to advantage and that such combination remains within the scope of aspects of the disclosed embodiments.
Claims
1. An autonomous guided vehicle, the autonomous guided vehicle comprising: a frame having a payload holder; a drive section coupled to the frame, the drive section including drive wheels that support the autonomous guided vehicle on a driving surface, the drive wheels providing vehicle movement on the driving surface and moving the autonomous guided vehicle on the driving surface at a facility; a payload handler coupled to the frame, the payload handler configured to transport a payload having a flat, non-deterministic seating surface seated on the payload holder to / from the payload holder of the autonomous guided vehicle and a storage location for the payload in a storage array; a vision system mounted on the frame, the vision system having at least one camera positioned to generate video stream data imaging of an object within a logistics space, the object being at least one of at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics item or structure within the logistics space external to the autonomous guided vehicle; a controller communicatively connected to record the video stream data imaging from the at least one camera and to at least one or more of a time-of-flight sensor and a distance sensor to detect the distance of the object; The autonomous guided vehicle, wherein the controller is configured to provide robust object detection and localization within a predetermined frame of reference from the video stream data imaging alternatively via both binocular and monocular vision from the video stream data imaging, wherein detections determined via the monocular vision have equivalent reliability to detections determined via the binocular vision.
2. The autonomous guided vehicle of claim 1 , wherein the controller is configured to use a deep machine learning model to provide object detection via the monocular vision.
3. The autonomous guided vehicle of claim 1 , wherein the controller is configured to interface with a machine learning module having object detection capabilities that determines object detection from the monocular vision and deep machine learning model.
4. 10. The autonomous guided vehicle of claim 1, further comprising a media server communicatively connected to the at least one camera and configured to record video stream data, the media server interfacing with the controller, the controller being located onboard the autonomous guided vehicle or remotely located from the autonomous guided vehicle.
5. The autonomous guided vehicle of claim 1 , wherein the payload handler is configured to underpick the payload from the storage location.
6. the at least one camera of the vision system includes two cameras forming a stereo vision camera pair; 10. The autonomous guided vehicle of claim 1, wherein the robust object detection and localization results in underpicking by a payload handler of payloads from more than two closely spaced payloads held in adjacent storage locations regardless of availability of stereoscopic vision from the stereo vision camera pair.
7. the at least one camera of the vision system includes two cameras forming a stereo vision camera pair; 10. The autonomous guided vehicle of claim 1, wherein the robust object detection and localization results in underpicking by a payload handler of a deformed payload from more than two closely spaced payloads held in adjacent storage locations regardless of availability of stereoscopic vision from the stereo vision camera pair.
8. the at least one camera of the vision system includes two cameras forming a stereo vision camera pair; 10. The autonomous guided vehicle of claim 1, wherein the robust object detection and localization results in underpicking by a payload handler of payloads from more than two payloads held in adjacent storage locations and having a dynamic Gaussian case size distribution within the facility, regardless of availability of stereo vision from the stereo vision camera pair.
9. The autonomous guided vehicle of claim 1 , wherein object localizations provided by the monocular vision are as reliable as object localizations determined via the binocular vision.
10. The autonomous guided vehicle of claim 1 , wherein the controller is configured to interchangeably select between detecting and locating objects using binocular vision and detecting and locating objects using monocular vision.
11. 10. The autonomous guided vehicle of claim 1, wherein the controller comprises a selector arranged to select between binocular vision object detection and location and monocular vision object detection and location on demand based on detection of predetermined operating characteristics of the autonomous guided vehicle.
12. The autonomous guided vehicle of claim 11 , wherein the predetermined characteristic is video stream data recorded by the controller that does not support object detection and localization using the binocular vision.
13. 1. A method for object detection and localization for an autonomous guided vehicle, said method comprising: For autonomous guided vehicles, a frame having a payload holder; a drive section coupled to the frame, the drive section including drive wheels that support the autonomous guided vehicle on a driving surface, the drive wheels providing vehicle movement on the driving surface and moving the autonomous guided vehicle on the driving surface at a facility; a payload handler coupled to the frame, the payload handler configured to transport a payload having a flat, non-deterministic seating surface seated on a payload holder to / from the payload holder of the autonomous guided vehicle and a storage location for the payload in a storage array; a vision system mounted on the frame, the vision system having at least one camera; a controller communicatively coupled to the vision system; providing generating video stream data imaging of an object within a logistics space using the at least one camera of the vision system, the object being at least one of at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics item or structure within the logistics space external to the autonomous guided vehicle; recording the video stream data imaging from the at least one camera with the controller; detecting the distance of the object using at least one or more of a time-of-flight sensor and a range sensor communicatively coupled to the controller; Including, The method, wherein the controller provides robust object detection and localization within a predetermined reference frame from the video stream data imaging alternatively via both binocular and monocular vision from the video stream data imaging, and detection determined via the monocular vision has equal reliability as detection determined via the binocular vision.
14. The method of claim 13 , wherein the controller uses a deep machine learning model to effect object detection via the monocular vision.
15. The method of claim 13 , wherein the controller interfaces with a machine learning module having object detection capabilities that determines object detection from the monocular vision and deep machine learning model.
16. 14. The method of claim 13, wherein a media server is communicatively connected to the at least one camera and records video stream data, the media server interfaces with the controller, the controller being located onboard the autonomous guided vehicle or remotely located from the autonomous guided vehicle.
17. The method of claim 13 , wherein the payload handler underpicks the payload from the storage location.
18. and wherein the at least one camera of the vision system includes two cameras forming a stereo vision camera pair, the method comprising:
14. The method of claim 13, further comprising: via said robust object detection and localization, causing under-picking by a payload handler of payloads from more than two closely spaced payloads held in adjacent storage locations regardless of availability of stereoscopic vision from said stereo vision camera pair.
19. and wherein the at least one camera of the vision system includes two cameras forming a stereo vision camera pair, the method comprising:
14. The method of claim 13, further comprising: via said robust object detection and localization, providing under-picking by a payload handler of a deformed payload from more than two closely spaced payloads held in adjacent storage locations, regardless of availability of stereoscopic vision from said stereo vision camera pair.
20. and wherein the at least one camera of the vision system includes two cameras forming a stereo vision camera pair, the method comprising:
14. The method of claim 13, further comprising: via the robust object detection and localization, providing under-picking by a payload handler of payloads from more than two payloads held in adjacent storage locations and having a dynamic Gaussian case size distribution within the facility, regardless of availability of stereoscopic vision from the stereo vision camera pair.
21. The method of claim 13 , wherein the object localization provided by the monocular vision is as reliable as the object localization determined via the binocular vision.
22. 14. The method of claim 13, wherein binocular object detection and location and monocular object detection and location are interchangeably selectable via the controller.
23. 14. The method of claim 13, wherein the controller comprises a selector arranged to select between binocular vision object detection and localization and monocular vision object detection and localization on demand based on detection of predetermined operating characteristics of the autonomous guided vehicle.
24. 24. The method of claim 23, wherein the predetermined characteristic is video stream data recorded by the controller that does not support object detection and location using the binocular vision.
25. An autonomous guided vehicle, the autonomous guided vehicle comprising: a frame having a payload holder; a drive section coupled to the frame, the drive section including drive wheels that support the autonomous guided vehicle on a driving surface, the drive wheels providing vehicle movement on the driving surface and moving the autonomous guided vehicle on the driving surface at a facility; a payload handler coupled to the frame, the payload handler configured to transport a payload having a flat, non-deterministic seating surface seated on the payload holder to / from the payload holder of the autonomous guided vehicle and a storage location for the payload in a storage array; a vision system mounted on the frame, the vision system having at least one camera positioned to generate video stream data imaging of an object within a logistics space, the object being at least one of at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics item or structure within the logistics space external to the autonomous guided vehicle; a controller communicatively connected to record video stream data imaging from the at least one camera and to at least one or more of a time-of-flight sensor and a distance sensor to detect the distance of the object; The controller is configured to cause object detection and localization within a predetermined reference frame to be derived from the video stream data imaging selectively using binocular vision and monocular vision from the video stream data imaging, and each of the binocular vision object detection and localization and the monocular vision object detection and localization is selectable on demand by the controller.
26. 26. The autonomous guided vehicle of claim 25, wherein the monocular vision detection and localization is as reliable as detection determined via the binocular vision object detection and localization.
27. 26. The autonomous guided vehicle of claim 25, wherein the controller is configured to interchangeably select between the binocular vision object detection and location and the monocular vision object detection and location.
28. 26. The autonomous guided vehicle of claim 25, wherein the controller comprises a selector arranged to select between the binocular vision object detection and location and the monocular vision object detection and location on demand based on detection of predetermined operating characteristics of the autonomous guided vehicle.
29. 30. The autonomous guided vehicle of claim 28, wherein the predetermined characteristic is video stream data imaging recorded by the controller that does not support object detection and location using the binocular vision.
30. 26. The autonomous guided vehicle of claim 25, wherein the controller is configured to use a deep machine learning model to effect object detection via the monocular vision.
31. 26. The autonomous guided vehicle of claim 25, wherein the controller is configured to interface with a machine learning module having object detection capabilities that determines object detection from the monocular vision and deep machine learning model.
32. 26. The autonomous guided vehicle of claim 25, further comprising a media server communicatively connected to the at least one camera and that records video stream data, the media server interfacing with the controller, the controller being located onboard the autonomous guided vehicle or remotely located from the autonomous guided vehicle.
33. 1. A method for object detection and localization for an autonomous guided vehicle, said method comprising: For autonomous guided vehicles, a frame having a payload holder; a drive section coupled to the frame, the drive section including drive wheels that support the autonomous guided vehicle on a driving surface, the drive wheels providing vehicle movement on the driving surface and moving the autonomous guided vehicle on the driving surface at a facility; a payload handler coupled to the frame, the payload handler configured to transport a payload having a flat, non-deterministic seating surface seated on the payload holder to / from the payload holder of the autonomous guided vehicle and a storage location for the payload in a storage array; a vision system mounted on the frame, the vision system having at least one camera positioned to generate video stream data imaging of an object within a logistics space, the object being at least one of at least a portion of the frame, at least a portion of the payload, at least a portion of the payload handler, and at least a portion of a logistics item or structure within the logistics space external to the autonomous guided vehicle; a controller communicatively connected to the vision system and to at least one or more of a time-of-flight sensor and a range sensor for detecting the distance of the object; providing recording the video stream data imaging from the at least one camera with the controller; and using the controller to selectively use binocular vision and monocular vision from the video stream data imaging to produce object detection and localization within a predetermined reference frame from the video stream data imaging, wherein each of the binocular vision object detection and localization and the monocular vision object detection and localization is selectable on demand by the controller.
34. 34. The method of claim 33, wherein the monocular detection and localization is as reliable as detection determined via the binocular object detection and localization.
35. 34. The method of claim 33, wherein the binocular object detection and location and the monocular object detection and location are interchangeably selectable by the controller.
36. 34. The method of claim 33, wherein the controller comprises a selector arranged to select between the binocular vision object detection and localization and the monocular vision object detection and localization on demand based on detection of predetermined operating characteristics of the autonomous guided vehicle.
37. 37. The method of claim 36, wherein the predetermined characteristic is video stream data imaging recorded by the controller that does not support object detection and location using the binocular vision.
38. 34. The method of claim 33, wherein the controller employs a deep machine learning model to effect object detection via the monocular vision.
39. 34. The method of claim 33, wherein the controller interfaces with a machine learning module having object detection capabilities that determines object detection from the monocular vision and deep machine learning model.
40. 34. The method of claim 33, wherein a media server is communicatively connected to the at least one camera and records video stream data, the media server interfaces with the controller, the controller being located onboard the autonomous guided vehicle or remotely located from the autonomous guided vehicle.