Logistics autonomous vehicles with robust object detection, localization, and monitoring
The autonomous guided vehicle uses a vision system with asynchronous cameras and advanced image analysis to overcome navigation and detection challenges in ultra-constrained environments, achieving high-accuracy localization and reduced picking failures.
Patent Information
- Application Number
- JP2025527813
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2023-11-14
- Publication Date
- 2025-11-28
AI Technical Summary
Existing autonomous vehicles in automated storage and retrieval systems face challenges in navigation and object detection due to limitations in stereo or binocular cameras, such as occlusion, obstruction, and degraded image processing, which affect their ability to accurately guide and locate vehicles in ultra-constrained environments with closely spaced, deformed, and irregularly arranged cases.
The autonomous guided vehicle employs a vision system with a pair of inexpensive, two-dimensional rolling-shutter, asynchronous cameras to generate binocular or stereo images, combined with a controller that analyzes image data to create dense depth maps and keypoint data for precise localization and object detection, enabling accurate pose and position determination in ultra-constrained systems.
This approach enhances the vehicle's ability to identify and locate cases with high accuracy, reducing picking failures to approximately 1 failure per 1 million picks, even in environments with closely spaced, deformed, and irregularly arranged cases, while maintaining efficient transfer times.
Smart Images

Figure 2025538388000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application is a non-provisional application of and claims the benefit of U.S. Provisional Patent Application No. 63 / 383,597, filed November 14, 2022, the entire disclosure of which is incorporated herein by reference.
[0002] [Technical field] The disclosed embodiments relate generally to material handling systems, and more particularly to conveyors for automated logistics systems. [Background technology]
[0003] Generally, automated logistics systems, such as automated storage and retrieval systems, utilize autonomous vehicles that transport goods within the automated storage and retrieval system. These autonomous vehicles are guided throughout the automated storage and retrieval system by location beacons, capacitive or inductive proximity sensors, line following sensors, reflective beam sensors, and other narrow-focus beam sensors. These sensors may provide limited information to effect navigation of the autonomous vehicles through the storage and retrieval system or may provide limited information regarding the identification and discrimination of hazardous materials that may be present throughout the automated storage and retrieval system.
[0004] Autonomous vehicles may also be guided throughout automated storage and retrieval systems by vision systems employing stereo or binocular cameras. However, the binocular cameras of these binocular vision systems are positioned at distances relative to each other that are not suitable for storing and retrieving warehouse logistics cases. In logistics environments, stereo or binocular cameras may be impaired, for example, by occlusion or obstruction and / or unclear visibility of one camera of a set of stereo cameras (e.g., by a payload carried by the autonomous vehicle, a storage structure, etc.), or may not always be available, or image processing quality may be degraded from processing overlapping image data or images that are otherwise not suitable (e.g., blurred, etc.) for guiding and locating an autonomous vehicle within an automated storage and retrieval system. Summary of the Invention
[0005] The foregoing aspects and other features of the disclosed embodiments are explained in the following description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0006] [Figure 1A] 1 is a schematic diagram of a logistics facility incorporating aspects of the disclosed embodiments; [Figure 1B] FIG. 1B is a schematic diagram of the logistics facility of FIG. 1A in accordance with aspects of the disclosed embodiment; [Figure 2] FIG. 1B is a schematic diagram of an autonomous guided vehicle of the logistics facility of FIG. 1A in accordance with aspects of the disclosed embodiment; [Figure 3A] FIG. 3 is a schematic diagram of a portion of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 3B] FIG. 3 is a schematic diagram of a portion of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 3C] FIG. 3 is a schematic diagram of a portion of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 4A] 3 is an example of image data captured by a vision system of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 4B] 3 is an example of image data captured by a vision system of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 4C] 3 is an example of image data captured by a vision system of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 5] FIG. 3 is a schematic diagram of a portion of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment; [Figure 6] 1 is an exemplary illustration of a dense depth map generated from a pair of stereo images in accordance with aspects of the disclosed embodiment; [Figure 7] FIG. 10 is an exemplary diagram of a stereo set of keypoints in accordance with aspects of the disclosed embodiment; [Figure 8] FIG. 10 illustrates an exemplary flow diagram of keypoint detection for one image of a stereo image pair in accordance with aspects of the disclosed embodiment; [Figure 9] FIG. 10 is an exemplary flow diagram of keypoint detection for a stereo image pair according to aspects of the disclosed embodiment; [Figure 10] FIG. 10 is an exemplary flow diagram of plane estimation for a facial surface of an object according to aspects of the disclosed embodiment; [Figure 11] 1B is a schematic diagram of a stereo vision calibration station of the logistics facility of FIG. 1A in accordance with aspects of the disclosed embodiment; [Figure 12] FIG. 12 is a schematic diagram of a portion of the calibration station of FIG. 11 in accordance with aspects of the disclosed embodiment; [Figure 13] FIG. 3 is an exemplary schematic diagram of a model of the autonomous guided vehicle of FIG. 2 in accordance with aspects of the disclosed embodiment. [Figure 14] FIG. 10 is an exemplary flow diagram of a method according to aspects of the disclosed embodiment; [Figure 15] FIG. 10 is an exemplary flow diagram of a method according to aspects of the disclosed embodiment; DETAILED DESCRIPTION OF THE INVENTION
[0007] 1A and 1B illustrate an exemplary automated storage and retrieval system 100 in accordance with aspects of the disclosed embodiment. While aspects of the disclosed embodiment will be described with reference to the drawings, it should be understood that aspects of the disclosed embodiment can be embodied in many forms. Furthermore, any suitable size, shape, or type of elements or materials may be used.
[0008] Aspects of the disclosed embodiments provide a logistics autonomous guided vehicle 110 (referred to herein as an autonomous guided vehicle) with intelligent autonomy and coordinated operation. For example, the autonomous guided vehicle 110 includes a vision system 400 (see FIG. 2) having at least one (or multiple) cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B arranged to generate a binocular or stereo image of a field commonly imaged by each of the at least one (or multiple) cameras that generates a binocular image (the binocular stereo image can be video stream data images or still image data) of a logistics space (such as the operating environment or space of the storage and retrieval system 100) including shelves 555 (see FIGS. 1B, 3A, and 4B) of a rack structure on which a plurality of objects (such as cases CU) are stored. The commonly imaged field is formed by the combination of the individual fields 410AF, 410BF, 420AF, 420BF, 430AF, 430BF, 460AF, 460BF, 477AF, 477BF of each camera pair 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, and / or 477A and 477B. For example, with respect to a stereo or binocular imaging camera (where "stereo or binocular imaging camera" is generally referred to herein as a "stereo imaging camera" which may be a camera pair or more than two cameras that generate a stereo image), such as case unit monitoring cameras 410A, 410B, the commonly imaged field is the combination of the respective fields of view 410AF, 410BF. For illustrative purposes, vision system 400 employs at least stereo or binocular vision configured to provide detection of cases CUs and objects (such as facility structures and unwanted foreign / transient objects) within a logistics facility such as automated storage and retrieval system 100. The stereo or binocular vision is also configured to provide localization of autonomous guided vehicles within automated storage and retrieval system 100.The vision system 400 also provides collaborative vehicle operations by providing images (still images or video streams, live or recorded) to the operator of the automated storage and retrieval system 100, where these images are provided, in some embodiments, via a user interface UI as augmented images as described herein.
[0009] As described in more detail herein, the autonomous guided vehicle 110 includes a controller 122 that is programmed to access data from the vision system 400 to provide robust case / object detection and localization within an ultra-constrained system or operating environment using at least a pair of inexpensive, two-dimensional rolling-shutter, asynchronous cameras (in other embodiments, the camera pair may include relatively more expensive, two-dimensional global shutter cameras that may or may not be synchronized with each other) and while the autonomous guided vehicle 110 is moving relative to the cases / objects. An ultra-constrained system includes, but is not limited to, at least the following constraints: closely spaced spacing between dynamically positioned adjacent cases (also referred to herein as tightly packed juxtaposition relative to each other), the autonomous guided vehicle being configured to under-pick cases (lift them from below), cases of various sizes being distributed in a Gaussian distribution within the storage array SA, cases may exhibit deformation, and cases may be arranged in an irregular manner on the support surface, all of which affect the transfer of case units CU between the storage shelves 555 (or other case holding locations) and the autonomous guided vehicle 110.
[0010] The cases CU stored in the storage and retrieval system have a Gaussian distribution (see FIG. 4A ) with respect to case size within picking aisle 130A and with respect to case size across storage array SA, such that the size of any given storage space on storage shelf 555 dynamically changes (e.g., a dynamic Gaussian case size distribution) as cases are picked and placed. As such, autonomous guided vehicle 110 is configured to determine or otherwise identify cases held in dynamically sized storage spaces (depending on the cases held), as described herein, regardless of movement of the autonomous guided vehicle relative to the stored cases.
[0011] Further, as seen in FIG. 4A , for example, the cases CU are arranged on the storage shelf 555 (or other storage station) in a closely coupled or closely spaced relationship, where the distance DIST between adjacent case units CU is approximately half the distance between the storage shelf hats 444. The distance / width DIST between the hats 444 of the support slats 520L is approximately 2.5 inches. The close spacing of the case CUs can be complicated (i.e., the spacing can be less than half the distance between the storage shelf hats 444) in that the case CUs (e.g., see FIGS. 4A-4C illustrating a deformed case—open flap case deformation) may exhibit deformations (e.g., bulging sides, open flaps, convex sides, etc.) and / or may be tilted relative to the hats 444 on which they rest (i.e., the front of the case may not be parallel to the front of the storage shelf 555 and the sides of the case may not be parallel to the hats 555 of the storage shelf 555—see FIG. 4A ). Case deformation and tilted case placement may further reduce the spacing between adjacent cases. As such, the autonomous guided vehicle is configured to determine or otherwise identify the pose and position of the cases using at least one pair of inexpensive, two-dimensional, rolling-shutter, asynchronous cameras in an ultra-constrained system, as described herein, for case transfer (e.g., picked from and placed in storage) without substantial interference between closely spaced adjacent cases, regardless of the movement of the autonomous guided vehicle relative to the cases / objects.
[0012] It is also noted that the height HGT of the hats 444 is approximately 2 inches, and the space envelope ENV between the hats 444, through which the tines 210AT of the transfer arm 210A of the autonomous guided vehicle 110 are inserted under the case units CU to pick / place cases from / to the storage shelves 555, has a width of approximately 1.7 inches and a height of approximately 1.2 inches (see, for example, Figures 3A, 3C, and 4A). Under-picking of a case CU by an autonomous guided vehicle requires interfacing with a case CU held on a storage shelf 555 at a picking / case support surface (defined by case seating surfaces 444S of hats 444 (see FIG. 4A )) without collision between the tines 210AT of the transfer arm 210A of the autonomous guided vehicle 110 and the hats 444 / slats 520L, without collision between the tines 210AT and adjacent cases (not being picked), and without collision between the case being picked and adjacent cases not being picked, all of which is effected by the placement of the tines 210AT in the envelope ENV between the hats 444. To this end, the autonomous guided vehicle is configured to detect and locate the spatial envelope ENV for inserting the tines 210AT of the transfer arm 210A under a given case CU to pick the case, using at least one pair of inexpensive, two-dimensional, rolling-shutter, asynchronous cameras as described herein.
[0013] Another constraint of the ultra-constrained system is the transfer time required for the autonomous guided vehicle 110 to transfer a case unit(s) between the payload platform 210B of the autonomous guided vehicle 110 and a case holding location (e.g., a storage space, a buffer, a transfer station, or other case holding location described herein), where the transfer time for a case transfer is about 10 seconds or less. Therefore, the vision system 400 identifies the location and pose of the case (or the location and pose of the holding station) in less than about 2 seconds, or less than about 0.5 seconds.
[0014] The above hyper-constraint system requires robustness of the vision system and can be considered to define the robustness of the vision system 400 since the vision system 400 is configured to accommodate the above-mentioned constraints, and can provide pose and localization information for the case CU and / or the autonomous guided vehicle 110 resulting in an autonomous guided vehicle picking failure rate of approximately 1 picking failure per approximately 1 million picks.
[0015] According to aspects of the disclosed embodiment, autonomous guided vehicle 110 includes a controller (e.g., controller 122 or a vision system controller 122VC communicatively coupled to or forming part of controller 122) that registers image data (e.g., video streams) from cameras in one or more camera pairs (e.g., a camera pair formed by each of cameras 410A, 410B, 420A, 420B, 430A, 430B, 460A, 460B, 477A, 477B). The controller is configured to analyze the registered (video) image data into individual registered (still) image frames to form a set of still stereoscopic image frames (as opposed to motion video from which images are analyzed) (e.g., see image frames 600A, 600B in FIG. 6 for an example) for each camera pair (e.g., camera pair 410A, 410B illustrated in FIG. 3A ).
[0016] As described herein, the controller generates dense depth maps of objects within the fields of view of the cameras in the camera pair from the stereo vision frames and determines the position and pose of the imaged objects from the dense depth maps. The controller also generates binocular keypoint data for the stereo vision frames, the keypoint data being separate and distinct from the dense depth maps, where the keypoint data provides (e.g., binocular, three-dimensional) determination of the position and pose of the objects within the fields of view of the cameras. Although the term "keypoint" is used herein, the keypoints described herein are also referred to in the art as "feature point(s)," "invariant feature(s)," "invariant point(s)," or "features" (such as corners, facet joints, or object surfaces). The controller weights the keypoint data and combines the dense depth map with the keypoint data to determine or otherwise identify the pose and position of the imaged object (e.g., within the logistics space and / or relative to the autonomous guided vehicle 110) with greater accuracy than the pose and position determination accuracy of the dense depth map alone and greater accuracy than the pose and position determination accuracy of the keypoint data alone.
[0017] According to aspects of the disclosed embodiment, the automated storage and retrieval system 100 of FIGS. 1A and 1B may be deployed in a retail distribution (logistics) center or warehouse to fulfill orders received from retailers for replenishment items shipped in cases, packages, and / or parcels, for example. The terms case, package, and parcel are used interchangeably herein and, as previously mentioned, may refer to any container that may be used for shipping and that may be filled with one or more product units by a manufacturer. A case(s) as used herein refers to a case, package, or parcel unit that is not stored (e.g., not contained) in a tray, on a tote, or the like. It is noted that a case unit CU (also referred to herein as a mixed case, case, and shipping unit) may include a case of items / units (e.g., a case of soup cans, cereal boxes, etc.) or individual items / units adapted to be removed from or placed on a pallet. According to an exemplary embodiment, shipping cases or case units (e.g., cartons, barrels, boxes, crates, jugs, shrink-wrapped trays or groups, or any other suitable device for holding case units) may have variable sizes, may be used to hold case units during shipping, and may be configured to be palletizable for shipping. Case units may also contain totes, boxes, and / or containers of one or more individual items (generally referred to as break-pack items) that have been opened / released from their original packaging and placed in totes, boxes, and / or containers (collectively referred to as totes) with one or more other individual items of a mixed or common type at an order filling station.For example, it is noted that when incoming bundles or pallets (e.g., from a case unit manufacturer or supplier) arrive at the automated storage and retrieval system 100 for replenishment, the contents of each pallet may be uniform (e.g., each pallet holds a predetermined number of the same items, i.e., one pallet holds soup, another pallet holds cereal). As can be appreciated, the cases in such a pallet load may be generally similar, or in other words, homogenous cases (e.g., similar dimensions) and may have the same SKUs (otherwise, as previously mentioned, the pallet may be a “rainbow” pallet with layers formed of homogenous cases). Once the pallet exits the automated storage and retrieval system, with the cases or totes filled with replenishment orders, the pallet may contain any suitable number and combination of different case units (e.g., each pallet may hold different types of case units, i.e., the pallet holds a combination of canned soup, cereal, drink cartons, cosmetics, and household cleaners). The cases combined on a single pallet may have different dimensions and / or different SKUs.
[0018] Automated storage and retrieval system 100 may generally be described as a storage and retrieval engine 190 coupled to a palletizer 162. Now in more detail, and still referring to Figures 1A and 1B, automated storage and retrieval system 100 may be configured for installation in an existing warehouse structure or adapted to a new warehouse structure, for example. As previously mentioned, automated storage and retrieval system 100 shown in Figures 1A and 1B is representative and may include, for example, infeed and outfeed conveyors terminating in respective transfer stations 170, 160, lift module(s) 150A, 150B, a storage structure 130, and several autonomous guided vehicles 110. It is noted that storage and retrieval engine 190 is formed by at least storage structure 130 and autonomous guided vehicle 110 (and in some embodiments also by lift modules 150A, 150B, while in other embodiments lift modules 150A, 150B may form a vertical sequencer in addition to storage and retrieval engine 190, as described in U.S. Patent Application No. 17 / 091,265, filed November 6, 2020, and entitled "Pallet Building System with Flexible Sequencing," the entire disclosure of which is incorporated herein by reference). In alternative embodiments, automated storage and retrieval system 100 may include a robot or bot transfer station (not shown) that may provide an interface between autonomous guided vehicle 110 and lift module(s) 150A, 150B. The storage structure 130 may include multiple levels of storage rack modules, where each storage structure level 130L of the storage structure 130 includes a respective picking aisle 130A and a transfer deck 130B for transporting case units between any of the storage areas of the storage structure 130 and the shelves of the(s) lift module(s) 150A, 150B.In one embodiment, picking aisle 130A is configured to provide guided movement of autonomous guided vehicle 110 (such as along rail 130AR), while in other embodiments, picking aisle 130A is configured to provide unconstrained movement of autonomous guided vehicle 110 (e.g., picking aisle 130A is open and non-deterministic with respect to the guidance / movement of autonomous guided vehicle 110). Transfer deck 130B has an open, non-deterministic bot-supporting movement surface along which autonomous guided vehicle 110 moves under guidance and control provided by any suitable bot steering. In one or more embodiments, transfer deck 130B has multiple lanes between which autonomous guided vehicle 110 transitions freely to access picking aisle 130A and / or lift modules 150A, 150B. As used herein, "open and non-deterministic" indicates that the movement surface of the picking aisle and / or transfer deck has no mechanical constraints (such as guide rails) that limit the movement of the autonomous guided vehicle 110 to any given path along the movement surface.
[0019] Picking aisle 130A and transfer deck 130B also enable autonomous guided vehicle 110 to place case units CU in picking stock and retrieve ordered case units CU (and define various locations at which the bot performs its autonomous tasks, although any number of locations within the storage structure (e.g., deck, aisle, storage rack, etc.) could be one or more of the various locations). In an alternative embodiment, each level may include a respective transfer station 140 that provides indirect case transfer between autonomous guided vehicle 110 and lift modules 150A, 150B. Autonomous guided vehicle 110 may be configured to place case units, such as the retail items described above, in picking stock in one or more storage structure levels 130L of storage structure 130, and then selectively retrieve ordered case units and ship the ordered case units, for example, to a store or other suitable location. The infeed transfer station 170 and the outfeed transfer station 160 may operate in conjunction with respective lift module(s) 150A, 150B to transfer case units CU bidirectionally to / from one or more storage structure levels 130L of the storage structure 130. While the lift modules 150A, 150B may be described as dedicated inbound lift module 150A and outbound lift module 150B, it is noted that in alternative embodiments, each of the lift modules 150A, 150B may be used for both inbound and outbound transfer of case units from the automated storage and retrieval system 100.
[0020] As can be appreciated, the automated storage and retrieval system 100 may include multiple infeed and outfeed lift modules 150A, 150B that are accessible, for example, by the autonomous guided vehicles 110 of the automated storage and retrieval system 100 (e.g., indirectly via transfer station 140 or via direct case transfer between the lift modules 150A, 150B and the autonomous guided vehicles 110) so that one or more uncontained case units (e.g., case unit(s) not held in a tray) or one or more contained case units (in a tray or tote) can be transferred from the lift modules 150A, 150B to each storage space on each level, and from each storage space on each level to any one of the lift modules 150A, 150B. Autonomous guided vehicle 110 may be configured to transfer cases CU (also referred to herein as case units) between storage space 130S (e.g., located in picking aisle 130A or other suitable storage space / case unit buffer located along transfer deck 130B) and lift modules 150A, 150B. Generally, lift modules 150A, 150B include at least one movable payload support that can move case unit(s) between infeed and outfeed transfer stations 160, 170 and respective levels of the storage space where the case unit(s) are stored and retrieved. The lift module(s) may have any suitable configuration, such as, for example, a reciprocating lift, or any other suitable configuration.The lift module(s) 150A, 150B may include any suitable controller (such as the control server 120 or other suitable controller coupled to the control server 120, the warehouse management system 2500, and / or the palletizer controllers 164, 164′) to form a sequencer or sorter in a manner similar to that described in U.S. patent application Ser. No. 16 / 444,592, filed June 18, 2019, and entitled “Vertical Sequencer for Product Order Fulfillment,” the disclosure of which is incorporated herein by reference in its entirety.
[0021] Automated storage and retrieval system 100 may include a control system comprising, for example, one or more control servers 120 communicatively connected to infeed and outfeed conveyors and transfer stations 170, 160, lift modules 150A, 150B, and autonomous guided vehicle 110 via a suitable communication and control network 180. Communication and control network 180 may have any suitable architecture, for example, incorporating various programmable logic controllers (PLCs), such as for commanding the operation of the automation of infeed and outfeed conveyors and transfer stations 170, 160, lift modules 150A, 150B, and other suitable systems. Control server 120 may include high-level programming to provide a case management system (CMS) that manages the case flow system. Network 180 may further include suitable communications to provide a bidirectional interface with autonomous guided vehicle 110. For example, autonomous guided vehicle 110 may include an on-board processor / controller 122. Network 180 may include a suitable two-way communication suite that enables autonomous guided vehicle controllers 122 to request or receive commands from control server 120 to effect the desired transport of case units (e.g., placement into or removal from a storage location) and to transmit desired autonomous guided vehicle 110 information and data to control server 120, including autonomous guided vehicle 110 ephemeris, status, and other desired data. As seen in FIGS. 1A and 1B , control server 120 may further be connected to a warehouse management system 2500 for, for example, providing inventory management and customer order fulfillment information to control server 120's CMS-level programs. As previously mentioned, control server 120 and / or warehouse management system 2500 enable at least some degree of collaborative control of autonomous guided vehicles 110 via a user interface UI, as described further below.A suitable example of an automated storage and retrieval system arranged to hold and store case units is described in U.S. Pat. No. 9,096,375, issued August 4, 2015, the entire disclosure of which is incorporated herein by reference.
[0022] 1A, 1B, and 2, autonomous guided vehicle 110 includes a frame 200 with an integrated payload support or payload platform 210B (also referred to as a payload holder or payload bay). Frame 200 has a front end 200E1 and a back end 200E2 that define a longitudinal axis LAX of autonomous guided vehicle 110. Frame 200 may be constructed of any suitable material (e.g., steel, aluminum, composite materials, etc.) and includes a case handling assembly 210 configured to handle cases / payloads to be transported by autonomous guided vehicle 110. Case handling assembly 210 includes payload platform 210B on which payloads are placed for transport and / or any suitable transfer arm 210A (also referred to as a payload handler) connected to the frame. Transfer arm 210A is configured to (autonomously) transfer a payload (such as a case unit CU) having a flat, non-deterministic seating surface that seats on payload platform 210B to / from payload platform 210B of autonomous guided vehicle 110, and to / from a storage location for payload CU within storage array SA (such as storage space 130S on storage shelf 555 (see FIG. 2), a shelf, buffer, transfer station, and / or any other suitable storage location), where storage location 130S within storage array SA is separate and distinct from transfer arm 210A and payload platform 210B. Transfer arm 210A is configured to extend in a lateral direction LAT and / or a vertical direction VER to transport payloads to / from payload platform 210B. Examples of suitable payload platforms 210B and transfer arms 210A and / or autonomous guided vehicles 110 to which aspects of the disclosed embodiments may be applied are described in U.S. Pat. No. 1,107,801, entitled "Automated Bot with Transfer Arm," issued Aug. 3, 2021, and entitled "Materials-Handling System Using Autonomous Transfer and Transport," the entire disclosure of which is incorporated herein by reference.U.S. Patent No. 7,591,630, issued September 22, 2009, entitled "Materials-Handling System Using Autonomous Transfer and Transport Vehicles," U.S. Patent No. 7,991,505, issued August 2, 2011, entitled "Autonomous Transport Vehicle," U.S. Patent No. 9,561,905, issued February 7, 2017, entitled "Autonomous Transport Vehicle," U.S. Patent No. 9,082,112, issued July 14, 2015, entitled "Autonomous Transport Vehicle Charging System," U.S. Patent No. 9,850,079, issued December 26, 2017, entitled "Storage and Retrieval System Transport Vehicle," U.S. Patent No. 9,187,244, issued November 17, 2015, entitled "Bot Payload Alignment and Sensing," and U.S. Patent No. 9,187,244, issued November 17, 2015, entitled "Automated Bot Transfer Arm Drive" No. 9,499,338, issued November 22, 2016, entitled "Bot Having High Speed Stability System," U.S. Patent No. 8,965,619, issued February 24, 2015, entitled "Bot Position Sensing," U.S. Patent No. 9,008,884, issued April 14, 2015, entitled "Bot Position Sensing," U.S. Patent No. 8,425,173, issued April 23, 2013, entitled "Autonomous Transports for Storage and Retrieval Systems," and U.S. Patent No. 8,696,010, issued April 15, 2014, entitled "Suspension System for Autonomous Transports."
[0023] Frame 200 includes one or more idler wheels or casters 250 positioned adjacent to front end 200E1. Suitable examples of casters can be found in U.S. Patent Application No. 17 / 664,948 (Attorney Docket No. 1127P015753-US(PAR)), filed May 25, 2022, entitled "Autonomous Transport Vehicle with Synergistic Vehicle Dynamic Response," and U.S. Patent Application No. 17 / 664,838 (Attorney Docket No. 1127P015753-US(PAR)), filed May 26, 2021, entitled "Autonomous Transport Vehicle with Steering," the entire disclosures of which are incorporated herein by reference. Frame 200 also includes one or more drive wheels 260 positioned adjacent to back end 200E2. In other aspects, the positions of caster 250 and drive wheel 260 may be reversed (e.g., drive wheel 260 is located on front end 200E1 and caster 250 is located on back end 200E2). It is noted that in some aspects, autonomous guided vehicle 110 is configured to move with front end 200E1 leading the direction of movement or with back end 200E2 leading the direction of movement. In one aspect, casters 250A, 250B (substantially similar to caster 250 described herein) are positioned at respective front corners at front end 200E1 of frame 200, and drive wheels 260A, 260B (substantially similar to drive wheel 260 described herein) are positioned at respective back corners at back end 200E2 of frame 200 (e.g., support wheels are positioned at each of the four corners of frame 200), thereby allowing autonomous guided vehicle 110 to stably travel on transfer deck 130B and picking aisle 130A of storage structure 130.
[0024] Autonomous guided vehicle 110 includes a drive section 261D connected to frame 200, with drive wheels 260 that support autonomous guided vehicle 110 on running / rolling surface 284, and drive wheels 260 provide vehicle movement on running surface 284 and move autonomous guided vehicle 110 on running surface 284 within a facility (e.g., warehouse, store, etc.). Drive section 261D has at least a pair of traction drive wheels 260 (also referred to as drive wheels 260 (see drive wheels 260A, 260B)) straddling drive section 261D. The drive wheels 260 have a fully independent suspension 280 that couples each drive wheel 260A, 260B of at least one pair of drive wheels 260 to the frame 200 and is configured to maintain a substantially steady-state traction contact patch between at least one drive wheel 260A, 260B and a rolling / moving surface 284 (also referred to as an autonomous vehicle moving surface 284) over rolling surface transients (e.g., bumps, surface transitions, etc.). A suitable example of a fully independent suspension 280 can be found in U.S. Patent Application No. 17 / 664,948, entitled "Autonomous Transport Vehicle with Synergistic Vehicle Dynamic Response," filed May 25, 2022 (having Attorney Docket No. 1127P015753-US(PAR)), the entire disclosure of which was previously incorporated herein by reference.
[0025] Autonomous guided vehicle 110 includes a physical property sensor system 270 (also referred to as an autonomous navigation operation sensor system) connected to frame 200. Physical property sensor system 270 has electromagnetic sensors. Each of the electromagnetic sensors is responsive to an interaction or interface between an electromagnetic beam or field emitted or generated by the sensor and a physical property (e.g., of a storage structure or case unit CU, debris, or other ephemeral object), where the electromagnetic beam or field is perturbed by the interaction or interface with the physical property. The perturbation of the electromagnetic beam is detected by the electromagnetic sensor, resulting in the sensing of a physical property, where physical property sensor system 270 is configured to generate sensor data embodying at least one of vehicle navigation pose or position information (with respect to a storage and retrieval system or facility in which autonomous guided vehicle 110 operates) and payload pose or position information (with respect to storage location 130S or payload platform 210B).
[0026] Physical property sensor system 270 includes, by way of example only, one or more of laser sensor(s) 271, ultrasonic sensor(s) 272, barcode scanner(s) 273, position sensor(s) 274, line sensor(s) 275, a plurality of case sensors 278 (e.g., for detecting case units in payload berth 210B onboard vehicle 110 or on storage shelves offboard vehicle 110), arm proximity sensor(s) 277, vehicle proximity sensor(s) 278, or any other suitable sensors for detecting the position of vehicle 110 or payload (e.g., case unit CU). In some aspects, supplemental navigation sensor system 288 may form part of physical property sensor system 270. Suitable examples of sensors that may be included in the physical property sensor system 270 are described in U.S. Pat. No. 8,425,173, entitled "Autonomous Transport for Storage and Retrieval Systems," issued April 23, 2013; U.S. Pat. No. 9,008,884, entitled "Bot Position Sensing," issued April 14, 2015; and U.S. Pat. No. 9,946,265, entitled "Bot Having High Speed Stability," issued April 17, 2018, the entire disclosures of which are incorporated herein by reference.
[0027] The sensors of physical property sensor system 270 may be configured to provide autonomous guided vehicle 110 with, for example, awareness of its environment and external objects, as well as monitoring and control of internal subsystems. For example, the sensors may provide guidance information, payload information, or any other suitable information for use in operating autonomous guided vehicle 110.
[0028] Barcode scanner(s) 273 may be mounted in any suitable location on autonomous guided vehicle 110. Barcode scanner(s) 273 may be configured to provide an absolute position of autonomous guided vehicle 110 within storage structure 130. Barcode scanner(s) 273 may be configured to verify aisle references and positions on the transfer deck, for example, by reading barcodes located on the transfer deck, picking aisles, and transfer station floors to verify the position of autonomous guided vehicle 110. Barcode scanner(s) 273 may also be configured to read barcodes located on items stored on shelves 555.
[0029] Position sensor(s) 274 may be mounted in any suitable location on autonomous guided vehicle 110. Position sensor 274 may be configured, for example, to detect reference datum features (or count slats 520L of storage shelf 555) (e.g., see FIG. 5A ) to determine the position of vehicle 110 relative to shelves in picking aisle 130A (or transfer deck 130B or a buffer / transfer station located adjacent to lift 150). Reference datum information may be used by controller 122, for example, to correct vehicle odometry and stop autonomous guided vehicle 110 with support tines 210AT of transfer arm 210A positioned for insertion into the spaces between slats 520L (e.g., see FIG. 5A ). In one exemplary embodiment, the vehicle 110 may include position sensors 274 at the driving (rear) end 200E2 and the driven (front) end 200E1 of the autonomous guided vehicle 110 to enable reference datum detection regardless of which end of the autonomous guided vehicle 110 is facing in the direction the autonomous guided vehicle 110 is moving.
[0030] The line sensor 275 may be any suitable sensor mounted to the autonomous guided vehicle 110 in any suitable location, such as, by way of example only, on the frame 200 located adjacent the drive (rear) end 200E2 and the driven (front) end 200E1 of the autonomous guided vehicle 110. By way of example only, the line sensor 275 may be a diffuse infrared sensor. The line sensor 275 may be configured to detect guide lines 199 (see FIG. 1B ), for example, provided on the floor of the transfer deck 130B. The autonomous guided vehicle 110 may be configured to follow the guide lines when traveling on the transfer deck 130B and to define the end of a turn as the vehicle transitions onto or off the transfer deck 130B. The line sensor 275 may also enable the vehicle 110 to detect an index reference to determine absolute localization, where the index reference is generated by the crossed guide lines 119 (see FIG. 1B ).
[0031] The case sensor 276 may include a case overhang sensor and / or other suitable sensor configured to detect the position / pose of the case unit CU within the payload platform 210B. The case sensor 276 may be any suitable sensor positioned on the vehicle such that the field of view(s) of the sensor(s) spans the payload platform 210B adjacent the top surface of the support tines 210AT (see FIGS. 3A and 3B). The case sensor 276 may be positioned at an edge of the payload platform 210B (e.g., adjacent the transport opening 1199 of the payload platform 210B) to detect a case unit CU that is at least partially extending outside the payload platform 210B.
[0032] Arm proximity sensor 277 may be mounted to autonomous guided vehicle 110 in any suitable location, such as, for example, on transfer arm 210A. Arm proximity sensor 277 may be configured to detect objects around transfer arm 210A and / or support tines 210AT of transfer arm 210A as transfer arm 210A is raised / lowered and / or support tines 210AT are extended / retracted.
[0033] Laser sensor 271 and ultrasonic sensor 272 may be configured to enable autonomous guided vehicle 110 to locate itself with respect to each case unit forming a load carried by autonomous guided vehicle 110 before the case unit is picked, for example, from storage shelf 555 and / or lift 150 (or any other location suitable for retrieving a payload). Laser sensor 271 and ultrasonic sensor 272 also enable the vehicle to position itself with respect to empty storage locations 130S and place case units in those empty storage locations 130S. Laser sensor 271 and ultrasonic sensor 272 also enable autonomous guided vehicle 110 to verify that a storage space (or other load placement location) is empty before a load carried by autonomous guided vehicle 110 is placed, for example, in storage space 130S. In one example, laser sensor 271 may be mounted on autonomous guided vehicle 110 at a suitable location for detecting the edge of an item being transferred to (or from) autonomous guided vehicle 110. Laser sensor 271 may operate in conjunction with, for example, retroreflective tape (or other suitable reflective surface, coating, or material) disposed, for example, on the back of shelf 555, to enable the sensor to "see" all the way to the back of storage shelf 555. The reflective tape disposed on the back of the storage shelf renders laser sensor 1715 substantially unaffected by the color, reflectivity, roundness, or other suitable characteristics of the items located on shelf 555. Ultrasonic sensor 272 may be configured to measure the distance from autonomous guided vehicle 110 to a first item within a predetermined storage area of shelf 555 to enable autonomous guided vehicle 110 to determine the picking depth (e.g., the distance support tine 210AT travels into shelf 555 to pick item(s) from shelf 555). One or more of laser sensor 271 and ultrasonic sensor 272 may enable detection of case orientation (e.g., tilt of a case within storage shelf 555) by, for example, measuring the distance between autonomous guided vehicle 110 and the front of the case unit to be picked when autonomous guided vehicle 110 stops adjacent to the case unit to be picked.The case sensor may allow verification of the placement of a case unit, for example, on a storage shelf 555, for example, by scanning the case unit after placing it on the shelf.
[0034] A vehicle proximity sensor 278 may also be disposed on frame 200 to determine the position of autonomous guided vehicle 110 within picking aisle 130A and / or relative to lift 150. Vehicle proximity sensor 278 is disposed on autonomous guided vehicle 110 to detect targets or position-determining features disposed on rail 130AR along which vehicle 110 travels through picking aisle 130A (and / or on the walls of transfer area 195 and / or on the lift 150 access location). The positions of the targets on rail 130AR are at known positions that form incremental or absolute encoders along rail 130AR. Vehicle proximity sensor 278 detects the targets and provides sensor data to controller 122, which then determines the position of autonomous guided vehicle 110 along picking aisle 130A based on the detected targets.
[0035] The sensors of physical characteristic sensing system 270 are communicatively coupled to controller 122 of autonomous guided vehicle 110. As described herein, controller 122 is operably connected to drive section 261D and / or transfer arm 210A. Controller 122 is configured to determine the vehicle's pose and position (e.g., in up to six degrees of freedom: X, Y, Z, Rx, Ry, Rz) from the information of physical characteristic sensor system 270 to provide independent guidance of autonomous guided vehicle 110 through storage and retrieval facility / system 100. Controller 122 is also configured to determine the pose and position (onboard or offboard autonomous guided vehicle 110) of payloads (e.g., case units CU) from the information of physical characteristic sensor system 270 to provide independent underpicking (e.g., lifting case units CU from under case units CU) and placement of payloads CU from / to storage location 130S, and independent underpicking and placement of payloads CU on payload platform 210B.
[0036] 1A, 1B, 2, 3A, and 3B, as described above, autonomous guided vehicle 110 includes a supplemental or auxiliary navigation sensor system 288 connected to frame 200. Supplemental navigation sensor system 288 supplements physical property sensor system 270. Supplemental navigation sensor system 288 is, at least in part, a vision system 400 comprising a camera positioned to capture image data informing at least one of a vehicle navigation pose or position (relative to a structure or facility of the storage and retrieval system in which vehicle 110 operates) and a payload pose or position (relative to a storage location or payload platform 210B), which supplements the information of physical property sensor system 270. It is noted that the term “camera” as used herein refers to a still and / or video imaging device, including one or more of a two-dimensional camera and a two-dimensional camera having RGB (red, green, blue) pixels, non-limiting examples of which are provided herein. For example, as described herein, the two-dimensional cameras (with or without RGB pixels) are inexpensive two-dimensional rolling-shutter asynchronous cameras (e.g., compared to global shutter cameras) (although in other embodiments the cameras may be global shutter cameras that may or may not be synchronized with each other). In other embodiments, the two-dimensional rolling-shutter cameras, for example in a camera pair, may be synchronized with each other. Non-limiting examples of two-dimensional cameras include commercially available (i.e., "off-the-shelf") USB cameras each having 0.3 megapixels and a resolution of 640x480, MIPI Camera Serial Interface 2 (MIPI CSI-2®) cameras each having 8 megapixels and a resolution of 1280x720, or any other suitable camera.
[0037] 2, 3A, and 3B, the vision system 400 includes one or more of case unit monitoring cameras 410A, 410B, forward navigation cameras 420A, 420B, rear navigation cameras 430A, 430B, one or more three-dimensional imaging systems 440A, 440B, one or more case edge detection sensors 450A, 450B, one or more traffic monitoring cameras 460A, 460B, and one or more out-of-plane (e.g., upward-facing or downward-facing) localization cameras 477A, 477B (it is noted that the downward-facing cameras supplement the line following sensor 275 of the physical property sensor system 270 and may provide a wider field of view than the line following sensor 275, thereby guiding / navigating the vehicle 110 to bring the guide line 199 (see FIG. 1B) back into the field of view of the line following sensor 275 if the vehicle path deviates from the guide line 199 and causes the guide line 199 to fall out of the field of view of the line following sensor 275). Images (still images and / or dynamic video images) from the various vision system 400 cameras are requested by controller 122 from vision system controller 122VC as needed for the task of any given autonomous guided vehicle 110. For example, images are acquired by controller 122 from at least one or more of forward and rear navigation cameras 420A, 420B, 430A, 430B to provide navigation of autonomous guided vehicle 110 along transfer deck 130B and picking aisle 130A.
[0038] The forward navigation cameras 420A, 420B may be paired to form a stereo camera system, and the rear navigation cameras 430A, 430B may be paired to form another stereo camera system. With reference to Figures 2 and 3A, the forward navigation cameras 420A, 420B are any suitable cameras (such as those described above) configured to provide object detection and ranging in the manner described herein. The forward navigation cameras 420A, 420B may be positioned on opposite sides of the longitudinal centerline LAXCL of the autonomous guided vehicle 110 and may be spaced apart by any suitable distance such that the forward fields of view 420AF, 420BF provide stereo vision for the autonomous guided vehicle 110. Forward navigation cameras 420A, 420B are any suitable high-resolution or low-resolution video cameras (such as those described above, where video images including more than approximately 480 vertical scan lines and captured at greater than approximately 50 frames per second are considered high-resolution), or any other suitable cameras configured to provide object detection and ranging as described herein for navigating the autonomous vehicle along transfer deck 130B and picking aisle 130A. Rear navigation cameras 430A, 430B may be generally similar to the forward navigation cameras. Forward navigation cameras 420A, 420B and rear navigation cameras 430A, 430B provide for navigation of autonomous guided vehicle 110, including obstacle detection and avoidance (either end 200E1 of autonomous guided vehicle 110 is at the beginning or end of the direction of travel), as well as for localizing the autonomous guided vehicle within storage and retrieval system 100. Localization of the autonomous guided vehicle 110 may be provided by one or more of the forward navigation cameras 420A, 420B and rear navigation cameras 430A, 430B, by detection of guidelines on the moving / rolling surface 284 and / or by detection of suitable storage structures, including but not limited to storage rack (or other) structures.The line detection and / or storage structure detection may be compared to floor map and structure information of the vision system controller 122VC (e.g., stored in the memory of or accessible by the vision system controller 122VC). The forward navigation cameras 420A, 420B and the rear navigation cameras 430A, 430B may also send a signal to the controller 122 (including or via the vision system controller 122VC) when an object approaches the autonomous guided vehicle 110 (whether the autonomous guided vehicle 110 is stopped or in operation) so that the autonomous guided vehicle 110 can maneuver (e.g., on a non-deterministic rolling surface of the transfer deck 130B or in the picking aisle 130A (which may have a deterministic or non-deterministic rolling surface)) to avoid the approaching object (e.g., another autonomous guided vehicle, a case unit, or other transient object within the storage and retrieval system 100).
[0039] The forward navigation cameras 420A, 420B and rearward navigation cameras 430A, 430B may also provide a convoy of vehicles 110 along a picking aisle 130A or a transfer deck 130B, where one vehicle 110 follows another vehicle 110A at a predetermined fixed distance. By way of example, FIG. 1B illustrates a convoy of three vehicles 110, where one vehicle closely follows another vehicle at a predetermined fixed distance.
[0040] As another example, the controller 122 may acquire images from one or more of the three-dimensional imaging systems 440A, 440B, the case edge detection sensors 450A, 450B, and the case unit monitoring cameras 410A, 410B (the case unit monitoring cameras 410A, 410B form stereoscopic or binocular imaging cameras) to effect case handling by the vehicle 110. Still referring to Figures 2 and 3A, the one or more case edge detection sensors 450A, 450B are any suitable sensors, such as laser measurement sensors, configured to scan the shelves of the storage and retrieval system 100 to verify whether the shelves are clear for placing a case unit CU or to verify the size and position of the case unit CU before picking it. Although one case edge detection sensor 450A, 450B is illustrated on each side of the centerline CLPB (see FIG. 3A ) of the payload platform 210B, more or less than two case edge detection sensors may be positioned at any suitable location on the autonomous guided vehicle 110 so that the vehicle 110 can pass and scan the case unit CU, with the front end 200E1 leading the direction of vehicle travel or the rear end / back end 200E2 leading the direction of vehicle travel. It is noted that case handling includes picking and placing case units from a case unit holding location (such as for locating the case unit, verifying the case unit, and verifying the placement of the case unit within the payload platform 210B and / or at a case unit holding location such as a storage shelf or buffer location).
[0041] Images from out-of-plane localization cameras 477A, 477B (which may each also form a stereo image camera) may be acquired by controller 122 to effect navigation of autonomous guided vehicle 110 and / or provide data (e.g., image data) that supplements the localization / navigation data from one or more of front and rear navigation cameras 420A, 420B, 430A, 430B. Images from one or more traffic monitoring cameras 460A, 460B may be acquired by controller 122 to effect movement transition of autonomous guided vehicle 110 from picking aisle 130A to transfer deck 130B (e.g., entry onto transfer deck 130B and merging of autonomous guided vehicle 110 with other autonomous guided vehicles traveling along transfer deck 130B).
[0042] One or more out-of-plane (e.g., upward-facing or downward-facing) positioning cameras 477A, 477B (which may each also form a stereo imaging camera) are positioned on the frame 200 of the autonomous guided vehicle 110 to sense / detect position references (e.g., position marks (e.g., barcodes), lines 199 (see FIG. 1B), etc.) located on the ceiling of the storage and retrieval system or on the rolling surface 284 of the storage and retrieval system. The position references have known positions within the storage and retrieval system and may provide unique identification marks / patterns that are recognized by the vision system controller 122VC (e.g., processed data obtained from the positioning cameras 477A, 477B). Based on the detected position references, the vision system controller 122VC compares the detected position references with known position references (e.g., stored in the memory of the vision system controller 122VC or accessible to the vision system controller 122VC) to determine the position of the autonomous guided vehicle 110 within the storage structure 130.
[0043] One or more traffic monitoring cameras 460A, 460B (which may also each form a stereo image camera) are positioned on the frame 200 so that their respective fields of view 460AF, 460BF are oriented sideways in a lateral direction LAT1. While the one or more traffic monitoring cameras 460A, 460B are illustrated adjacent the transfer opening 1199 of the transfer bed 210B (e.g., on the picking side from which the arm 210A of the autonomous guided vehicle 110 extends), in other embodiments, the traffic monitoring cameras may be positioned on the non-picking side of the frame 200 so that their fields of view are oriented sideways in a direction LAT2. The traffic monitoring cameras 460A, 460B provide for autonomous merging of the autonomous guided vehicle 110, for example, as it exits the picking aisle 130A or lift transfer area 195 onto the transfer deck 130B (see FIG. 1B ). For example, autonomous guided vehicle 110V leaving lift transfer area 195 (FIG. 1B) detects autonomous guided vehicle 110T traveling along transfer deck 130B. Controller 122 then autonomously strategizes its merge onto the transfer deck (e.g., entering the transfer deck ahead of or behind autonomous guided vehicle 110T, accelerating onto the transfer deck based on the speed of the approaching vehicle 110T, etc.) based on information (e.g., distance, speed, etc.) of autonomous guided vehicle 110T collected by traffic monitoring cameras 460A, 460B and communicated to vision system controller 122VC for processing.
[0044] The case unit monitoring cameras 410A, 410B are any suitable two-dimensional rolling shutter high-resolution or low-resolution video cameras (wherein video images including greater than about 480 vertical scan lines and captured at greater than about 50 frames per second are considered high resolution), such as those described herein. The case unit monitoring cameras 410A, 410B are positioned relative to one another to form a stereo vision camera system configured to monitor the entry and exit of case units CU into and out of payload platform 210B. The case unit monitoring cameras 410A, 410B are coupled to frame 200 in any suitable manner and are focused on at least payload platform 210B. 3A , one camera 410A of the camera pair is positioned at or near one end or edge of payload platform 210B (e.g., adjacent end 200E1 of autonomous guided vehicle 110), and the other camera 410B of the camera pair is positioned at or near the other end or edge of payload platform 210B (e.g., adjacent end 200E2 of autonomous guided vehicle 110). It is noted that the distance between the cameras on either side of payload platform 210B, for example, can be such that the disparity between cameras 410A, 410B in the stereo image camera is approximately 700 pixels (in other embodiments, the disparity can be greater or less than approximately 700 pixels, although it is noted that the disparity between conventional stereo image cameras is significantly less than approximately 255 pixels, and typically approximately 96 pixels). The increased disparity between cameras 410A, 410B compared to conventional stereo image cameras can improve the resolution of disparity from pixel matching (such as when generating a depth map as described herein), where rectification of pixel matching improves the accuracy of resolution for pixels in the camera's field of view for objects located closer to the camera (near field) and objects located farther from the camera (far field) compared to conventional binocular camera systems.For example, and referring also to FIG. 4B, increasing the parallax between cameras 410A, 410B in accordance with aspects of the disclosed embodiment provides a resolution of approximately 1 mm (e.g., a parallax error of approximately 1 mm) at the front side (e.g., the side of the holding location closest to autonomous guided vehicle 110) of a case holding location (such as a storage shelf of a storage rack or other case holding location of storage and retrieval system 100) and a resolution of approximately 3 mm (e.g., a parallax error of approximately 3 mm) at the rear side (e.g., the side of the holding location further away from autonomous guided vehicle 110) of the case holding location.
[0045] The robustness of the vision system 400 allows for the determination or identification of object position and pose, taking into account the aforementioned disparity between the stereo image cameras 410A, 410B. In one or more embodiments, the case unit monitoring (stereo image) cameras 410A, 410B are coupled to the transfer arm 210A for movement in the direction LAT with the transfer arm 210A (such as when picking and placing a case unit CU) and are positioned to be focused on the payload platform 210B and support tines 210AT of the transfer arm 210A. In one or more embodiments, an off-the-shelf camera pair that is closely spaced (e.g., with a disparity of less than about 255 pixels) may be utilized.
[0046] 5A, the case unit monitoring cameras 410A, 410B at least partially provide one or more of case unit determination, case unit location identification, case unit position verification, and verification of case unit positioning features (e.g., positioning blade 471 and pusher 470) and case transfer features (e.g., tines 210AT, puller 472, and payload platform floor 473). For example, the case unit monitoring cameras 410A, 410B detect one or more of case unit lengths CL, CL1, CL2, CL3, case unit heights CH1, CH2, CH3, and case unit yaw YW (e.g., relative to the extension / retraction direction LAT of transfer arm 210A). Data from case handling sensors (e.g., described above) can also provide the location / position of the pusher 470, puller 472, and positioning blade 471, such as when payload platform 210B is empty (e.g., not holding a case unit).
[0047] Case unit monitoring cameras 410A, 410B are also configured, using vision system controller 122VC, to provide determination of a front case center point FFCP relative to a reference position of autonomous guided vehicle 110 (e.g., in the X, Y, and Z directions relative to autonomous guided vehicle 110's frame of reference BREF (see FIG. 3A) with the case unit positioned on an off-board shelf or other holding area of vehicle 110). The reference position of autonomous guided vehicle 110 may be defined by one or more positioning planes of payload platform 210B or a centerline CLPB of payload platform 210B. For example, front case center point FFCP may be determined along longitudinal axis LAX (e.g., in the Y direction) relative to centerline CLPB (FIG. 3A) of payload platform 210B. The front case center point FFCP may be determined along a vertical axis VER (e.g., in the Z direction) relative to a case unit support surface PSP of the payload platform 210B (FIGS. 3A and 3B—formed by one or more of the tines 210AT of the transfer arm 210A and the payload platform floor 473). The front case center point FFCP may be determined along a horizontal axis LAT (e.g., in the X direction) relative to a positioning surface JPP of the pusher 470 (FIG. 3B). Determining the front case center point FFCP of a case unit CU located on a storage shelf 555 (see Figures 3A and 4A) or other case unit holding location can, by way of non-limiting example, provide location identification for the autonomous guided vehicle 110 relative to the case unit CU to be picked, mapping of the case unit's location within the storage structure (e.g., in a manner similar to that described in U.S. Pat. No. 9,242,800, issued January 26, 2016, entitled "Storage and retrieval system case unit detection," the entire disclosure of which is incorporated herein by reference), and / or accuracy of picking and placing relative to other case units on the storage shelf 555 (e.g., to maintain a predetermined gap size between case units).
[0048] Determining the front case center point FFCP also results in a comparison of the “real world” environment in which the autonomous guided vehicle 110 is operating with a virtual model 400VM of that operating environment, thereby allowing the controller 122 of the autonomous guided vehicle 110 to make a substantially direct comparison of what the vision system 400 “sees” with what the autonomous guided vehicle 110 expects to “see” based on a simulation of the storage and retrieval system structure, in a manner similar to that described in U.S. patent application Ser. No. 17 / 804,026, filed May 25, 2022, entitled “Autonomous Transport Vehicle with Vision System” (having Attorney Docket No. 1127P016037-US(PAR)), the entire disclosure of which is incorporated herein by reference. 5A , the object (case unit) and its characteristics determined by the vision system controller 122VC are fitted (combined and overlaid) to the virtual model 400VM to improve resolution with up to six degrees of resolution freedom of the object's pose relative to the facility or global reference frame GREF (see FIG. 2). As can be appreciated, alignment of the cameras of the vision system 400 with the global reference frame GREF (as described herein) enables improved resolution of the pose and / or position of the vehicle 110 relative to both the global reference (the facility's features rendered in the virtual model 400VM) and the imaged object. More specifically, any discrepancies or anomalies in object position (e.g., edge spacing between the reference edges of the case unit or tilt or skew of the case unit relative to the rack slats 520L of the virtual model 400VM) revealed and identified upon fitting the object image with the virtual model 400VM exceed a predetermined nominal threshold and represent an incorrect pose of one or more of the case, rack, and / or vehicle 110. Determining whether the error is due to one or more poses / positions of the case, rack, or vehicle 110 is determined through a comparison with pose data from sensors 270 and supplemental navigation sensor system 288 .
[0049] As an example of the above-mentioned resolution enhancement, if a case unit placed on a shelf imaged by vision system 400 is rotated compared to a case unit and virtual model 400VM placed next to it on the same shelf (also imaged by the vision system), vision system 400 may determine that the case is tilted (see FIG. 4A ) and provide enhanced case position information to controller 122 to operate and position transfer arm 210A to pick the case based on the enhanced resolution of the case's pose and position. As another example, if the edge of the case is offset from the edge of slat 520L (see FIGS. 4A-4C ) beyond a predetermined threshold, it is noted that vision system 400 may generate a position error for the case, but if the offset is within the threshold, supplemental information from supplemental navigation sensor system 288 will enhance the pose / position resolution (e.g., an offset approximately equal to the determined pose / position of the case relative to the frame of slat 520L and transfer arm 210A of payload platform 210B of vehicle 110). It is further noted that while the vision system may generate a case position error if only one case is tilted / offset relative to the edge of the slat 520L, if two or more side-by-side cases are determined to be tilted relative to the edge of the slat 520L, the vision system may generate a pose error for the vehicle 110 and cause a repositioning of the vehicle 110 (e.g., correcting the position of the vehicle 110 based on an offset determined from supplemental information from the supplemental navigation sensor system 288) or send a service message to the operator (e.g., where the vision system 400 provides a "dashboard camera" collaboration mode (as described herein) that provides remote control of the vehicle 110 by the operator with images (still images and / or real-time video) from the vision system being communicated to the operator to effect remote control operation). The vehicle 110 may be stopped (e.g., not traveling down the picking aisle 130A or transfer deck 130B) until the operator initiates remote control of the vehicle 110.
[0050] Case unit monitoring cameras 410A, 410B may also provide feedback regarding the location of the case unit's positioning features and case transfer features of autonomous guided vehicle 110, for example, before and / or after picking / placing the case unit from a storage shelf or other storage location (e.g., to verify the location / position of the positioning features and case transfer features to result in picking / placing the case unit by transfer arm 210A without obstruction of the transfer arm). For example, as described above, case unit monitoring cameras 410A, 410B have a field of view that encompasses payload platform 210B. The vision system controller 122VC is configured to receive sensor data from the case unit monitoring cameras 410A, 410B and, using any suitable image recognition algorithm stored in the memory of or accessible to the vision system controller 122VC, determine the position of the pusher 470, position adjustment blade 471, puller 472, tines 210AT, and / or any other features of the payload platform 210B that engage with a case unit held on the payload platform 210B. The positions of the pusher 470, position adjustment blade 471, puller 472, tines 210AT, and / or other features of the payload platform 210B may be utilized by the controller 122 to verify the respective positions of the pusher 470, position adjustment blade 471, puller 472, tines 210AT, and / or other features of the payload platform 210B determined by the motor encoders or other respective position sensors, although in some embodiments the positions determined by the vision system controller 122VC may be utilized as a redundancy measure in the event of an encoder / position sensor failure.
[0051] The adjusted position of the case unit CU within the payload platform 210B can also be verified by the case unit monitoring cameras 410A, 410B. For example, referring also to FIG. 3C , the vision system controller 122VC is configured to receive sensor data from the case unit monitoring cameras 410A, 410B and, using any suitable image recognition algorithm stored in the memory of or accessible by the vision system controller 122VC, determine the position of the case unit in the X, Y, and Z directions, for example, relative to one or more of the centerline CLPB of the payload platform 210B, the reference / home position of the positioning surface JPP ( FIG. 3B ) of the pusher 470, and the case unit support surface PSP ( FIGS. 3A and 3B ). Here, the position determination of the case unit CU within the payload platform 210B affects at least its placement accuracy relative to other case units on the storage shelf 555 (e.g., to maintain a predetermined gap size between the case units).
[0052] 2, 3A, 3B, and 5, the one or more three-dimensional imaging systems 440A, 440B may include any suitable three-dimensional imager(s), including, but not limited to, a time-of-flight camera, an imaging radar system, or a light detection and ranging (LIDAR). The one or more three-dimensional imaging systems 440A, 440B may provide, for example, improved localization of the autonomous guided vehicle 110 relative to a global reference frame GREF (see FIG. 2) of the storage and retrieval system 100. For example, one or more three-dimensional imaging systems 440A, 440B, together with the vision system controller 122VC, may provide for determining the size (e.g., height and width) and front case center point FFCP (e.g., in the X, Y, and Z directions) of the front face (i.e., front surface) of the case unit CU relative to a reference position of the autonomous guided vehicle 110, as well as determining the invariants of the shelf supporting the case unit CU (e.g., one or more three-dimensional imaging systems 440A, 440B may provide the position of the case unit CU (which position of the case unit CU within the automated storage and retrieval system 100 is defined in the global reference frame GREF) without reference to the shelf supporting the case unit CU, and may provide a determination of whether the case unit is supported on a shelf via determining the invariant characteristics of the case unit). Here, the determination of the front surface and case center point FFCP also results in a comparison of the "real world" environment in which the autonomous guided vehicle 110 is operating with the virtual model 400VM, thereby allowing the controller 122 of the autonomous guided vehicle 110 to make a substantially direct comparison of what the vision system 400 "sees" with what the autonomous guided vehicle 110 expects to "see" based on a simulation of the storage and retrieval system structure, as described in U.S. Patent Application No. 17 / 804,026, filed May 25, 2022, entitled "Autonomous Transport Vehicle with Vision System" (having Attorney Docket No. 1127P016037-US(PAR)), the entire disclosure of which has previously been incorporated by reference herein.Image data obtained from one or more three-dimensional imaging systems 440A, 440B may supplement and / or enhance image data from cameras 410A, 410B when the data from cameras 410A, 410B is incomplete or missing. Here, object detection and localization relative to the pose of autonomous guided vehicle 110 in the global reference frame GREF may be determined with high accuracy and reliability by one or more three-dimensional imaging systems 440A, 440B, although in other aspects object detection and localization may be provided by one or more of physical property sensor system 270 and / or wheel encoders / inertial sensors of autonomous guided vehicle 110.
[0053] As illustrated in FIG. 5, one or more three-dimensional imaging systems 440A, 440B have respective fields of view extending beyond the payload platform 210B in a substantially LAT direction such that each three-dimensional imaging system 440A, 440B is positioned to detect case units CU adjacent to but outside the payload platform 210B (such as case units CU arranged in one or more rows extending along the length of the picking aisle 130A (see FIG. 5A), or substrate buffer / transfer stations arranged along the transfer deck 130B (similar in configuration to the storage rack 599 and its shelves 555 arranged along the picking aisle 130A). The fields of view 440AF, 440BF of each three-dimensional imaging system 440A, 440B encompass spatial volumes 440AV, 440BV extending to the picking range height 670 of the autonomous guided vehicle 110 (e.g., the range / height in the direction VER (Figure 2) at which the arm 210A can move to pick / place case units on shelves or stacked shelves accessible from the common rolling surface 284 on which the autonomous guided vehicle 110 rides (e.g., on the transfer deck 130B or the picking aisle 130A (see Figure 2)).
[0054] It is noted that data from one or more three-dimensional imaging systems 440A, 440B may supplement the object determination and location described herein with respect to the stereo camera pair. For example, the three-dimensional imaging systems 440A, 440B may be utilized for pose and position verification to supplement the pose and position determination made with the stereo camera pair, such as during stereo image camera calibration or autonomous guided vehicle pick and place operations. The three-dimensional imaging systems 440A, 440B may also provide reference frame transformations so that the pose and position of an object determined in the autonomous guided vehicle's reference frame BREG can be transformed to a pose and position in the global reference frame GREF, and vice versa. In other aspects, the autonomous guided vehicle may not include a three-dimensional imaging system.
[0055] The vision system 400 may also cooperate with the operator to provide operational control of the autonomous guided vehicle 110. The vision system 400 provides data (images) that is registered by the vision system controller 122VC to determine either (a) the characteristics of the information (which is in turn provided to the controller 122) or (b) the information is sent to the controller 122 without being characterized (objects within predetermined criteria) and characterized by the controller 122. In either case (a) or (b), it is the controller 122 that determines the choice to switch to a collaborative state. After switching, collaborative operation is effected by a user accessing the vision system 400 via the vision system controller 122VC and / or the controller 122 via the user interface UI. However, in its simplest form, the vision system 400 can be thought of as providing a collaborative mode of operation for the autonomous guided vehicle 110. Here, the vision system 400 complements the autonomous navigation / operation sensor system 270 to provide coordinated identification and mitigation of objects / hazards 299 (see FIG. 3A; such objects / hazards include fluids, cases, solid debris, etc.) intruding onto the driving / rolling surface 284, as described, for example, in U.S. patent application Ser. No. 17 / 804,026, filed May 25, 2022, entitled "Autonomous Transport Vehicle with Vision System" (having attorney docket number 1127P016037-US(PAR)), the entire disclosure of which was previously incorporated by reference herein.
[0056] In one aspect, an operator may select or switch control of the autonomous guided vehicle from automatic operation to collaborative operation (e.g., via a user interface UI) (e.g., the operator remotely controls the operation of the autonomous guided vehicle 110 via the user interface UI). For example, the user interface UI may include a capacitive touchpad / screen, joystick, tactile screen, or other input device that communicates kinematic direction commands (e.g., turn, accelerate, decelerate, etc.) from the user interface UI to the autonomous guided vehicle 110 to provide operator control input in a collaborative operation mode of the autonomous guided vehicle 110. For example, the vision system 400 provides a “dashboard camera” (or dash camera) that transmits video and / or still images from the autonomous transport vehicle 110 to an operator (via a user interface UI) to enable remote operation or monitoring of an area relative to the autonomous transport vehicle 110, in a manner similar to that described in U.S. Patent Application No. 17 / 804,026, filed May 25, 2022, entitled “Autonomous Transport Vehicle with Vision System” (having Attorney Docket No. 1127P016037-US(PAR)), the entire disclosure of which was previously incorporated herein by reference, providing a “dashboard camera” (or dash camera).
[0057] 1A , as described above, autonomous guided vehicle 110 is provided with vision system 400 having an architecture based on camera pairs (e.g., camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, 477A and 477B, etc.) arranged for stereo or binocular object detection and depth determination (e.g., via utilization of both disparity maps / dense depth maps from recorded video frames / images captured by each camera and keypoint data determined from recorded video frames / images captured by each camera). Object detection and depth determination results in localization (e.g., pose and position determination or identification) of the object (e.g., at least the case storage location, such as on a shelf or lift, and the case to be picked) relative to autonomous guided vehicle 110. As described herein, the vision system controller 122VC is communicatively connected to the vision system 400 to register (in any suitable memory) binocular images BIM (examples of binocular images are illustrated in Figures 6 and 7) from the vision system 400. As described herein, the vision system controller 122VC is configured to provide stereo mapping (also called disparity mapping) from the binocular images BIM to resolve a dense depth map 620 (see Figure 6) of imaged objects within the field of view. As also described herein, the vision system controller 122VC is configured to detect a stereo set of keypoints KP1-KP12 (see FIG. 7) from the binocular image BIM, each set of keypoints (see the set of keypoints in image frame 600A and the set of keypoints in image frame 600B—see FIG. 7) being separate and different from each other set of keypoints and indicating a common predetermined feature of each imaged object (e.g., a corner, an edge, a portion of text, a portion of a barcode, etc.), whereby the vision system controller 122VC derives a depth analysis result for each object from the stereo set of keypoints KP1-KP12 that is separate and different from the dense depth map 620.
[0058] 3A, 3B, 6 and 7, as mentioned above, the vision system controller 122VC is part of or communicatively connected to the controller 122 and registers image data from the camera pairs. The camera pair 410A and 410B positioned to view at least the payload platform 210B is referenced for illustrative purposes, but it should be understood that the other camera pairs 420A and 420B, 430A and 430B, 460A and 460B, 477A and 477B also provide detection (or identification) of the pose and position of the object in a generally similar manner. Here, the controller 122VC registers binocular images from the cameras 410A, 410B (e.g., in the form of video stream data) and analyzes the video stream data into pairs of stereo image frames (e.g., still images), although it is noted that the cameras 410A, 410B are again not synchronized with one another (note that camera synchronization occurs when the cameras are configured to capture corresponding image frames (still or video) simultaneously with respect to one another, i.e., when the camera shutters for each camera in the pair are synchronously activated and deactivated at the same time). The vision system controller 122VC is configured to process the image data from each camera 410A, 410B such that image frames 600A, 600B are analyzed from the respective video stream data as a time-matched stereo image pair 610.
[0059] 3A, 3B, and 6, the vision system controller 122VC is configured with an object extractor 1000 (FIG. 10) that includes a dense depth estimator 666. The dense depth estimator 666 configures the vision system controller 122VC to generate a depth map 620 from the stereo image pair 610, where the depth map 620 embodies objects in the fields of view of the cameras 410A, 410B. The depth map 620 may be a dense depth map generated in any suitable manner, such as from a point cloud obtained by disparity mapping the pixels / image points of each image 600A, 600B in the stereo image pair 610. The image points in the images 600A, 600B may be obtained and matched (e.g., pixel matching) in any suitable manner and using any suitable algorithm that is stored and executed by the controller 122VC (or controller 122). Exemplary algorithms include, but are not limited to, RAFT-Stereo, HITNet, AnyNet®, StereoNet, StereoDNN (also known as Stereo Depth DNN), Semi-Global Matching (SGM), or any other suitable method, such as employing any of the stereo matching methods listed in Appendix A, all of which are incorporated herein by reference in their entirety, one or more of which may be deep learning methods (e.g., involving training with a suitable model) or learning-free approaches (e.g., no training). Here, the cameras 410A, 410B in the stereo image camera are calibrated with respect to each other (by any suitable method, such as those described herein), so that the epipolar geometry representing the relationship between the stereo image pair captured by the cameras 410A, 410B is known, and the images 600A, 600B of the image pair are aligned (rectified) with respect to each other, resulting in the generation of a depth map. As described above, the dense depth map is generated using an unsynchronized camera pair 410A and 410B (eg, images 600A, 600B are close in time but unsynchronized).Thus, if one image 600A, 600B in an image pair is blocked, blurred, or otherwise unusable, the vision system controller 122VC obtains another image pair of the analyzed binocular image pair from the registered image data (i.e., one subsequent to the image pair 600A, 600B), and if such subsequent image is unblocked (e.g., the loop continues until an unblocked analyzed image pair is obtained), a dense depth map 620 is generated (it is noted that keypoints as described herein are determined from the same image pair used to generate the depth map).
[0060] The dense depth map 620 is “dense” (e.g., has depth resolution for all or nearly all pixels in the image) compared to a sparse depth map (e.g., stereo-matched keypoints) and has a similar level of detail for identifying objects within the camera's field of view that provides resolution for the autonomous guided vehicle 110's pick-and-place operations. Here, the density of the dense depth map 620 may depend on (or be defined by) the processing power and processing time available for object identification. By way of example, as described above, the transfer of an object (such as a case unit CU) to / from the payload platform 210B of the autonomous guided vehicle 110 is performed within approximately 10 seconds from the time the bot stops moving to the time the bot starts moving. For the transfer of an object, movement of transfer arm 210A is initiated before autonomous guided vehicle 110 stops traveling, the autonomous guided vehicle is positioned adjacent to the pick / place location where the object is to be transferred (e.g., the position and pose of a holding station, the position and pose of an object / case unit, etc.), and transfer arm 210A is extended substantially simultaneously with the stopping of the autonomous guided vehicle. Here, at least a portion of the images captured by vision system 400 (e.g., to identify the object to be picked, the case holding location, or other objects in storage and retrieval system 100) are captured while the autonomous guided vehicle is traveling on the traveling surface (i.e., while autonomous guided vehicle 110 is traveling along transfer deck 130B or picking aisle 130A and moving past the object). Identification of the object occurs approximately simultaneously with stopping of the autonomous guided vehicle (e.g., at least partially while the autonomous guided vehicle 110 is operating and decelerating from a driving speed to a stop) so that generation of a high-density depth map for object identification is completed (e.g., in less than about 2 seconds, or less than about 0.5 seconds) approximately coinciding with stopping of the autonomous guided vehicle and the start of movement of the transfer arm 210A.The analysis of the dense depth map 620 renders (notifies) the vision system controller 122VC (and controller 122) of object anomalies (see open case flaps and tape on the case, or other anomalies including but not limited to tears on the case front, appliqués (such as tape or other adhered overlays) on the case front) due to the object's surface, etc., in association with autonomous guided vehicle commands (e.g., pick / place commands, etc.). The analysis of the dense depth map 620 also results in stock keeping unit (SKU) identification, where the vision system controller 122VC determines the front dimensions of the case and determines the SKU based on the front dimensions (e.g., SKUs are stored in a table along with their respective front dimensions, whereby the SKUs are correlated to their respective front dimensions, and the vision system controller 122VC or controller 122 compares the determined front dimensions to the front dimensions in the table to identify which SKUs are correlated to the determined front dimensions).
[0061] 3A, 3B, 6, 7, 8, and 9, the vision system controller 122VC is configured with an object extractor 1000 (see FIG. 10) that includes a binocular case keypoint detector 999. The binocular case keypoint detector 999 configures the vision system controller 122VC to detect a stereo set of keypoints from the binocular images 600A, 600B (see FIG. 7, where exemplary keypoints KP1, KP2 form one keypoint set, keypoints KP3-KP7 form another keypoint set, and keypoints KP8-KP12 form yet another keypoint set, although it is noted that keypoints may also be referred to as "feature points," "invariant features," "invariant points," or "features" (such as corners or facet joints or object surfaces)). Each set of keypoints is distinct from and indicates common, predetermined features of each image object (here, CU1, CU2, CU3) and is separate and distinct from each other set of keypoints, thereby allowing vision system controller 122VC to derive a depth analysis result for each object CU1, CU2, CU3 from the stereo set of keypoints, which is distinct and distinct from dense depth map 620. A keypoint detection algorithm may be located within the residual network backbone (see FIG. 8) of vision system 400, where a feature pyramid network for feature / object detection (see FIG. 8) is used to predict or otherwise resolve keypoints for each image 600A, 600B separately. Keypoints for each image 600A, 600B may be determined in any suitable manner using any suitable algorithm stored in controller 122VC (or controller 122), including, but not limited to, Harris Corner Detector, Microsoft COCO (Common Objects in Context), other deep learning and logistics models, or other corner detection methods. It is noted that suitable examples of corner detection methods may be informed by deep learning methods or may be corner detection approaches that do not use deep learning.As described above, at least some of the images captured by vision system 400 (e.g., to identify an object to be picked, a case holding location, or other object in storage and retrieval system 100) are captured while the autonomous guided vehicle is traveling on the travel surface (i.e., while autonomous guided vehicle 110 is moving along transfer deck 130B or picking aisle 130A and moving past the object). Identification of the object occurs substantially simultaneously with stopping of the autonomous guided vehicle (e.g., at least partially while autonomous guided vehicle 110 is operating and decelerating from a travelling speed to a stop) such that keypoint detection for object identification is completed (e.g., in less than about 2 seconds, or less than about 0.5 seconds) substantially coinciding with the autonomous guided vehicle stopping travel and transfer arm 210A starting movement.
[0062] As can be seen from FIGS. 8 and 9, keypoint detection is provided separately and differently from the dense depth map 620. For each camera image 600A, 600B, keypoints are detected within the image frame separately and differently from each other camera image 600A, 600B to form a stereo pair or set of keypoints from the stereo images 600A, 600B. While FIG. 8 illustrates an example flow diagram of keypoint determination for image 600B, it is noted that such keypoint determination is generally similar for image 600A. In keypoint detection, a residual network backbone and a feature pyramid network provide predictions (FIG. 8, block 800) for region proposals (FIG. 8, block 805) and regions of interest (FIG. 8, block 810). Bounding boxes are provided for objects in image 600B (FIG. 8, block 815), and suspect cases are identified (FIG. 8, block 820). Non-maximum suppression (NMS) is applied to the bounding boxes (and the suspect cases or portions thereof identified in the bounding boxes) (FIG. 8, block 825) to filter the results, where such filtered results and regions of interest are input into a keypoint logit mask (e.g., using deep learning, or in other aspects, such as in the exemplary methods described herein without deep learning) (FIG. 8, block 830) to determine keypoints (FIG. 8, block 835).
[0063] 9 illustrates an exemplary keypoint determination flow diagram for keypoint determination in both images 600A and 600B (keypoints for image 600A are determined separately from keypoints for image 600B, where keypoints in each image are determined in a manner generally similar to that described above with respect to FIG. 8). Keypoint determination for images 600A and 600B may be performed in parallel or sequentially (e.g., blocks 800, 900, 805-830 may be performed in parallel or sequentially), such that the output of each keypoint logit mask 830 is utilized by vision system controller 122VC as input to a matched stereo logit mask (FIG. 9, block 910) for determination of stereo (three-dimensional) keypoints (FIG. 9, block 920). High-resolution regions of interest (FIG. 9, block 905) may be determined / predicted by a residual network backbone and a feature pyramid network based on the respective regions of interest (block 810), where the high-resolution regions of interest (block 905) are input into the respective keypoint logic mask 830. The vision system controller 122VC generates matched stereo regions of interest (FIG. 9, block 907) based on the regions of interest (block 905) for each image 600A, 600B, where the matched stereo regions of interest (block 907) are input into the matched stereo logit mask (block 910) for determination of stereo (three-dimensional) keypoints (block 920). The high-resolution regions of interest from each image 600A, 600B may be matched by pixel matching or any other suitable method to generate the matched stereo regions of interest (block 907). Another non-maximum suppression (NMS) is applied (FIG. 9, block 925) to filter the keypoints to obtain a final set of stereo-matched keypoints 920F, such as the stereo-matched keypoints KP1-KP12 illustrated in FIG. 7 (also referred to herein as a stereo set of keypoints).
[0064] 4A, 6, and 7, the stereo-matched keypoints KP1-KP12 are matched to generate a best fit (e.g., a depth discrimination for each keypoint), where the stereo-matched keypoints KP1-KP12 resolve the depth of each stereo-matched keypoint KP1-KP12 relative to a predetermined reference frame (e.g., the autonomous guided vehicle's reference frame BREF (see FIG. 3A) and / or the automated storage and retrieval system 100's global reference frame GREF (see FIG. 4A), where the autonomous guided vehicle's reference frame BREF is associated with the global reference frame GREF (i.e., a transformation is determined as described herein) such that the pose and position of objects detected by the autonomous guided vehicle 110 are known in both the global reference frame GREF and the autonomous guided vehicle's reference frame BREF). As described herein, the solved stereo matched keypoints KPI to KPI2 are separate and distinct from the dense depth map 620 and provide a solution separate and distinct from the solution provided by the dense depth map 620 to determine the pose and depth / position of an object (such as case CU), although both solutions are provided from a common set of stereo images 600A, 600B.
[0065] 6, 7, and 10, the vision system controller 122VC has an object extractor 1000 configured to determine the position and pose of each imaged object (such as a case CU or other object of the automated storage and retrieval system 100 located within the field of view of the cameras 410A, 410B) from both the dense depth map 620 resolved from the binocular images 600A, 600B and the depth analysis results from the matched stereo keypoints 620F. For example, the vision system controller 122VC is configured to combine the dense depth map 620 (from the dense depth estimator 666) and the matched stereo keypoints 920F (from the binocular case keypoint detector 999) in any suitable manner. The depth information from the matched stereo keypoints 920F is combined with depth information from the dense depth map 620 for one or more objects in the images 600A, 600B, such as the case CU2, to determine an initial estimate of a point within the case plane CF (FIG. 10, block 1010). An outlier detection loop ( FIG. 10 , block 1015) is performed on the initial estimates of the points in the case surface CF to generate a valid plane for the case surface ( FIG. 10 , block 1020). The outlier detection loop may be any suitable outlier algorithm (e.g., RANSAC or any other suitable outlier / inlier detection method, etc.) that identifies points in the initial estimates of the points in the case surface as inliers and outliers, where inliers are within a predetermined best-fit threshold and outliers are outside the predetermined best-fit threshold. The valid plane for the case surface may be defined by the valid plane for the case surface containing a best-fit threshold of approximately 75% of the points in the initial estimates of the points in the case surface (in other embodiments, the best-fit threshold may be higher or lower than approximately 75%).Any suitable statistical test (similar to, but using less stringent criteria than, the outlier detection loop described above) is performed on the valid plane of the case face (again, the best-fit point based on a subsequent predetermined best-fit threshold) ( FIG. 10 , block 1025), and approximately 95% (in other embodiments, the subsequent best-fit threshold may be higher or lower than approximately 95%) of the points (some of which may be outliers in the outlier detection loop) are included in the points on the case face to define a final estimate (e.g., best fit) of the points ( FIG. 10 , block 1030). The remaining points (greater than approximately 95%) may also be analyzed so that points at a predetermined distance from the determined case face CF are included in the final estimate of the points on the face. For example, the predetermined distance may be approximately 2 cm (in other embodiments, the predetermined distance may be greater or less than approximately 2 cm), so that points corresponding to open flaps or other case deformations / anomalies are included in the final estimate of the face, notifying the vision system controller 122VC that an open flap or other case deformation / anomaly is present.
[0066] The final value of the points in the case plane (best fit) (block 1030) may be verified (e.g., with a weighted verification where weights are assigned to the matched stereo keypoints 920F (see also keypoints KP1-KP12, which are examples of matched stereo keypoints 920F)). For example, the object extractor 1000 is configured to identify the position and pose of each imaged object (e.g., relative to a predetermined reference frame, such as the global reference frame GREF and / or the autonomous guided vehicle reference frame BREF) based on the registration of the (set of) matched stereo keypoints (and their depth analysis results) with the depth map 620. Here, the matched or matching stereo keypoints KP1-KP12 are superimposed with the final estimate of the point in the case plane (block 1030) (e.g., the point cloud forming the final estimate of the point in the case plane is projected onto the plane formed by the matching stereo keypoints KP1-KP12) and resolved for comparison with the point in the case plane to determine whether the final estimate of the point in the case plane is within a predetermined threshold distance from the matching stereo keypoints KP1-KP12 (and the case plane formed thereby). If the final estimate of the point in the plane is within the predetermined threshold distance, the final estimate of the point in the plane (defining the determined case plane CF) is verified to form a plane estimate of the matching stereo keypoints ( FIG. 10 , block 1040). If the final estimate of the point in the plane is outside the predetermined threshold, the final estimate of the point in the plane is discarded or readjusted (e.g., by lowering the best fit threshold described above or in any other suitable manner). In this way, the determined pose and position of the case face CF is weighted to the matching stereo keypoints KP1-KP12.
[0067] 4A and 10, the vision system controller 122VC is configured to determine a front face and a front face dimension of at least one extracted object based on the planar estimation of the matching stereo keypoints (block 1040). For example, FIG. 4 illustrates planar estimations of matching stereo keypoints (block 1040) for various extracted objects (e.g., cases CU1, CU2, CU3, storage shelf hut 444, support slat 520L, storage shelf 555, etc.). Referring to case CU2 as an example, the vision system controller 122VC determines a case face CF and case face CF dimensions CL2, CH2 of case CU2. As described herein, the determined dimensions CL2, CH2 of case CU2 are stored in a table such that the vision system controller 122VC is configured to determine the logistics identification information (e.g., stockkeeping unit) of the extracted object (e.g., case CU2) based on the front face or case face CF dimensions CL2, CH2 in a manner similar to that described herein.
[0068] From the matching stereo keypoint planar estimates (block 1040), the vision system controller 122VC may also determine the front case center point FFCP and other dimensions / features (e.g., the spatial envelope ENV between the hats 444, the case support surface, the distance DIST between the cases, the case tilt, the case deformations / anomalies, etc.) that affect the case transfer between the storage shelf 555 and the autonomous guided vehicle 110, as described herein. For example, the vision system controller 122VC is configured to characterize the front planar surface PS (of the extracted object) and the orientation of the planar surface PS with respect to a predetermined reference frame (such as the autonomous guided vehicle reference frame BREF and / or the global reference frame GREF). Referring again to case CU2 as an example, the vision system controller 122VC characterizes a planar surface PS of the case face CF of case CU2 from the plane estimates of the matching stereo keypoints (block 1040) and determines an orientation (e.g., tilt or yaw YW—see also FIG. 3A) of the planar surface PS relative to one or more of the global reference frame GREF and the autonomous guided vehicle's reference frame BREF. The vision system controller 122VC is configured to characterize a picking surface BE (e.g., a lower edge that defines the position of the picking surface of the case unit CU to be picked—see FIG. 4A) of the extracted object (such as case unit CU) based on the features of the planar surface PS from the plane estimates of the matching stereo keypoints (block 1040), where the picking surface BE interfaces with the payload handler or transfer arm 210A (see FIGS. 2 and 3A) of the autonomous guided vehicle 110.
[0069] As noted above, determining the plane estimate of the matching stereo keypoints (block 1040) includes points located a predetermined distance in front of the plane / surface formed by the matched stereo keypoints KP1-KP12. Here, the vision system controller 122VC is configured to ascertain the presence and characteristics of anomalies relative to the planar surface PS (e.g., tape on the case face CF (see FIG. 6), an open case flap (see FIG. 6), a tear in the case face, etc.) from the plane estimate of the matching stereo keypoints.
[0070] Vision system controller 122VC is configured to generate at least one of a run command and a stop command for an actuator of autonomous guided vehicle 110 (e.g., an actuator of transfer arm 210A, an actuator of drive wheel 260, or any other suitable actuator of autonomous guided vehicle 110) based on the identified position and pose of the case CU to be picked. For example, if the pose and position of the case identify that the case CU to be picked extends beyond shelf 555 and the case cannot be picked substantially without interference or obstruction (e.g., substantially without error), vision system controller 122VC may generate a stop command that prevents extension of transfer arm 210A. As another example, if the pose and position of the case determine that the case CU to be picked is tilted and not aligned with the transfer arm 210A, the vision system controller 122VC may generate an execution command that causes the autonomous guided vehicle to travel along the travel surface to position the transfer arm 210A relative to the case CU to be picked so that the tilted case is aligned with the transfer arm 210A and can be picked without error.
[0071] It is noted that as the flip side of, or in conjunction with, robustly resolving the pose and position of the case CU relative to either or both of the autonomous guided vehicle reference frame BREF and the global reference frame GREF, resolving the reference frame BREF (e.g., pose and position) of the autonomous guided vehicle 110 relative to the global reference frame GREF is possible and can be resolved using three-dimensional imaging systems 440A, 440B (see FIG. 3A ). For example, three-dimensional imaging systems 440A, 440B can be utilized to detect a global reference datum (e.g., a portion of a storage and retrieval system structure having a known position, such as a calibration station, a case transfer station, etc., as described herein), whereby vision system controller 122VC determines the pose and position of the autonomous guided vehicle 110 relative to the global reference datum. Determining the pose and position of the autonomous guided vehicle 110 and the pose and position of the case CU informs controller 122 whether the pick / place operation can be performed substantially without interference or obstruction.
[0072] 1A, 2, 3A, 3B, and 11, stereo pairs of cameras 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, and 477A and 477B are calibrated to acquire video stream data imaging using vision system 400. The stereo pairs of cameras 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, and 477A and 477B may be calibrated in any suitable manner (e.g., by intrinsic and extrinsic camera calibration, etc.) to provide detection of case units CU, storage structures (shelves, columns, etc.), and other structural features of the storage and retrieval system. Calibration of the stereo pairs of cameras may be performed at calibration station 1110 of storage structure 130. As seen in FIG. 11 , the calibration station 1110 may be located at or adjacent to an autonomous guided vehicle entry or exit location 1190 of the storage structure 130. The autonomous guided vehicle entry or exit location 1190 allows for the introduction and exit of the autonomous guided vehicle 110 into and from one or more storage levels 130L of the storage structure 130 in a manner generally similar to that described in U.S. Patent No. 9,656,803, issued May 23, 2017, entitled “Storage and Retrieval System Rover Interface,” the entire disclosure of which is incorporated herein by reference. For example, the autonomous guided vehicle entry or exit location 1190 includes a lift module 1191 such that the autonomous guided vehicle 110 can enter and exit each storage level 130L of the storage structure 130. The lift module 1191 may interface with the transfer deck 130B of one or more storage levels 130L. The interface between the lift module 1191 and the transfer deck 130B may be located at a predetermined location on the transfer deck 130B such that the entry and exit of the autonomous guided vehicles 110 to and from each transfer deck 130B is substantially decoupled from the throughput of the automated storage and retrieval system 100 (e.g., the entry and exit of the autonomous guided vehicles 110 to and from each transfer deck does not affect the throughput).In one aspect, the lift module 1191 may interface with a spur or staging area 130B1-130Bn (e.g., a load platform for an autonomous guided vehicle) connected to or forming part of the transfer deck 130B of each storage level 130L. In other aspects, the lift module 1191 may interface substantially directly with the transfer deck 130B. It is noted that the transfer deck 130B and / or the staging area 130B1-130Bn may include any suitable barrier 1120 that substantially prevents the autonomous guided vehicle 110 from moving away from the transfer deck 130B and / or the staging area 130B1-130Bn at the lift module interface. In one aspect, the barrier may be a movable barrier 1120 that may be movable between a deployed position to substantially prevent the autonomous guided vehicle 110 from moving away from the transfer deck 130B and / or the staging areas 130B1-130Bn and a stowed position to allow the autonomous guided vehicle 110 to transition between the lift platform 1192 of the lift module 1191 and the transfer deck 130B and / or the staging areas 130B1-130Bn. In addition to transporting the autonomous guided vehicle 110 into or out of the storage structure 130, in one aspect, the lift module 1191 may also transport the rover 110 between storage levels 130L without removing the autonomous guided vehicle 110 from the storage structure 130.
[0073] Each staging area 130B1-130Bn includes a respective calibration station 1110 positioned so that the autonomous guided vehicle 110 may repeatedly calibrate the stereo camera pairs 410A and 410B, 420A and 420B, 430A and 430B, 460A and 460B, and 477A and 477B. Calibration of the stereo camera pairs may occur automatically upon introduction of the autonomous guided vehicle into the storage structure 130 (via an autonomous guided vehicle entry or exit position 1190 in a manner similar to that described in U.S. Pat. No. 9,656,803, previously incorporated by reference). In other aspects, calibration of the stereo camera pairs may occur manually (such as when the calibration station is located on a lift 1192) or may be performed prior to insertion of the autonomous guided vehicle 110 into the storage structure 130 in a manner similar to that described herein with respect to the calibration station 1110.
[0074] To calibrate the stereo camera pair, the autonomous guided vehicle is positioned (manually or automatically) at a predetermined location in the calibration station 1110 ( FIG. 14 , block 1400). Automatic positioning of the autonomous guided vehicle 110 at the predetermined location may utilize detection of any suitable features of the calibration station 1110 using the vision system 400 of the autonomous guided vehicle 110. For example, the calibration station 1110 includes any suitable location flags or positions 1110S disposed on one or more surfaces 1200 of the calibration station 1110. The location flags 1110S are disposed on one or more surfaces within the field of view of at least one camera 410A, 410B of each camera pair. The vision system controller 122VC is configured to detect the location flags 1110S, and detection of one or more of the location flags 1110S causes the autonomous guided vehicle to be roughly positioned relative to a calibration or known object in the calibration station 1110. In other aspects, in addition to or instead of the position flag 1110S, the calibration station 1110 may include a buffer or physical stop against which the autonomous guided vehicle 110 abuts to position itself at a predetermined location in the calibration station 1110. The buffer or physical stop may be, for example, a barrier 1120 or other suitable fixed or deployable feature of the calibration station. Automatic positioning of the autonomous guided vehicle 110 at the calibration station 1110 may occur upon introduction of the autonomous guided vehicle 110 into the storage and retrieval system 100 (e.g., the autonomous guided vehicle exits the lift 1192) and / or at any appropriate time when the autonomous guided vehicle enters the calibration station 1110 from the transfer deck 130. Here, the autonomous guided vehicle 110 may be programmed with calibration instructions that effect stereo vision calibration upon introduction into the storage structure 130, or the calibration instructions may be initialized at any appropriate time while the autonomous guided vehicle 110 is operating (i.e., in service) within the storage structure 130.
[0075] The one or more surfaces 1200 of each calibration station 1110 include any suitable number of known objects 1210-1218. The one or more surfaces 1200 may be any surface visible by the stereo pair of cameras, including, but not limited to, the sidewalls 1111 of the calibration station 1110, the ceiling 1112 of the calibration station 1110, the floor / running surface 1115 of the calibration station 1110, and the barrier 1120 of the calibration station 1110. The objects 1210-1218 (also referred to as visual datums or calibration objects) included with each surface 1200 may be raised structures, apertures, appliqués (e.g., paint, stickers, etc.), each having known physical characteristics such as shape, size, etc.
[0076] While calibration of case unit monitoring (stereo imaging) cameras 410A, 410B using calibration station 1110 is described for illustrative purposes, it should be understood that other stereo imaging cameras may be calibrated in a substantially similar manner. With autonomous guided vehicle 110 remaining continuously stationary at a predetermined position at calibration station 1110 (where objects 1210-1218 are within the field of view of cameras 410A, 410B) throughout the calibration process, each camera 410A, 410B of the stereo imaging cameras images objects 1210-1218 (FIG. 14, block 1405). These images of the object are registered by the vision system controller 122VC (or controller 122), which is configured to calibrate the stereo vision of the stereo image cameras by determining the epipolar geometry of the camera pair (FIG. 14, block 1410) in any suitable manner (such as the methods described in Wheeled Mobile Robotics from Fundamentals Towards Autonomous Systems, 1st Ed., 2017, ISBN 9780128042045, the entire disclosure of which is incorporated herein by reference). The vision system controller 122VC is also configured to calibrate the disparity between the cameras 410A, 410B in the stereo camera using the objects 1210-1218, where the disparity between the cameras 410A, 410B is determined by matching pixels from an image captured by the camera 410A with pixels in a corresponding image captured by the camera 410B (FIG. 14, block 1415), and the distance for each pair of matching pixels is calculated. The disparity and epipolar geometry calibration may be further refined in any suitable manner, for example, using data obtained from images of the objects 1210-1218 captured by the three-dimensional imaging systems 440A, 440B of the autonomous guided vehicle 110 (FIG. 14, block 1420).
[0077] Further, the binocular vision reference frame may be transformed or otherwise resolved into a predetermined reference frame, such as the reference frame BREF of the autonomous guided vehicle 110 and / or the global reference frame GREF, using three-dimensional imaging systems 440A, 440B (FIG. 14, block 1425), where a portion of the autonomous guided vehicle 110 (such as frame 200 having known dimensions or a portion of transfer arm 210A at a known pose relative to frame 200) is imaged relative to a known global reference frame datum (e.g., a global datum target GDT located at calibration station 1110, which in some embodiments may be the same as objects 1210-1218). 13 , a computer model 1300 (such as a computer-aided drafting model or CAD model) of autonomous guided vehicle 110 and / or a computer model 400VM of the operating environment of storage structure 130 (see also FIG. 1A ) may be utilized by vision system controller 122VC to transform the binocular vision reference frame of each camera pair into the reference frame BREF of autonomous guided vehicle 110 or the global reference frame GREF of storage structure 130. As seen in FIG. 13 , feature dimensions, such as any appropriate feature of payload platform 210B depending on which camera pair is calibrated (which in this example is a feature of the payload platform fence relative to the reference frame BREF, or any other appropriate feature of autonomous guided vehicle 110 via autonomous guided vehicle model 1300, and / or an appropriate feature of the storage structure via the virtual model of the operating environment 400VM), may be extracted by vision system controller 122VC for portions of autonomous guided vehicle 110 that are within the field of view of the camera pairs. These characteristic dimensions of payload platform 210B are determined from the origin of the reference frame BREF of autonomous guided vehicle 110. These known dimensions of autonomous guided vehicle 110, along with the image pairs or depth maps produced by the stereo imaging cameras, are utilized by vision system controller 122VC to correlate each camera's reference frame (or camera pair's reference frame) to the reference frame BREF of autonomous guided vehicle 110.Similarly, the characteristic dimensions of the global datum target GDT are determined from the origin (e.g., the origin of the global reference frame GREF) of the storage structure 130. These known dimensions of the global datum target GDT, along with the image pairs or depth maps produced by the stereo imaging cameras, are utilized by the vision system controller 122VC to correlate each camera's reference frame (stereoscopic reference frame) to the global reference frame GREF.
[0078] If stereoscopic calibration of the autonomous guided vehicle 110 is performed manually, the autonomous guided vehicle 110 is manually positioned at a calibration station. For example, the autonomous guided vehicle 110 is manually positioned on a lift 1191 that includes surface(s) 1111 (one of which is shown, but other surfaces may be located at the end of the lift platform or may be located on the lift platform in an orientation similar to the surface of the calibration station 1110 (e.g., the lift platform is configured as the calibration station)). The surface(s) include known objects 1210-1218 and / or global datum targets GDTs such that stereoscopic calibration is performed in a manner generally similar to that described above.
[0079] 1A, 2, 3A, 3B, 6, and 7, an exemplary method for determining the pose and position of an imaged object is described in accordance with aspects of the disclosed embodiment. In the method, an autonomous guided vehicle 110 as described herein is provided (FIG. 15, block 1500). The vision system 400 generates binocular images 600A, 600B (FIG. 15, block 1505) of a field of logistics space (defined by the combined fields of view of cameras in a camera pair, such as cameras 410A and 410B—see FIGS. 6 and 7) (e.g., formed by storage structures 130) including rack structure shelves 555 on which a plurality of objects (e.g., case units CU) are stored. A controller (e.g., vision system controller 122VC or controller 122) communicatively coupled to vision system 400 registers binocular images 600A, 600B (e.g., in any suitable memory of the controller) (FIG. 15, block 1510) and performs stereo matching from the binocular images to resolve a dense depth map 620 of imaged objects in the field (FIG. 15, block 1515). The controller detects a stereo set of keypoints KP1-KP12 from the binocular images (FIG. 15, block 1520), each set of keypoints (each image 600A, 600B having a set of keypoints) distinct and different from each other set, indicating common predetermined characteristics of each imaged object, whereby the controller derives a depth resolution for each object from the stereo set of keypoints KP1-KP12, distinct and different from the dense depth map 620 (FIG. 15, block 1525). The controller uses the controller's object extractor 1000 to determine or identify the position and pose of each imaged object from both the dense depth map 620 solved from the binocular images 600A, 600B and the depth analysis results from the stereo set of keypoints KP1-KP12 (FIG. 15, blocks 1530 and 1535).
[0080] In accordance with one or more aspects of the disclosed embodiments, there is provided an autonomous guided vehicle comprising: a frame having a payload holding portion; a drive section coupled to the frame having drive wheels for supporting the autonomous guided vehicle on a driving surface, the drive wheels effecting vehicle movement on the driving surface and moving the autonomous guided vehicle over the driving surface at a facility; a payload handler coupled to the frame, the payload handler configured to transport a payload having a flat, non-deterministic landing surface seated on the payload holding portion to / from the payload holding portion of the autonomous guided vehicle and a storage location for the payload within a storage array; and a vision system mounted on the frame, the vision system having a plurality of cameras positioned to generate a binocular image of a field of a logistics space including shelves in a rack structure on which a plurality of objects are stored, the vision system registering the binocular image. and a controller communicatively connected to the vision system to perform stereo matching from the binocular images and configured to effect stereo matching from the binocular images to resolve a dense depth map of imaged objects in the field, the controller configured to detect stereo sets of keypoints from the binocular images, each set of keypoints being distinct and different from each other set and indicating a common predetermined characteristic of each imaged object, whereby the controller derives a depth analysis result for each object from the stereo set of keypoints that is distinct and different from the dense depth map, the controller having an object extractor configured to determine a position and pose of each imaged object from both the dense depth map resolved from the binocular images and the depth analysis result from the stereo set of keypoints.
[0081] In accordance with one or more aspects of the disclosed embodiment the plurality of cameras are rolling shutter cameras.
[0082] In accordance with one or more aspects of the disclosed embodiment, multiple cameras generate video streams and registered images are analyzed from the video streams.
[0083] In accordance with one or more aspects of the disclosed embodiment, the cameras are not synchronized with one another.
[0084] In accordance with one or more aspects of the disclosed embodiment, the binocular images are generated during operation of the vehicle passing the object.
[0085] In accordance with one or more aspects of the disclosed embodiment, multiple objects on a rack structure are dynamically positioned in close packed juxtaposition relative to one another.
[0086] In accordance with one or more aspects of the disclosed embodiment the controller is configured to determine a front surface and a dimension of the front surface of the at least one extracted object.
[0087] In accordance with one or more aspects of the disclosed embodiment the controller is configured to characterize a planar surface of the front surface and an orientation of the planar surface relative to a predetermined frame of reference.
[0088] In accordance with one or more aspects of the disclosed embodiment, the controller is configured to characterize a picking surface of the extracted object based on characteristics of a planar surface that interfaces with the payload handler.
[0089] In accordance with one or more aspects of the disclosed embodiment the controller is configured to ascertain a presence and characteristics of an anomaly relative to the planar surface.
[0090] In accordance with one or more aspects of the disclosed embodiment the controller is configured to determine a logistics identity of the extracted object based on a dimension of the front surface.
[0091] In accordance with one or more aspects of the disclosed embodiment, the controller is configured to generate at least one of a run command and a stop command for the bot actuator based on the determined position and pose.
[0092] In accordance with one or more aspects of the disclosed embodiments, there is provided an autonomous guided vehicle comprising: a frame having a payload holding portion; a drive section coupled to the frame having drive wheels supporting the autonomous guided vehicle on a driving surface, the drive wheels effecting vehicle movement on the driving surface and moving the autonomous guided vehicle over the driving surface at a facility; a payload handler coupled to the frame, the payload handler configured to transport a payload having a flat, non-deterministic landing surface seated on the payload holding portion to / from the payload holding portion of the autonomous guided vehicle and a storage location for the payload within a storage array; and a vision system mounted on the frame, the vision system having a binocular imaging camera for generating a binocular image of a field of a logistics space including shelves in a rack structure on which a plurality of objects are stored, the vision system registering the binocular image. a controller communicatively connected to the vision system such that the stereo matching is performed and configured to perform stereo matching from the binocular images to elucidate a dense depth map of imaged objects in the field, the controller being configured to detect stereo sets of keypoints from the binocular images, each set of keypoints being distinct and different from each other set of keypoints and indicating common predetermined characteristics of each imaged object, whereby the controller derives a depth analysis result for each object from the stereo set of keypoints, the depth analysis result being distinct and different from the dense depth map; and the controller having an object extractor configured to identify a position and pose of each imaged object based on a superposition of the depth analysis result and the depth map of the stereo set of keypoints.
[0093] In accordance with one or more aspects of the disclosed embodiment the plurality of cameras are rolling shutter cameras.
[0094] In accordance with one or more aspects of the disclosed embodiment, multiple cameras generate video streams and registered images are analyzed from the video streams.
[0095] In accordance with one or more aspects of the disclosed embodiment, the cameras are not synchronized with one another.
[0096] In accordance with one or more aspects of the disclosed embodiment, the binocular images are generated during operation of the vehicle passing the object.
[0097] In accordance with one or more aspects of the disclosed embodiment, multiple objects on a rack structure are dynamically positioned in close packed juxtaposition relative to one another.
[0098] In accordance with one or more aspects of the disclosed embodiment the controller is configured to determine a front surface and a dimension of the front surface of the at least one extracted object.
[0099] In accordance with one or more aspects of the disclosed embodiment the controller is configured to characterize a planar surface of the front surface and an orientation of the planar surface relative to a predetermined frame of reference.
[0100] In accordance with one or more aspects of the disclosed embodiment, the controller is configured to characterize a picking surface of the extracted object based on characteristics of a planar surface that interfaces with the payload handler.
[0101] In accordance with one or more aspects of the disclosed embodiment the controller is configured to ascertain a presence and characteristics of an anomaly relative to the planar surface.
[0102] In accordance with one or more aspects of the disclosed embodiment the controller is configured to determine a logistics identity of the extracted object based on a dimension of the front surface.
[0103] In accordance with one or more aspects of the disclosed embodiment, the controller is configured to generate at least one of a run command and a stop command for the bot actuator based on the identified position and pose.
[0104] According to one or more aspects of the disclosed embodiments, there is provided a method including the steps of providing an autonomous guided vehicle, the autonomous guided vehicle including a frame with a payload holding portion, a drive section coupled to the frame including drive wheels supporting the autonomous guided vehicle on a driving surface, the drive wheels effecting vehicle movement on the driving surface and moving the autonomous guided vehicle over the driving surface at a facility, and a payload handler coupled to the frame, the payload handler configured to transport payloads having a flat, non-deterministic landing surface seated on the payload holding portion to / from the payload holding portion of the autonomous guided vehicle and a storage location for the payload in a storage array; and generating, using a vision system mounted on the frame and having a plurality of cameras, a binocular image of a field of logistics space including shelves of a rack structure on which a plurality of objects are stored. using a controller communicatively connected to the vision system to register the binocular images and generate stereo matching from the binocular images to resolve a dense depth map of imaged objects in the field; using the controller to detect stereo sets of keypoints from the binocular images, each set of keypoints being distinct and different from each other set and indicating common predetermined characteristics of each imaged object, whereby the controller derives a depth analysis result for each object from the stereo set of keypoints that is distinct and different from the dense depth map; and using an object extractor of the controller to determine the position and pose of each imaged object from both the dense depth map resolved from the binocular images and the depth analysis result from the stereo set of keypoints.
[0105] In accordance with one or more aspects of the disclosed embodiment the plurality of cameras are rolling shutter cameras.
[0106] In accordance with one or more aspects of the disclosed embodiment the method further includes analyzing registered images from video streams generated by the multiple cameras.
[0107] In accordance with one or more aspects of the disclosed embodiment, the cameras are not synchronized with one another.
[0108] In accordance with one or more aspects of the disclosed embodiment the method further includes generating a binocular image during movement of the vehicle past the object.
[0109] In accordance with one or more aspects of the disclosed embodiment, multiple objects on a rack structure are dynamically positioned in close packed juxtaposition relative to one another.
[0110] In accordance with one or more aspects of the disclosed embodiment the method further includes determining, with the controller, a front surface of the at least one extracted object and a dimension of the front surface.
[0111] In accordance with one or more aspects of the disclosed embodiment the method further includes characterizing, with the controller, a planar surface of the front surface and an orientation of the planar surface relative to a predetermined frame of reference.
[0112] In accordance with one or more aspects of the disclosed embodiment, the method further includes using the controller to characterize a picking surface of the extracted object based on characteristics of a planar surface that interfaces with the payload handler.
[0113] In accordance with one or more aspects of the disclosed embodiment the method further includes ascertaining, with the controller, the presence and characteristics of anomalies relative to the planar surface.
[0114] In accordance with one or more aspects of the disclosed embodiment the method further includes determining, with the controller, a logistics identity of the extracted object based on a dimension of the front surface.
[0115] In accordance with one or more aspects of the disclosed embodiment, the method further includes using a controller to generate at least one of a run command and a stop command for the bot actuator based on the determined position and pose.
[0116] According to one or more aspects of the disclosed embodiments, there is provided a method comprising the steps of: providing an autonomous guided vehicle, the autonomous guided vehicle including a frame with a payload holding portion, a drive section coupled to the frame with drive wheels supporting the autonomous guided vehicle on a driving surface, the drive wheels effecting vehicle movement on the driving surface and moving the autonomous guided vehicle over the driving surface at a facility, and a payload handler coupled to the frame, the payload handler configured to transport a payload having a flat, non-deterministic landing surface seated on the payload holding portion to / from the payload holding portion of the autonomous guided vehicle and a storage location for the payload in a storage array; and generating, using a vision system having a binocular imaging camera, a binocular image of a field of logistics space including shelves of a rack structure on which a plurality of objects are stored. registering the binocular images using a controller communicatively connected to the vision system, and using the controller to generate stereo matching from the binocular images to resolve a dense depth map of imaged objects in the field; detecting, using the controller, stereo sets of keypoints from the binocular images, each set of keypoints being distinct and different from each other set and indicating common predetermined characteristics of each imaged object, whereby the controller derives, from the stereo sets of keypoints, a depth analysis result for each object that is distinct and different from the dense depth map; and using an object extractor of the controller to identify the position and pose of each imaged object based on the overlay of the depth analysis result and the depth map of the stereo sets of keypoints.
[0117] In accordance with one or more aspects of the disclosed embodiment the plurality of cameras are rolling shutter cameras.
[0118] In accordance with one or more aspects of the disclosed embodiment the method further includes analyzing registered images from video streams generated by the multiple cameras.
[0119] In accordance with one or more aspects of the disclosed embodiment, the cameras are not synchronized with one another.
[0120] In accordance with one or more aspects of the disclosed embodiment the method further includes generating a binocular image during movement of the vehicle past the object.
[0121] In accordance with one or more aspects of the disclosed embodiment, multiple objects on a rack structure are dynamically positioned in close packed juxtaposition relative to one another.
[0122] In accordance with one or more aspects of the disclosed embodiment the method further includes determining, with the controller, a front surface of the at least one extracted object and a dimension of the front surface.
[0123] In accordance with one or more aspects of the disclosed embodiment the method further includes characterizing, with the controller, a planar surface of the front surface and an orientation of the planar surface relative to a predetermined frame of reference.
[0124] In accordance with one or more aspects of the disclosed embodiment, the method further includes using the controller to characterize a picking surface of the extracted object based on characteristics of a planar surface that interfaces with the payload handler.
[0125] In accordance with one or more aspects of the disclosed embodiment the method further includes ascertaining, with the controller, the presence and characteristics of anomalies relative to the planar surface.
[0126] In accordance with one or more aspects of the disclosed embodiment the method further includes determining, with the controller, a logistics identity of the extracted object based on a dimension of the front surface.
[0127] In accordance with one or more aspects of the disclosed embodiment, the method further includes using the controller to generate at least one of a run command and a stop command for the bot actuator based on the identified position and pose.
[0128] [Table 1] TIFF2025538388000003.tif204164 TIFF2025538388000004.tif210164 TIFF2025538388000005.tif202164 TIFF2025538388000006.tif206164 TIFF2025538388000007.tif198164 TIFF2025538388000008.tif206164 TIFF2025538388000009.tif210164 TIFF2025538388000010.tif200164 TIFF2025538388000011.tif208164 TIFF2025538388000012.tif92164
[0129] It should be understood that the foregoing description is merely illustrative of aspects of the disclosed embodiments. Various substitutions and modifications may be contemplated by those skilled in the art without departing from the aspects of the disclosed embodiments. Accordingly, aspects of the disclosed embodiments are intended to embrace all such substitutions, modifications, and variations that fall within the scope of any claims appended hereto. Furthermore, the mere fact that different features are recited in mutually different dependent or independent claims does not indicate that a combination of these features cannot be used to advantage and that such combination remains within the scope of aspects of the disclosed embodiments.
Claims
1. An autonomous guided vehicle, the autonomous guided vehicle comprising: a frame having a payload holder; a drive section coupled to the frame, the drive section including drive wheels that support the autonomous guided vehicle on a driving surface, the drive wheels providing vehicle movement on the driving surface and moving the autonomous guided vehicle on the driving surface at a facility; a payload handler coupled to the frame, the payload handler configured to transport a payload having a flat, non-deterministic seating surface seated on the payload holding portion to / from the payload holding portion of the autonomous guided vehicle and a storage location for the payload in a storage array; a vision system mounted on the frame, the vision system having a plurality of cameras positioned to generate a binocular image of a field of logistics space including shelves of a rack structure on which a plurality of objects are stored; a controller communicatively connected to the vision system to register the binocular images and configured to effect stereo matching from the binocular images to elucidate a dense depth map of imaged objects in the field, the controller configured to detect stereo sets of keypoints from the binocular images, each set of keypoints being distinct and different from each other set and indicative of a common predetermined characteristic of each imaged object, whereby the controller derives a depth analysis result for each object from the stereo sets of keypoints that is distinct and different from the dense depth map; the controller having an object extractor configured to determine the position and pose of each imaged object from both the dense depth map resolved from the binocular images and the depth analysis results from the stereo set of keypoints.
2. The autonomous guided vehicle of claim 1 , wherein the plurality of cameras are rolling shutter cameras.
3. The autonomous guided vehicle of claim 1 , wherein the plurality of cameras generate video streams and registered images are analyzed from the video streams.
4. The autonomous guided vehicle of claim 1 , wherein the cameras are not synchronized with one another.
5. The autonomous guided vehicle of claim 1 , wherein the binocular images are generated as the vehicle moves past the object.
6. The autonomous guided vehicle of claim 1 , wherein the plurality of objects on the rack structure are dynamically positioned in close packed juxtaposition relative to one another.
7. The autonomous guided vehicle of claim 1 , wherein the controller is configured to determine a front face of at least one extracted object and a dimension of the front face.
8. The autonomous guided vehicle of claim 7 , wherein the controller is configured to characterize a planar surface of the front face and an orientation of the planar surface relative to a predetermined frame of reference.
9. The autonomous guided vehicle of claim 8 , wherein the controller is configured to characterize a picking surface of the extracted object based on characteristics of the planar surface that interfaces with the payload handler.
10. The autonomous guided vehicle of claim 8 , wherein the controller is configured to ascertain the presence and characteristics of anomalies to the planar surface.
11. The autonomous guided vehicle of claim 7 , wherein the controller is configured to determine a logistics identity of the extracted object based on a dimension of the front surface.
12. The autonomous guided vehicle of claim 1 , wherein the controller is configured to generate at least one of run commands and stop commands for bot actuators based on the determined position and pose.
13. An autonomous guided vehicle, the autonomous guided vehicle comprising: a frame having a payload holder; a drive section coupled to the frame, the drive section including drive wheels that support the autonomous guided vehicle on a driving surface, the drive wheels providing vehicle movement on the driving surface and moving the autonomous guided vehicle on the driving surface at a facility; a payload handler coupled to the frame, the payload handler configured to transport a payload having a flat, non-deterministic seating surface seated on the payload holding portion to / from the payload holding portion of the autonomous guided vehicle and a storage location for the payload in a storage array; a vision system mounted on the frame, the vision system having a binocular imaging camera that generates a binocular image of a field of logistics space including shelves of a rack structure on which a plurality of objects are stored; a controller communicatively connected to the vision system to register the binocular images and configured to effect stereo matching from the binocular images to elucidate a dense depth map of imaged objects in the field, the controller configured to detect stereo sets of keypoints from the binocular images, each set of keypoints being distinct and different from each other set of keypoints and indicative of a common predetermined characteristic of each imaged object, whereby the controller derives a depth analysis result for each object from the stereo sets of keypoints that is distinct and different from the dense depth map; The autonomous guided vehicle, wherein the controller has an object extractor configured to identify the position and pose of each imaged object based on depth analysis of a stereo set of keypoints and an overlay of a depth map.
14. The autonomous guided vehicle of claim 13 , wherein the plurality of cameras are rolling shutter cameras.
15. The autonomous guided vehicle of claim 13 , wherein the plurality of cameras generate video streams and registered images are analyzed from the video streams.
16. The autonomous guided vehicle of claim 13 , wherein the cameras are not synchronized with one another.
17. The autonomous guided vehicle of claim 13 , wherein the binocular images are generated as the vehicle moves past the object.
18. The autonomous guided vehicle of claim 13 , wherein the objects on the rack structure are dynamically positioned in close packed juxtaposition relative to one another.
19. The autonomous guided vehicle of claim 13 , wherein the controller is configured to determine a front face of at least one extracted object and a dimension of the front face.
20. 20. The autonomous guided vehicle of claim 19, wherein the controller is configured to characterize a planar surface of the front face and an orientation of the planar surface relative to a predetermined frame of reference.
21. 21. The autonomous guided vehicle of claim 20, wherein the controller is configured to characterize a picking surface of the extracted object based on characteristics of the planar surface that interfaces with the payload handler.
22. The autonomous guided vehicle of claim 20 , wherein the controller is configured to ascertain the presence and characteristics of anomalies to the planar surface.
23. 20. The autonomous guided vehicle of claim 19, wherein the controller is configured to determine a logistics identity of the extracted object based on a dimension of the front surface.
24. The autonomous guided vehicle of claim 13 , wherein the controller is configured to generate at least one of run commands and stop commands for bot actuators based on the identified positions and poses.