Autonomous robotic movement in storage sites
An autonomous robot navigates storage sites by recognizing structures and counting rows/columns to efficiently manage inventory, addressing labor-intensive issues and improving storage site efficiency.
Patent Information
- Application Number
- JP2023501392
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-07-09
- Filing Date
- 2021-06-28
- Publication Date
- 2025-12-22
- Estimated Expiration
- 2041-06-28
AI Technical Summary
Inventory management in storage sites, such as warehouses, is labor-intensive due to misplaced items that occupy space and require time-consuming relocation, as existing systems struggle to efficiently locate and manage inventory.
A robot equipped with an image sensor and processors navigates autonomously by recognizing regularly shaped structures, counting rows and columns, and determining its location to accurately identify target inventory locations, assisted by a base station and inventory management system for data processing and command execution.
The robot efficiently manages inventory by reducing time spent on item relocation and ensuring accurate location of items, thereby optimizing storage space and improving inventory management efficiency.
Smart Images

Figure 0007789390000001 
Figure 0007789390000002 
Figure 0007789390000003
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of priority to U.S. Non-Provisional Application No. 16 / 925,241, filed July 9, 2020, which is incorporated herein by reference in its entirety for all purposes.
[0002] (Technical field) The present disclosure relates generally to robots used at storage sites, and more particularly to robots that navigate autonomously within storage sites. [Background technology]
[0003] Inventory management at storage sites, such as warehouses, can be a complex and labor-intensive process. Inventory items are typically placed in their designated locations, with the item's barcode easily scannable for fast item location. However, in some cases, items can be misplaced, making them difficult to find. Misplaced items may also occupy space reserved for upcoming inventory items, forcing storage site personnel to spend time relocating items and tracking down missing items. Summary of the Invention [Means for solving the problem]
[0004] An embodiment relates to a robot that includes an image sensor for capturing images of a storage site and one or more processors for executing various sets of instructions. The robot receives a target location for the storage site. The robot moves along a path from a base station to the target location. As the robot moves along the path, it receives images captured by the image sensor. The robot determines its current location on the path by analyzing the images captured by the image sensor and tracking the number of regularly shaped structures in the storage site that the robot has passed. The regularly shaped structures may be racks, horizontal bars on the racks, and vertical bars on the racks. The robot may determine the target location by counting the number of rows and columns that the robot has passed. [Brief explanation of the drawings]
[0005] [Figure 1] FIG. 1 is a block diagram illustrating a system environment of an exemplary storage site, according to an embodiment. [Figure 2] FIG. 1 is a block diagram illustrating components of an example robot and an example base station, according to an embodiment. [Figure 3] 1 is a flowchart illustrating an example method for managing inventory at a storage site, according to an embodiment. [Figure 4] FIG. 1 is a conceptual diagram of an example layout of a storage site with a robot, according to an embodiment. [Figure 5] 1 is a flowchart illustrating an example navigation process of a robot, according to an embodiment. [Figure 6A] 1 is a flowchart illustrating an exemplary visual referencing operation, according to an embodiment. [Figure 6B] FIG. 10 is a conceptual diagram illustrating the results of image segmentation according to an embodiment. [Figure 7] 1 is a flowchart illustrating a planner algorithm, according to an embodiment. [Figure 8] FIG. 1 is a block diagram illustrating an example machine learning model, according to an embodiment. [Figure 9]FIG. 1 is a block diagram illustrating components of an exemplary computing machine, according to an embodiment.
[0006] For purposes of illustration only, the figures depict and the detailed description sets forth various non-limiting embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0007] The figures and the following description relate only to preferred embodiments by way of example, and those skilled in the art may recognize alternative embodiments of the structures and methods disclosed herein as practical alternatives that may be employed without departing from the essence of what is disclosed.
[0008] Reference will now be made to the details of several embodiments, examples of which are shown in the accompanying figures. It should be noted that wherever feasible, similar or similar reference numerals are used in numerals, the similar or similar reference numerals may indicate similar or similar functionality. The figures represent embodiments of the disclosed system (or method) for illustrative purposes only. Those skilled in the art will readily appreciate from the following description that alternative embodiments of the structures and methods shown herein may be employed without departing from the essence described herein.
[0009] An embodiment relates to a robot that navigates a storage site by visually recognizing objects in the environment, including racks and rows and columns of racks, and counting the number of objects the robot passes by. The robot may use an image sensor to continuously capture the storage site environment. The images are analyzed using image segmentation techniques to identify the contours of readily identifiable objects, such as racks and rows and columns of racks. By counting the number of racks the robot passes by, the robot can identify the aisle in which the target location lies. The robot may also count the number of rows and columns to identify the target rack location.
[0010] System Overview FIG. 1 is a block diagram illustrating an exemplary robot-assisted or fully autonomous storage site system environment 100, according to an embodiment. By way of example, the system environment 100 includes a storage site 110, a robot 120, a base station 130, an inventory management system 140, a computer server 150, a data store 160, and user devices 170. The entities and components of the system environment 100 communicate with each other via a network 180. In various embodiments, the system environment 100 may include different, fewer, or additional components. Also, even when each component of the system environment 100 is referred to in the singular, the system environment 100 may include one or more of each component. For example, the storage site 110 may include one or more robots 120 and one or more base stations 130. Each robot 120 may have a corresponding base station 130, or multiple robots 120 may share a base station 130.
[0011] The storage site 110 may be any suitable facility for storing, selling, or displaying inventory, such as goods, merchandise, groceries, items, and collectibles. Exemplary storage sites 110 may include warehouses, inventory lots, bookstores, shoe stores, outlets, other retail stores, libraries, museums, and the like. The storage site 110 may include any number of regularly shaped structures. A regularly shaped structure may be a structure, fixture, equipment, furniture, frame, shell, rack, or other suitable object within the storage site that has a readily identifiable regular shape or contour, whether the object is permanent or temporary, fixed or movable, load-bearing or not. Regularly shaped structures are often used to store inventory at the storage site 110. For example, racks (including metal racks, shells, frames, or other similar structures) are often used to store goods and merchandise in warehouses. However, regularly shaped structures need not necessarily be used for inventory storage. A storage site 110 may include a particular layout in which various items are placed and systematically stored. For example, in a warehouse, racks may be organized into sections and separated by aisles. Each rack may include multiple pallet locations that may be identified using row and column numbers. A storage site may include high and low racks and, in some cases, may carry most inventory items primarily near ground level.
[0012] Storage site 110 may include one or more robots 120 used to track and manage inventory within storage site 110. For ease of reference, robot 120 may be referred to in the singular even if more than one robot 120 is used. Also, in some embodiments, more than one type of robot 120 may be present at storage site 110. For example, some robots 120 may be specialized in scanning inventory at storage site 110, while other robots 120 may be specialized in moving items. Robot 120 may also be referred to as an autonomous robot, an inventory cycle count robot, an inventory inspection robot, an inventory detection robot, or an inventory control robot. Inventory robots may be used to track inventory items, move inventory items, and perform other inventory control tasks. The degree of autonomy may vary between embodiments. For example, in one embodiment, robot 120 may be fully autonomous, such that robot 120 automatically performs assigned tasks. In other embodiments, robot 120 may be semi-autonomous, such that robot 120 can move around storage site 110 with minimal human command or control. In some embodiments, regardless of the degree of autonomy robot 120 has, robot 120 may be remotely operated or switched to a manual mode. Robot 120 may take various forms, such as an aerial drone, a ground robot, a vehicle, a forklift, or a mobile picking robot.
[0013] The base station 130 may be a device to which the robot 120 returns or to which the aerial robot lands. The base station 130 may include one or more return locations. The base station 130 may be used to repower the robot 120. Various methods of repowering the robot 120 may be used in various embodiments. For example, in one embodiment, the base station 130 serves as a battery exchange station that replaces the robot 120's batteries when the robot arrives at the base station so that the robot 120 can quickly resume its mission. The replaced batteries may be charged at the base station 130 via a wired or wireless connection. In another embodiment, the base station 130 serves as a charging station that includes one or more charging terminals that mate with the robot 120's charging terminals to recharge the robot 120's batteries. In yet another embodiment, the robot 120 may use fuel for power, and the base station 130 may repower the robot 120 by filling its fuel tank.
[0014] The base station 130 may also serve as a communications station for the robots 120. For example, in certain types of storage sites 110, such as warehouses, network coverage may not exist or may only exist in certain locations. The base station 130 may communicate with other components in the system environment 100 using wireless or wired communications channels, such as Wi-Fi or an Ethernet cable. When the robot 120 returns to the base station 130, the robot 120 may communicate with the base station 130. The base station 130 may send input, such as commands, to the robot 120 and download data captured by the robot 120. In embodiments where multiple robots 120 are used, the base station 130 may include a swarm controller or algorithms for coordinating movements among the robots. The base station 130 and the robots 120 may communicate using any suitable method, such as radio frequency, Bluetooth, near-field communication (NFC), or wired communication. In one embodiment, robot 120 communicates primarily with a base station, while in other embodiments, robot 120 may have the ability to communicate directly with other components within system environment 100. In one embodiment, base station 130 may serve as a wireless signal amplifier for robot 120 that communicates directly with network 180.
[0015] The inventory management system 140 may be a computer system operated by an administrator (e.g., a company that owns inventory, a warehouse management administrator, or a retailer selling inventory) using the storage site 110. The inventory management system 140 may be a system used to manage the location of inventory items. The inventory management system 140 may include a database that stores data about inventory items and item-related information, such as quantity, metadata tags, asset type tags, barcode labels, and item location coordinates within the storage site 110. The inventory management system 140 may provide both front-end and back-end software for the administrator to access the central database and inventory metrics, analyze data, generate reports, forecast future demand, and manage the location of inventory items to ensure that items are correctly located. The administrator may rely on the item coordinate data from the inventory management system 140 to ensure that items are correctly located at the storage site 110 so that they can be quickly retrieved from the storage location. This prevents misplaced items from taking up space reserved for upcoming inventory and also reduces the time it takes to locate missing items during the outbound process.
[0016] The computer server 150 may be a server tasked with analyzing data provided by the robot 120 and providing instructions for the robot 120 to perform various inventory recognition and management tasks. The robot 120 may be controlled by the computer server 150, the user device 170, or the inventory management system 140. For example, the computer server 150 may instruct the robot 120 to scan and capture images of inventory stored in various locations at the storage site 110. Based on the data provided by the inventory management system 140 and the ground truth data captured by the robot 120, the computer server 150 checks for discrepancies between the two sets of data and determines whether an item should be flagged for various reasons, such as being misplaced, lost, or damaged. The computer server 150 may then instruct the robot 120 to resolve potential problems, such as moving a misplaced item to its correct location. In one embodiment, the computer server 150 may generate a report of flagged items so that site personnel can manually resolve the issues.
[0017] The computer server 150 may include one or more computers operating at various locations. For example, a portion of the computer server 150 may be a local server residing at the storage site 110. Computer hardware, such as a processor, may be associated with a site computer or may be included in the base station 130. Another portion of the computer server 150 may be a geographically distributed cloud server. The computer server 150 may serve as a ground control station (GCS), providing data processing and maintaining end-user software used by the user devices 170. The GCS may be responsible for controlling, monitoring, and maintaining the robot 120. In some embodiments, the GCS may reside on-site as part of the base station 130. The data processing pipeline and end-user software server may reside remotely or on-site.
[0018] The computer server 150 may host software applications that allow users to manage inventory, base stations 130, and robots 120. The computer server 150 and the inventory management system 140 may or may not be operated by the same entity. In some embodiments, the computer server 150 may be operated by an entity separate from the storage site administrator. For example, the computer server 150 may be operated by a robotics service provider that supplies robots 120 and related systems to modernize and automate the storage site 110. The software applications hosted by the computer server 150 may take various forms. In some embodiments, the software applications may be integral to the inventory management system 140 or integrated as add-ons to the inventory management system 140. In other embodiments, the software applications may be separate applications that supplement or replace the inventory management system 140. In one embodiment, the software application may be provided as software as a service (SaaS) to an administrator of the storage site 110 by a robot service provider supplying the robot 120.
[0019] Data store 160 includes one or more storage devices, such as memory, in the form of a non-transitory, non-volatile computer storage medium that stores various data uploaded by robot 120 and inventory management system 140. For example, data stored in data store 160 may include images, sensor data, and other data captured by robot 120. The data may also include inventory data maintained by inventory management system 140. A computer-readable storage medium is a medium that does not involve a transitory medium, such as a propagated signal or carrier wave. Data store 160 may take various forms. In some embodiments, data store 160 communicates with other components over network 180. This type of data store 160 may be referred to as a cloud storage server. Exemplary cloud storage service providers may include AWS, Azure Storage, Google Cloud Storage, etc. In other embodiments, data store 160 is a storage device controlled by and connected to computer server 150, instead of a cloud storage server. For example, data store 160 may take the form of memory (e.g., hard drive, flash memory, disk, ROM, etc.) used by computer server 150, such as storage in a storage server room operated by computer server 150.
[0020] The user device 170 may be used by an administrator of the storage site 110 to provide commands to the robot 120 and manage inventory at the storage site 110. For example, the administrator can use the user device 170 to provide task commands to the robot for the robot to automatically complete a task. In one case, the administrator can specify a specific target location or a range of storage locations for the robot 120 to scan. The administrator may also specify specific items for the robot 120 to locate or confirm placement. Examples of user devices 170 include personal computers (PCs), desktop computers, laptop computers, tablet computers, smartphones, wearable electronic devices such as smartwatches, or other suitable electronic devices.
[0021] User device 170 may include user interface 175 taking the form of a graphical user interface (GUI). A software application provided by computer server 150 or inventory management system 140 may be represented as user interface 175. User interface 175 may take various forms. In some embodiments, user interface 175 is part of a front-end software application that includes a GUI displayed on user device 170. In one case, the front-end software application is a software application that may be downloaded and installed on user device 170 via, for example, an application store (e.g., App Store) on user device 170. In other cases, user interface 175 takes the form of a web interface of computer server 150 or inventory management system 140 through which a client may perform actions via a web browser. In other embodiments, the user interface 175 does not include graphical elements and communicates with the computer server 150 or the inventory control system 140 through other suitable methods, such as a command window or application program interfaces (APIs).
[0022] Communications between the robot 120, base station 130, inventory control system 140, computer server 150, data store 160, and user devices 170 may be transmitted over a network 180, such as the Internet. In some embodiments, the network 180 uses standard communication technologies and / or protocols. As such, the network 180 may include links using technologies such as Ethernet, 802.11, worldwide interoperability for microwave access (WiMAX), 3G, 4G, LTE, 5G, digital subscriber line (DSL), asynchronous transfer mode (ATM), InfiniBand, PCI Express, etc. Similarly, network protocols used on network 180 may include multiprotocol label switching (MPLS), transmission control protocol / Internet protocol (TCP / IP), user datagram protocol (UDP), hypertext transport protocol (HTTP), simple mail transfer protocol (SMTP), file transfer protocol (FTP), etc. Data exchanged on network 180 may be represented using technologies and / or formats including hypertext markup language (HTML), extensible markup language (XML), etc.Additionally, all or portions of the links may be encrypted using conventional encryption techniques such as secure sockets layer (SSL), transport layer security (TLS), virtual private networks (VPNs), Internet protocol security (IPsec), etc. Network 180 also includes links and packet switching networks such as the Internet. In some embodiments, two computer servers, such as computer server 150 and inventory management system 140, may communicate via an API. For example, computer server 150 may retrieve inventory data from inventory management system 140 via the API.
[0023] Illustrative robot and base station FIG. 2 is a block diagram illustrating components of an example robot 120 and an example base station 130, according to an embodiment. The robot 120 may include an image sensor 210, a processor 215, a memory 220, a flight control unit (FCU) 225 including an inertia measurement unit (IMU) 230, a state estimator 235, a visual reference engine 240, a planner 250, a communications engine 255, an I / O interface 260, and a power source 265. The functionality of the robot 120 may be distributed among various components in various ways than those described below. In various embodiments, the robot 120 may include different and / or fewer and / or additional components. Also, while each component in FIG. 2 is described in the singular, multiple components may be included. For example, the robot 120 may include one or more image sensors 210 and one or more processors 215.
[0024] The image sensor 210 captures images of the storage environment for navigation, localization, collision avoidance, object recognition and identification, and inventory person recognition purposes. The robot 120 may include one or more image sensors 210 and one or more types of such image sensors 210. For example, the robot 120 may include a digital camera that captures optical images of the environment for the state estimator 235. For example, data captured by the image sensor 210 may be provided to a VIO unit 236 included in the state estimator 235 for localization purposes, such as determining the position and orientation of the robot 120 with respect to an inertial system, such as a global frame, whose position is known and fixed. The robot 120 may also include a stereo camera with two or more lenses so that the image sensor 210 can capture three-dimensional images through stereo photography. The stereo camera may generate point cloud data for each image frame, including, for example, red, green, and blue (RGB) pixel values and depth information. Images captured by the stereo camera may be provided to the visual reference engine 240 for object recognition purposes. The image sensor 210 may be other types of image sensors, such as a light detection and ranging (LIDAR) sensor, an infrared camera, and a 360-degree depth camera. The image sensor 210 may also capture images of item labels (e.g., barcodes) for inventory cycle counting purposes. In some embodiments, a single stereo camera may be used for multiple purposes. For example, the stereo camera may provide image data to the visual reference engine 240 for object recognition. The stereo camera may be used to capture images of labels (e.g., barcodes). In some embodiments, the robot 120 includes a gimbal-like rotating stage that can rotate the image sensor 210 to various angles and stabilize the images captured by the image sensor 210. In some embodiments, the image sensor 210 may capture data along a path for the purpose of mapping a storage site.
[0025] The robot 120 includes one or more processors 215 and one or more memories 220 that store one or more sets of instructions. The one or more sets of instructions, when executed by the one or more processors, cause the one or more processors to execute processes implemented as one or more software engines. Various components of the robot 120, such as the FCU 225 and state estimator 235, may be implemented as a combination of software and hardware (e.g., sensors). The robot 120 may use a single general-purpose processor to execute the various software engines or may use separate, more specialized processors for different functions. In some embodiments, the robot 120 may use a general-purpose computer (e.g., CPU) capable of executing various sets of instructions for the various components (e.g., the FCU 225, the visual reference engine 240, the state estimator 235, and the planner 250). The general-purpose computer may run on a suitable operating system, such as LINUX, ANDROID, etc. For example, in one embodiment, the robot 120 may carry a smartphone containing an application used to control the robot. In other embodiments, robot 120 includes multiple processors specialized for various functions. For example, some of the functional components, such as FCU 225, visual reference engine 240, state estimator 235, and planner 250, may be modularized, each including its own processor, memory, and instruction set. Robot 120 may include a central processor unit (CPU) to coordinate and communicate with each modularized component. Therefore, depending on the embodiment, robot 120 may include a single processor or multiple processors 215 to perform various tasks. Memory 220 may store images and video captured by image sensor 210. Images may include captured images of the surrounding environment and images of inventory items, such as barcodes and labels.
[0026] The flight control unit (FCU) 225 may control the movement of the robot 120 through a combination of software and hardware, such as an inertial measurement unit (IMU) 230 and other sensors. For a ground robot 120, the flight control unit 225 may be referred to as a microcontroller unit (MCU). The FCU 225 controls the movement of the robot 120 based on information provided by other components. For example, a planner 250 determines the path of the robot 120 from a start point to a destination and provides commands to the FCU 225. Based on the commands, the FCU 225 generates electrical signals to various mechanical components of the robot 120 (e.g., actuators, motors, engines, wheels) to adjust the movement of the robot 120. The precise mechanical components of the robot 120 may vary depending on the embodiment and the type of robot 120.
[0027] The IMU 230 may be part of the FCU 225 or may be a separate component. The IMU 230 may include one or more accelerometers, gyroscopes, and other suitable sensors to generate measurements of forces, linear acceleration, and rotation of the robot 120. For example, an accelerometer measures forces acting on the robot 120 and senses linear acceleration. Multiple accelerometers work together to sense acceleration of the robot 120 in three-dimensional space. For example, a first accelerometer senses acceleration in the x-direction, a second accelerometer senses acceleration in the y-direction, and a third accelerometer senses acceleration in the z-direction. The gyroscope senses rotation and angular acceleration of the robot 120. Based on the measurements, the processor 215 may obtain an estimated position of the robot 120 by integrating the movement and rotation data of the IMU 230 over time.
[0028] State estimator 235 may represent a set of software instructions stored in memory 220 that may be executed by processor 215. State estimator 235 may be used to generate localization information for robot 120 and may include various sub-components for estimating the state of robot 120. For example, in one embodiment, state estimator 235 may include a visual-inertial odometry (VIO) unit 236 and an altitude estimator 238. In other embodiments, other modules, sensors, and algorithms may be used by state estimator 238 to determine the position of robot 120.
[0029] The VIO unit 236 receives image data from the image sensor 210 (e.g., a stereo camera) and measurements from the IMU 230 and generates localization information, such as the position and pose of the robot 120. Position data, obtained by double integration of acceleration measurements from the IMU 230, is often prone to drift errors. The VIO unit 236 extracts image features and tracks them in image sequences to generate optical flow vectors that represent the movement of edges, boundaries, and surfaces of objects in the environment captured by the image sensor 210. Various signal processing techniques, such as filtering (e.g., Wiener filter, Kalman filter, bandpass filter, particle filter) and optimization, and data / image transformations may be used to reduce various errors in determining localization information.
[0030] The altitude estimator 238 may be a combination of software and hardware used to measure the robot 120's absolute and relative altitude (e.g., distance from an object on the floor). The altitude estimator 238 may include a downward range finder that measures the altitude of an object below the robot 120 relative to the ground. The range finder may include an IR (or suitable signal) emitter and a sensor that senses the round-trip time of the IR reflected from an object. The altitude estimator 238 may receive data from the VIO unit 236, which estimates the robot 120's absolute altitude, but this is typically a less accurate method than a range finder. When the robot 120 flies over various objects and inventory placed on the floor or other horizontal surface, the altitude estimator 238 may include a software algorithm for combining data generated by the range finder with data generated by the VIO unit 236. Data generated by the altitude estimator 238 may be used for collision avoidance and target location. The altitude estimator 238 may set a global maximum altitude to prevent the robot 120 from hitting the ceiling. The altitude estimator 238 also provides information about how many rows of racks are below the robot 120 so that the robot 120 can locate its target position. The altitude data may be used along with the number of rows the robot 120 has passed to determine the vertical level of the robot 120.
[0031] The visual referencing engine 240 may represent a set of software instructions stored in memory 220 that may be executed by the processor 215. The visual referencing engine 240 may include various image processing and position algorithms to determine the current position of the robot 120, identify objects, edges, and surfaces in the environment near the robot 120, and determine an estimated distance and pose (e.g., yaw) of the robot 120 relative to nearby surfaces of objects. The visual referencing engine 240 may receive pixel data of successive images and point cloud data from the image sensor 210. The position information generated by the visual referencing engine 240 may include the distance and yaw from the object and the offset of the center from a target point (e.g., the center of the target object).
[0032] The visual reference engine 240 may include one or more algorithms and machine learning models for creating image segmentations from images captured by the image sensor 210. The image segmentation may include one or more segments separating frames (e.g., vertical or horizontal bars of a rack) or contours of regularly shaped structures appearing in the captured image from other objects and the environment. Algorithms used for image segmentation may include convolutional neural networks (CNNs). Other image segmentation algorithms may be used to perform the segmentation, such as edge detection algorithms (e.g., Canny, Laplace, Sobel, and Prewitt operators), corner detection algorithms, Hough transforms, and other suitable feature detection algorithms.
[0033] The visual reference engine 240 also performs object recognition (e.g., object detection and further analysis) and keeps track of the relative movement of objects across successive images. The visual reference engine 240 may track the number of regularly shaped structures at the storage site 110 that the robot 120 has passed over. For example, the visual reference engine 240 may identify a reference point (e.g., the center of gravity) on the frame of a rack and, across successive images, determine whether the reference point passes through a particular location in the image (e.g., whether the reference point passes through the center of the image). If so, the visual reference engine 240 increments the number of regularly shaped structures that the robot 120 has passed over.
[0034] The robot 120 may use various components to generate various position information, including position information and localization information for nearby objects. For example, in one embodiment, the state estimator 235 may process data from the VIO unit 236 and the altitude estimator 238 to provide localization information to the planner 250. The visual referencing engine 240 may count the number of regularly shaped structures passed by the robot 120 to determine its current position. The visual referencing engine 240 may generate position information for nearby objects. For example, when the robot 120 reaches a target position on a rack, the visual referencing engine 240 may use point cloud data to reconstruct the surface of the rack and may use depth data from the point cloud to more accurately determine the yaw and distance between the robot 120 and the rack. The visual referencing engine 240 may determine a center offset corresponding to the distance between the robot 120 and the center of the target position (e.g., the center point of the target position on the rack). The planner 250 uses the centroid offset information to control the robot 120 to move to a target location and take a picture of the inventory at the target location. When the robot 120 changes direction (e.g., turns, transitions from horizontal to vertical motion, transitions from vertical to horizontal motion, etc.), the centroid offset information may be used to determine the exact position of the robot 120 relative to an object.
[0035] Planner 250 may correspond to a set of software instructions stored in memory 220 that may be executed by processor 215. Planner 250 may include various path planning algorithms for planning a path for robot 120 as the robot moves from a first location (e.g., a start location, the current location of robot 120 after completing a previous journey) to a second location (e.g., a target location). Robot 120 may receive input, such as a user command, to perform specific actions (e.g., scanning inventory, moving items, etc.) at specific locations. Planner 250 may include two types of routes, corresponding to spot checks and range scans. In spot checks, planner 250 may receive input including coordinates of one or more specific target locations. In response, planner 250 plans a path for robot 120 to move to the target locations and perform the actions. In range scans, the input may include a range of coordinates corresponding to a range of target locations. In response, the planner 250 plans a path for the robot 120 to perform actions for a full scan or range of target locations.
[0036] Planner 250 may plan a route for robot 120 based on data provided by visual reference engine 240 and data provided by state estimator 235. For example, visual reference engine 240 may estimate the current location of robot 120 by tracking the number of regularly shaped structures in storage site 110 that robot 120 has passed. Based on the location information provided by visual reference engine 240, planner 250 may determine a route for robot 120 and adjust the movement of robot 120 as robot 120 moves along the route.
[0037] Planner 250 may include fail-safe mechanisms in case the movement of robot 120 deviates from the plan. For example, if planner 250 determines that robot 120 has passed the target path and is traveling too far from the target path, planner 250 may send a signal to FCU 225 in an attempt to correct the path. If the error is not corrected after a timeout or within a reasonable distance, or if planner 250 cannot accurately determine the current location, planner 250 may instruct the FCU to land or stop robot 120.
[0038] Relying on various position information, planner 250 may also include algorithms for collision avoidance purposes. In one embodiment, planner 250 relies on distance information, yaw angle, and center offset information relative to nearby objects to plan the movement of robot 120 to provide sufficient clearance between robot 120 and nearby objects. Alternatively or additionally, robot 120 may include one or more depth cameras, such as a set of 360-degree depth cameras, that generate distance data between robot 120 and nearby objects. Planner 250 uses position information from the depth cameras to perform collision avoidance.
[0039] The communications engine 255 and I / O interface 260 are communications components that enable the robot 120 to communicate with other components within the system environment 100. The robot 120 may communicate with external components, such as the base station 130, wirelessly or via wires using various communications protocols. Exemplary communications protocols may include Wi-Fi, Bluetooth, NFC, USB, etc., to connect the robot 120 to the base station 130. The robot 120 may transmit various types of data, such as image data, flight logs, position data, inventory data, and robot status information. The robot 120 may receive input from external sources to identify actions the robot 120 needs to perform. Commands may be generated automatically or manually by an administrator. The communications engine 255 may include algorithms for various communications protocols and standards, encoding, decoding, multiplexing, traffic control, data encryption, etc., for various communication processes. The I / O interface 260 may include software and hardware components, such as hardware interfaces, antennas, etc., for communication.
[0040] The robot 120 also includes a power source 265 used to power the various components and movements of the robot 120. The power source 265 may be one or more batteries or fuel tanks. Exemplary batteries may include lithium-ion batteries, lithium polymer (LiPo) batteries, fuel cells, and other suitable types of batteries. The batteries may be permanently installed or easily replaced. For example, the batteries may be removable so that they can be replaced when the robot 120 returns to the base station 130.
[0041] While Figure 2 shows various example components, robot 120 may include additional components. For example, mechanical features and components of robot 120 are not shown in Figure 2. Robot 120 may include various types of motors, actuators, robotic arms, lifts, other moving components, and other sensors, depending on the type of robot, to perform various tasks.
[0042] 2, the exemplary base station 130 includes a processor 270, a memory 275, an I / O interface 280, and a repowering unit 285. In various embodiments, the base station 130 may include different and / or fewer and / or additional components.
[0043] The base station 130 includes a processor 270 and one or more memories 275 containing one or more sets of instructions that cause the processor 270 to execute various processes implemented as one or more software modules. The base station 130 may provide input and commands to the robot 120 to perform various inventory control tasks. The base station 130 may include instruction sets for performing group control among multiple robots 120. Group control may include task allocation, routing and planning, and coordinating movements between the robots to avoid collisions. The base station 130 may serve as a central control unit for coordinating the robots 120. The memory 275 may also include various sets of instructions for performing analysis of data and images downloaded from the robot 120. The base station 130 may provide various degrees of data processing, from formatting raw data to complete data processing to generate information useful for inventory management. Alternatively or additionally, base station 130 may directly upload data downloaded from robot 120 to a data store, such as data store 160. Base station 130 may provide operational, management, and administrative commands to robot 120. In some embodiments, base station 130 may be remotely controlled by user device 170, computer server 150, or inventory management system 140.
[0044] The base station 130 may include various types of I / O interfaces 280 for communication with the robot 120 and for communication with the Internet. The base station 130 may continuously communicate with the robot 120 using wireless protocols such as Wi-Fi or Bluetooth. In some embodiments, one or more components of the robot 120 of FIG. 2 may reside within the base station, which may provide commands to the robot 120 for movement and navigation. Alternatively or additionally, when the robot 120 lands or stops at the base station 130, the base station 130 may communicate with the robot 120 via a short-range communication protocol such as NFC or a wired connection. The base station 130 may be connected to a network 180, such as the Internet. The storage site 110's wireless network (e.g., a LAN) may not have sufficient coverage. The base station 130 may be connected to the network 180 via an Ethernet cable.
[0045] The repowering unit 285 includes components used to detect the power level of the robot 120 and repower the robot 120. Repowering can be accomplished by replacing batteries, recharging batteries, refilling fuel tanks, etc. In one embodiment, the base station 130 includes a mechanical actuator, such as a robot arm, to replace batteries in the robot 120. In other embodiments, the base station 130 may be used as a charging station for the robot 120, either via wired or inductive charging. For example, the base station 130 may include a landing or resting pad with an induction coil underneath to wirelessly charge the robot 120 via an induction coil within the robot. Other suitable methods of repowering the robot 120 are also possible.
[0046] Example inventory management method 3 is a flowchart illustrating an exemplary method for managing inventory at a storage site, according to an embodiment. The method may be performed by a computer, which may be a single operating unit in the traditional sense (e.g., a single personal computer) or a set of distributed computers (e.g., a virtual machine, a distributed computer system, cloud computing, etc.) collaborating to execute a set of instructions. Also, although a computer is described in the singular, the computer performing the method of FIG. 3 may include one or more computers associated with computer server 150, inventory management system 140, robot 120, base station 130, or user device 170.
[0047] According to an embodiment, a computer receives 310 a configuration of a storage site 110. The storage site 110 may be a warehouse, a retail store, or other suitable location. The configuration information for the storage site 110 may be uploaded to the robot 120 for the robot to navigate the storage site 110. The configuration information may include the total number of regularly shaped structures in the storage site 110 and dimensional information for the regularly shaped structures. The provided configuration information may be in the form of a computer-aided design (CAD) drawing or other type of file format. The configuration may include a layout of the storage site 110, such as the layout of racks and the arrangement of other regularly shaped structures. The layout may be a two-dimensional layout. The computer extracts the number of sections, aisles, and racks, and the number of rows and columns for each rack, from the CAD drawing by counting their appearance in the CAD drawing. The computer may also extract the height and width of the rack cells from the CAD drawing or other source. In some embodiments, the computer does not need to extract the exact distance between a given set of racks or the width of each aisle or the total length of the racks. Instead, the robot 120 may measure the aisle, rack, and cell dimensions from depth sensor data, or may use a counting method implemented by the planner 250 in conjunction with the visual referencing engine 240 to traverse the storage site 110 by counting the number of rows and columns that the robot 120 passes. Therefore, in some embodiments, the exact dimensions of the racks may not be required.
[0048] Some configuration information may be manually entered by an administrator of storage site 110. For example, the administrator may provide the number of sections, the number of aisles and racks in each section, and the cell dimensions of the racks. The administrator may also enter the number of rows and columns for each rack.
[0049] Alternatively or additionally, configuration information may be obtained from a mapping process, such as pre-flight mapping or a mapping process performed as the robot 120 performs inventory control tasks. For example, for a storage site 110 newly implementing an automated management method, an administrator may provide the size of the navigable space of the storage site to one or more mapping robots, which may count sections, aisles, and rows and columns of regularly shaped structures within the storage site 110. Also, in one embodiment, the mapping or configuration information does not need to measure precise distances between racks or other structures within the storage site 110. Instead, the robot 120 may navigate the storage site 110 with only a rough layout of the storage site 110 by counting regularly shaped structures along a path to identify target locations. As the robot 120 continues to perform various inventory control tasks, the robotic system may gradually perform mapping or estimates of the size and location of various structures.
[0050] The computer receives (320) inventory control data for inventory control operations at the storage site 110. Certain inventory control data may be manually entered by an administrator, while other data may be downloaded from the inventory management system 140. The inventory control data may include scheduling and planning for inventory control activities, including operation frequency, time windows, etc. For example, the management data may specify that each location on the racks at the storage site 110 should be scanned at predetermined intervals (e.g., daily) and that the inventory scanning process should be performed by the robot 120 overnight after the storage site is closed. The inventory management system 140 data may provide information such as the item's barcode and label, the exact coordinates of the inventory item, and information about the rack and other storage space that needs to be cleared for upcoming inventory. The inventory control data may also include items that need to be retrieved from the storage site 110 each day (e.g., purchase order items that need to be shipped), and the robot 120 may need to focus on those items.
[0051] The computer creates 330 a plan to perform the inventory control. For example, the computer may create an automated plan that includes various commands to instruct the robot 120 to perform various scans. The commands may specify a range of locations the robot 120 needs to scan or one or more specific locations the robot 120 needs to go to. The computer may estimate the time per scanning trip and design a plan for each work interval based on the time available for robotic inventory control. For example, at a particular storage site 110, robotic inventory control may not be performed during business hours.
[0052] The computer generates 340 various commands to operate one or more robots 120 to move through the storage site 110 based on the plan and information derived from the configuration of the storage site 110. The robots 120 may move through the storage site 110 by at least visually recognizing and counting the number of regularly shaped structures within the storage site. In one embodiment, in addition to location techniques such as those used by VIO, the robot 120 counts the number of racks, rows, and columns that it has passed through to determine its current location along a path from a start location to a destination location without knowing the exact distance and direction that the robot 120 has traveled.
[0053] Inventory scanning or other inventory control tasks may be performed autonomously by the robot 120. In one embodiment, a scanning task begins at the base station 130, where the robot 120 receives input (342) including coordinates of a target location or a range of target locations within the storage site 110. The robot 120 departs (344) from the base station 130. The robot 120 navigates (346) the storage site 110 by visually recognizing regularly shaped structures. For example, the robot 120 tracks the number of regularly shaped structures it passes. The robot 120 turns and transforms based on the recognized regularly shaped structures captured by the image sensor 210. Upon arriving at the target location, the robot 120 may align itself with a reference point (e.g., a center position) of the target location. At the target location, the robot 120 captures target location data (e.g., measurements, photographs, etc.) including the inventory item, barcode, and label on the inventory item's box (348). If the initial command before the robot 120 departs includes multiple target locations or a range of target locations, the robot 120 continues to the next target location by moving up, down, or sideways to the next location to continue the scanning operation.
[0054] After completing a scanning trip, the robot 120 returns to the base station 130 by counting the number of regularly shaped structures it has passed in reverse. The robot 120 may implicitly recognize the structures it has passed as it moves to its target location. Alternatively, the robot 120 may return to the base station 120 by reversing its path without counting. The base station 130 repowers the robot 120. For example, the base station 130 may provide the robot 120 with a next command and replace the robot's 120's batteries (352) so that the robot 120 can quickly return to service another scanning trip. Used batteries may be charged at the base station 130. The base station 130 may download the data and images captured by the robot 120 and upload the data and images to the data store 160 for further processing. Alternatively, the robot 120 may include wireless communication components for transmitting its data and images to the base station 130 or directly to the network 180 .
[0055] The computer performs (360) an analysis of the data and images captured by the robot 120. For example, the computer may compare barcodes (including serial numbers) in images captured by the robot 120 with data stored in the inventory management system 140 to determine if items are misplaced or missing within the storage site 110. The computer may also determine other conditions of the inventory. The computer may generate reports for display on the administrator user interface 175 to take remedial action for misplaced or missing inventory. For example, reports may be generated daily so that personnel at the storage site 110 manually locate and remove misplaced items. Alternatively or additionally, the computer may generate an automation plan for the robot 120 to remove misplaced inventory. The data and images captured by the robot 120 may be used to verify the movement or arrival of inventory items.
[0056] Example of movement method FIG. 4 is a conceptual diagram of an example layout of a storage site 110 equipped with a robot 120, according to an embodiment. FIG. 4 illustrates a two-dimensional layout of the storage site 110, with a close-up view of an example rack shown in inset 405. The storage site 110 may be divided into various regions based on regularly shaped structures. In this example, the regularly shaped structures are racks 410. The storage site 110 may be divided by sections 415, aisles 420, rows 430, and columns 440. For example, a section 415 is a collection of racks. Each aisle may have racks on both sides. Each rack 410 may include one or more columns 440 and multiple rows 430. A storage unit of a rack 410 may be referred to as a cell 450. Each cell 450 may hold one or more pallets 460. In this particular example, two pallets 460 are placed in each cell 450. Inventory for storage site 110 is transported on pallets 460. The division and designation shown in Figure 4 is used as an example only; storage site 110 in other embodiments may be divided differently.
[0057] Each inventory item at storage site 110 may reside on a pallet 460. The target location (e.g., pallet location) of an inventory item may be identified using a coordinate system. For example, an item placed on pallet 460 may have an aisle number (A), a rack number (K), a row number (R), and a column number (C). For example, pallet location coordinates of [A3, K1, R4, C5] mean that pallet 460 is located at rack 410 in the north rack of aisle 3. Pallet 460's location in rack 410 is the fourth row and fifth column (counting from the ground). In some cases, such as the particular layout shown in FIG. 4, an aisle 420 may have racks 410 on both sides. Additional coordinate information may be used to identify racks 410 on the north side of aisle 420 and racks 410 on the south side of aisle 420. Alternatively, the upper and lower racks may have different aisle numbers. For spot checks, the robot may be provided with a single coordinate for a single location or multiple coordinates for more than one location. For a range scan to check for pallets 460 within a range, the robot may be provided with a range of coordinates, such as an aisle number, a rack number, a starting row, a starting column, an ending row, and an ending column. In some embodiments, the pallet location coordinates may be referred to in different ways. For example, in one case, the coordinate system may adopt the format "aisle-rack-shelf-location." The shelf number may correspond to the row number, and the location number may correspond to the column number.
[0058] Referring to FIG. 5 in conjunction with FIG. 4, FIG. 5 is a flowchart illustrating an example method of movement of the robot 120, according to an embodiment. The robot 120 receives 510 a target location 474 at the storage site 110. The target location 474 may be expressed in a coordinate system, as described above with respect to FIG. 4. The target location 474 may be received as an input command from the base station 130. The input command may also include an action the robot 120 needs to take, such as taking a picture at the target location 474 to capture the barcode and label of an inventory item. The robot 120 may generate location information based on the VIO unit 236 and the altitude estimator 238. In some cases, the starting location of the route is the base station 130. In some cases, the starting location of the route may be anywhere at the storage site 110. For example, the robot 120 may have recently completed a task and may begin another task without returning to the base station 130.
[0059] A processor of the robot 120, such as one executing the planner 250, controls 520 the robot 120 along a path 470 to a target location 474. The path 470 may be determined based on the coordinates of the target location 474. The robot 120 may rotate so that the image sensor 210 faces a regularly shaped structure (e.g., a rack). Movement of the robot 120 to the target location 474 may include moving to a particular aisle, rotating to enter the aisle, moving horizontally to a target column, moving vertically to a target row, and rotating to the correct angle to face the target location 474 to take a picture of the inventory items on the pallet 460. An example of the precision planning and actions taken by the robot 120 is described in further detail with respect to FIG. 7.
[0060] As the robot 120 moves to the target location 474, the robot 120 captures (530) images of the storage site 110 using the image sensor 210. The captured images may be sequential images. The robot 120 receives (540) images captured by the image sensor 210 as the robot 120 moves along the path 470. The images may capture objects in the environment, including regularly shaped structures such as racks. For example, the robot 120 may use algorithms in the visual reference engine 240 to visually recognize regularly shaped structures.
[0061] The robot 120 analyzes (550) the images captured by the image sensor 210 to determine its current location on the path 470 by tracking the number of regularly shaped structures in the storage site that the robot 120 has passed. The robot 120 may use various image processing and object recognition techniques, described in more detail with respect to FIG. 6A, to identify the regularly shaped structures and track the number of structures that the robot 120 has passed. With respect to the path 470 shown in FIG. 4, the robot 120 may move toward the rack 410 until it reaches the turning point 476. The robot 120 determines that it has reached the target aisle because it has passed two racks 410. In response, the robot 120 rotates counterclockwise and enters the target aisle facing the target rack. The robot 120 counts the number of rows it has passed until it reaches the target row. Following the line of targets, the robot 120 may move vertically up or down to reach the target location. Once at the target location, the robot 120 performs the action specified in the input command, such as taking a picture of the inventory item at the target location.
[0062] Visual reference tasks 6A is a flowchart illustrating example visual referencing operations performed by one or more algorithms of robot 120 to recognize objects in an environment, track the number of structures passed by robot 120 on its path, and identify various types of location information, according to an embodiment. The operations illustrated in FIG. 6A may be performed by algorithms consistent with visual referencing engine 240 of robot 120.
[0063] The robot 120 receives (600) image data from an image sensor 210 capturing an image of the environment of the storage site 110. In one embodiment, the image sensor 210 is a stereo camera including multiple lenses to capture depth information. The image data received by the robot 120 may include pixel data in any suitable format, such as RGB, or may also include a point cloud with spatial information, such as x, y, and z dimensions, where z represents depth information (e.g., the distance between the robot 120 and an object). For example, the image sensor 210 may include a depth camera that generates depth data of objects within the depth camera's field of view. The image data may take the form of a series of sequential images capturing gradual changes in the relative positions of objects in the environment. In some embodiments, the image sensor 210 may be other types of sensors, such as a standard digital camera, a LIDAR sensor, an infrared sensor, etc., that generate various types of image data. For example, the robot 120 in one embodiment may use a standard digital camera that generates only RGB pixel data.
[0064] Based on the received image data, the robot 120 may determine various types of information related to the position of the robot 120, such as the number of regularly shaped structures the robot 120 has passed over, the yaw angle and distance between the robot 120 and the object, and the offset of the center of gravity of the robot 120 relative to the structures. The position information may be generated by the visual reference engine 240 and provided to the planner 250 so that the planner 250 can control the movement of the robot 120. The position information may be provided to the planner 250 in various ways. Certain information may be provided to the planner 250 constantly or periodically. Other information may be provided to the planner 250 upon request from the planner 250. For example, in one embodiment, the count of structures the robot 120 has passed over and yaw angle and distance information may be provided to the planner 250 constantly, while the offset of the center of gravity may be provided to the planner 250 upon request. When the robot 120 changes direction (eg, turns, transitions from horizontal to vertical motion, transitions from vertical to horizontal motion), the planner 250 often requires center offset information.
[0065] The robot 120 may use wireframe segmentation 605 to recognize nearby structures, whether they are regularly shaped structures or other types of structures. The robot 120 creates 610 an image segmentation of the image captured by the image sensor 210. The robot 120 segments the image based on the type of object captured in the image. For example, the robot 120 may separate readily identifiable objects, such as regularly shaped structures (e.g., racks), from the rest of the environment by image segmentation that identifies the contours of the structures. The image segmentation may further identify various portions of regularly shaped structures, such as vertical and horizontal bars, in each image. Based on the identified structures (or regions of structures), the robot 120 may mark pixels in the image with various landmarks. For example, pixels corresponding to the identified vertical bars may be marked with a "1" or a first appropriate symbol, pixels corresponding to the identified horizontal bars may be marked with a "2" or a second appropriate symbol, and pixels consistent with the rest of the environment may be marked with a "0" or a third appropriate symbol.
[0066] Image segmentation may be performed by any suitable algorithm. In one embodiment, the robot 120 stores a trained machine learning model to perform image segmentation and assign labels to pixels. The robot 120 inputs images into the machine learning model, which outputs various types of landmarks for image segments. In one embodiment, the robot 120 inputs a series of images, and the machine learning model outputs an image segmentation from the series of images by considering objects that appear consecutively in the series of images. In one embodiment, the machine learning model is a convolutional neural network. The convolutional neural network may be trained using a training set of images of various storage site environments containing various types of structures. For example, the convolutional neural network may be trained by iteratively reducing the error in assigning segments to the training set of images. An exemplary structure of a convolutional neural network used to perform image segmentation is described in further detail in FIG. 8. Alternatively or additionally to convolutional neural networks, other types of machine learning models may be used in the image segmentation process, such as other types of neural networks, clustering, Markov random fields (MRFs), etc. Alternatively or additionally to using machine learning techniques, other image segmentation algorithms may be used, such as edge detection algorithms (e.g., Canny operator, Laplace operator, Sobel operator, Prewitt operator), corner detection algorithms, Hough transforms, and other suitable feature detection algorithms.
[0067] From the pixels labeled with symbols indicating regularly shaped structures, the robot 120 creates a contour of the area of those pixels (615). Each contour may be referred to as a cluster of pixels. The clustering of those pixels may be based on the distance between pixels, the color of the pixels, the intensity of the pixels, and / or other suitable characteristics of the pixels. The robot 120 may cluster nearby pixels that are similar (e.g., in terms of distance and / or color and / or intensity) to create a shape that is likely to correspond to part of the regularly shaped structure. Multiple clusters corresponding to various subregions may be created for each regularly shaped structure (e.g., each vertical bar). The robot 120 identifies a reference point for each contour (620). The reference point may be the centroid, the most extreme point in one direction, or any relevant reference point. For example, the robot 120 may determine the average pixel position of the pixels within the contour to identify the centroid.
[0068] The robot 120 performs noise filtering and contour merging (625), which merges contours based on their respective reference points. In noise filtering, the robot 120 may consider contours whose area is less than a threshold (e.g., too small contours) as noise. The robot 120 preserves the sufficient size of the contours. In some cases, contour merging results in pixels matching a structure (e.g., a vertical bar) that the clustering algorithm classifies into multiple contours (e.g., multiple regions of a vertical bar). The robot 120 merges contours that may represent the same portion of the structure. Merging may be based on the location of the contour's reference points and the contour's boundaries. For example, in a horizontal bar, the robot 120 may identify contours with reference points at similar vertical levels and merge those contours. In some cases, when two contours are merged, pixels between the two contours, which may belong to a smaller cluster or may not be identified in any cluster, may be classified as the same structure. Merging may be based on the distance between two reference points of the two contours. If the distance is less than a threshold level, the robot 120 may merge the two contours. Figure 6B shows an example of a regularly shaped structure identified by image segmentation, contouring, and shape merging. Figure 6B shows that the robot 120 can separate the image of the rack from the rest of the environment.
[0069] Based on the identified features, such as the vertical and horizontal bars of the regularly shaped structure, the robot 120 tracks (630) the number of regularly shaped structures that the robot 120 has passed along the path to the goal location. This information may be identified by the visual referencing engine 240 and output to the planner 250. The robot 120 may determine the count of regularly shaped structures that the robot 120 has passed by by monitoring the relative motion of the regularly shaped structures captured in the image sequence. For example, the robot 120 may use the wireframe segmentation 605 to identify that a portion (e.g., a vertical bar) of the regularly shaped structure (e.g., a rack) was captured in one image in the image sequence. This image in the image sequence may be referred to as the first image. The robot 120 then considers the next image in the image sequence. The robot 120 identifies the same vertical bar in the image following the first image. For example, the robot 120 may identify the vertical bar in a second image immediately following a first image that is closest to the position of the vertical bar that appeared in the first image as the same vertical bar. As the robot 120 continues to review successive images, the vertical bar may continue to be tracked until it moves out of the image boundary. Alternatively or additionally, when tracking the same portion of the same structure, the robot 120 may predict the movement of the regularly shaped structure. For example, if the robot 120 moves from left to right, the robot 120 expects the same regularly shaped structure to appear from right to left in successive images. Alternatively or additionally, the robot 120 may predict the change in position of the same regularly shaped structure from one image to another, taking into account the planned movement speed and the frame rate at which the images are captured.
[0070] To determine whether the robot 120 has passed a regularly shaped structure, the robot 120 may establish one or more objective criteria. The robot 120 may track the movement of a reference point on a portion of the target regularly shaped structure across a first image and a second image in a sequence of images. In response to the reference point on the portion of the structure passing a specific location in the image (e.g., the midpoint of the image), the robot 120 increments the count of regularly shaped structures (or subtypes, such as vertical bars, horizontal bars, etc.) that the robot 120 has passed. For example, in a sequence of images, the robot 120 may track the center of gravity of a horizontal bar to determine whether the center of gravity has moved through the center of the image. When this condition is met, the robot 120 increments the count of the vertical bar that the robot 120 has passed and continues to track the next vertical bar captured in the image. By counting these structures, the robot 120 can track the number of racks, aisles, rows, and columns that the robot 120 has passed through at the storage site 110 .
[0071] In some embodiments, the robot 120 may use other information in addition to the data from the wireframe partition 605 to track the number of regularly shaped structures that the robot 120 has passed over. For example, the robot 120 may use the altitude estimator 238 to identify the row level at which the robot is located. A count of the horizontal bars passed by the robot 120, generated from the wireframe partition 605, can be used to verify the identification from the altitude estimator 238.
[0072] In some embodiments, the visual referencing engine 240 may use the information generated by the wireframe segmentation 605 to determine other position data for the robot 120. For example, information from the wireframe segmentation 605 can be used in conjunction with the point cloud data to determine the yaw and distance of the robot 120 relative to an object. The determined vertical and horizontal bars may also be used to determine a center offset, which will be described in further detail.
[0073] In addition to analyzing the image data for wireframe segmentation 605, the robot 120 may use the image data to generate position information such as yaw and distance. For example, the robot 120 may receive point cloud data (635) from an image sensor 210, such as a stereo camera. The robot 120 then performs a point cloud filtering operation (640). Based on the information from the wireframe segmentation 605, the filtering operation removes points that are not regularly shaped structures and retains points identified as likely to correspond to regularly shaped structures. The robot 120 may identify points that correspond to regularly shaped structures appearing in the captured image at the target location. The robot 120 may then place the remaining points in a histogram based on the depth value of each point. The robot 120 may select one or more bars in the histogram with the most points. In other words, the robot 120 may retain the dominant points near the peak, which are likely to correspond to regularly shaped structures. Other filtering techniques may be used to select points for further processing.
[0074] Based on the filtered point cloud, the robot 120 fits (645) the point cloud to a plane to estimate one or more surfaces of the regularly shaped structure. Various plane adaptation techniques, such as least-squares plane adaptation, may be used. The robot 120 determines (650) yaw and distance information using the adapted plane. For example, the robot 120 uses depth data of the plane-adapted points to determine the distance between the robot 120 and the regularly shaped structure represented by the plane. The robot 120 may also determine an estimated yaw angle of the robot relative to the regularly shaped structure represented by the plane. The robot 120 may determine the yaw angle using geometric operations. The position information, including the yaw and distance, may be provided to the planner 250.
[0075] The position information determined by the visual referencing engine 240 may also include the center offset of the robot 120. The center offset may be the offset of the robot 120 relative to the center of the location (usually the location of a pallet, but may also be the location of a cell or other suitable location). The visual referencing engine 240 determines 660 the center offset of the robot 120 relative to the location of the front of the robot 120. The visual referencing engine 240 may determine the center offset based on the wireframe division 605 and yaw and distance information 650 determined from the point cloud data. FIG. 6B shows an example image annotated with the robot 120 to illustrate the determination of the center offset. The dot 670 shown in FIG. 6B indicates the projection of the robot 120's camera onto the regularly shaped structure. The center 690 of the image in FIG. 6B is assumed to be the location of the robot 120's camera. The center 690 is projected horizontally and vertically onto the regularly shaped structure, appearing as four dots 670. The visual referencing engine 240 uses information from the wireframe segmentation 605 to generate a merged contour 680. Based on the merged contour 680, the robot 120 identifies a center 692 of a pallet location 694. A center offset may be identified based on the center 692 of the pallet location 694 and the center 690 of the image.
[0076] The robot 120 may be instructed to take a picture of the pallet location. The robot 120 uses position information such as center offset, yaw, and distance to determine the offset of the robot 120 from the center of the target location. In response to determining that the center offset is within a threshold distance from the center of the pallet location (or other suitable target point), the robot 120 may take a picture of the target pallet location with the image sensor 210 so that it can take the inventory items currently placed on the pallet.
[0077] Example planner work FIG. 7 is a flowchart illustrating an example of a planning algorithm for the robot 120 to determine a path to a target location. The process may be performed by the planner 250. The planning algorithm may include a rack mode 710 and a cell mode 745. In rack mode 710, the robot 120 first identifies the number of racks (or other suitable regularly shaped structures) that the robot has passed and identifies the appropriate aisle for the robot to enter. After the robot 120 enters the aisle, the robot 120 may change to cell mode 745, which identifies the correct cell containing the pallet corresponding to the target location. The robot 120 may rely on the object recognition and number of regularly shaped structures generated by the visual referencing task described in FIG. 6A.
[0078] The robot 120 receives (705) a command for a spot check or range scan mission. A spot check mission may involve visiting one or more separate target locations that are not necessarily related to each other. A range scan mission may involve scanning a range of target locations, such as several columns and rows. For example, a range scan command may include a specific range, including a starting spot and an ending spot. In some embodiments, the robot 120 may distinguish between a spot check and a range scan mission by the number of spots in the command. For either a spot check or range scan, the command includes at least one coordinate specifying the target location for the spot check or the starting spot for the range scan. The coordinate may specify an aisle number, a rack number, a row number, and a column number. For a range scan mission, the robot 120 may also be provided with a range of target location coordinates. Based on the range, the robot 120 may determine a route for the range scan that is efficient for completing the range scan. For example, the robot 120 may determine that it is more efficient to move horizontally across a row first, and then move up to scan another row.
[0079] In response to receiving a spot check command or a range scan command, the robot 120 enters rack mode 710. The robot 120 receives position information (715), such as data generated by the visual reference engine 240. The robot 120 may identify a first target location. The first target location may be one of the spots in a spot check mission or the initial location in a range scan mission. In rack mode 710, the robot 120 moves toward a target aisle based on the coordinates of the first target location (720). The robot 120 visually identifies regularly shaped structures, for example, counting the number of regularly shaped structures the robot 120 has passed. Based on the visual reference information, the robot 120 determines (725) whether the current aisle count is equal to the target aisle number (or whether the rack count is equal to the target rack number). If not, the robot 120 continues moving toward the target aisle and keeps track of the rack count. If the aisle count is equal to the target aisle number, the robot 120 enters the target aisle (735). Based on the position information, the robot 120 determines a sufficient clearance distance between the robot 120 and the rack to avoid a collision with the rack when the robot 120 enters the target aisle (735).
[0080] Upon entering the target aisle, the robot 120 may enter cell mode 745. The robot 120 moves to a first position in the aisle (750). The robot 120 moves horizontally toward the target location (755). The robot 120 counts the number of rows the robot 120 has passed based on the visual reference data. The robot 120 determines whether the row count is equal to the target row number (760). If not, the robot 120 continues moving horizontally toward the target location. If the row count is equal to the target row number, the robot 120 determines that the robot 120 has reached the target row and may move vertically toward the target location (765). For example, the robot may move one or more rows, upward or downward. The robot 120 determines whether the row count is equal to the target row number (770). If not, the robot 120 continues moving vertically toward the target location. If they are equal, the robot 120 determines that it has reached the target location. Based on the position information, such as the center offset, the robot 120 receives position center alignment information to align itself with the target point of the target location (775). The robot 120 performs the assigned inventory control task, such as taking a picture of the target location (780). Upon completing the inventory control task, the robot 120 may continue moving according to the remaining spot checks or range scans (785).
[0081] The order shown in the various flowcharts in this disclosure may be changed. For example, the order of horizontal movement (755) and vertical movement (765) may be swapped. If a range or multiple locations of target locations are provided (e.g., in a range scan mission), the robot 120 may perform the assigned task at the first target location and then move to the next location to perform another task.
[0082] Example machine learning model In various embodiments, a wide variety of machine learning techniques may be used. Examples include various forms of supervised, unsupervised, and semi-supervised learning, such as decision trees, support vector machines (SVMs), regression, Bayesian networks, and genetic algorithms. Deep learning techniques such as neural networks, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), and long short-term memory networks (LSTMs), may also be used. For example, the image segmentation process described in FIG. 6A and various object recognition, localization, and other processes may apply one or more machine learning and deep learning techniques. In one embodiment, image segmentation is performed using a CNN, an example structure of which is shown in FIG. 8.
[0083] In various embodiments, training techniques for machine learning models may be supervised, semi-supervised, or unsupervised. In supervised learning, the machine learning model may be trained with a set of labeled training examples. For example, for a machine learning model trained to classify objects, the training examples may be various photographs of objects labeled with the object type. The labels for each training example may be binary or multi-class. In training a machine learning model for image segmentation, the training examples may be photographs of regularly shaped objects from various repository sites with manually identified portions of the images. In some cases, unsupervised learning techniques may be used. The samples used for training are unlabeled. Various unsupervised learning techniques, such as clustering, may be used. In some cases, training may be semi-supervised, with the training set having a mixture of labeled and unlabeled examples.
[0084] Machine learning models may be associated with an objective function that generates a metric that describes the desired goal of the training process. For example, training may be intended to reduce the model's error rate in generating predictions. In this case, the objective function may monitor the error rate of the machine learning model. In object recognition (e.g., object detection and classification), the objective function of a machine learning algorithm may be the training error rate in classifying objects in a training set. Such an objective function may be called a loss function. Other forms of objective functions may be used, particularly for unsupervised learning models that are unlabeled and therefore have an error rate that cannot be easily determined. In image segmentation, the objective function may correspond to the difference between the segments predicted by the model and the segments manually identified in the training set. In various embodiments, the error rate may be measured as a cross-entropy loss, an L1 loss (e.g., the sum of absolute differences between predicted and observed values), or an L2 loss (e.g., the sum of squared distances).
[0085] Referring to FIG. 8, an exemplary CNN structure is shown, according to an embodiment. CNN 800 may receive input 810 and generate output 820. CNN 800 may include various types of layers, such as convolutional layers 830, pooling layers 840, recurrent layers 850, fully connected layers 860, and custom layers 870. A convolutional layer 830 convolves one or more kernels with the layer's input (e.g., an image) to generate various types of images filtered by the kernels and generate feature maps. Each convolution result may be associated with an activation function. A convolutional layer 830 may be followed by a pooling layer 840 that selects the maximum value (max pooling) or the average value (average pooling) from the portion of the input covered by the kernel size. The pooling layer 840 reduces the spatial size of the extracted features. In some embodiments, a pair of convolutional layers 830 and pooling layers 840 may be followed by a recurrent layer 850 that includes one or more feedback loops 855. Feedback 855 may be used to evaluate spatial relationships between image features or temporal relationships between objects in the image. Layers 830, 840, and 850 may be followed by multiple fully connected layers 860 with interconnected nodes (represented by squares in FIG. 8 ). Fully connected layers 860 may be used for classification and object detection. In some embodiments, one or more custom layers 870 may be present for generating output 820 in a specific format. For example, a custom layer may be used in image segmentation for labeling pixels of an image input with various segment labels.
[0086] The layer order and number of layers in the CNN 800 of FIG. 8 are for illustrative purposes only. In various embodiments, the CNN 800 includes one or more convolutional layers 830, but may or may not include pooling layers 840 or recurrent layers 850. When a pooling layer 840 is present, not all convolutional layers 830 are necessarily followed by a pooling layer 840. Recurrent layers may be variously positioned in other locations in the CNN. The size of the kernels (e.g., 3×3, 5×5, 7×7, etc.) and the number of kernels allowed to be trained in each convolutional layer 830 may be different from the other convolutional layers 830.
[0087] A machine learning model may include specific layers and / or nodes and / or kernels and / or coefficients. Training a neural network such as the CNN800 may include forward propagation and backpropagation. Each layer of a neural network may include one or more nodes that are fully or partially connected to other nodes in adjacent layers. In forward propagation, a neural network performs a forward calculation based on the output of the previous layer. The work of a node may be defined by one or more functions. The functions defining the work of a node may include various computational tasks such as convolution of data with one or more kernels, pooling, recurrent loops in RNNs, and various gates in LSTMs. The functions may also include activation functions that adjust the weights of the node's output. Nodes in different layers may be associated with different functions.
[0088] Each of the neural network's functions may be associated with various coefficients (e.g., weights and kernel coefficients) that can be adjusted during training. Additionally, some of the nodes in a neural network may be associated with activation functions that determine the node's output weights in forward propagation. Common activation functions may include step functions, linear functions, sigmoid functions, hyperbolic tangent (tanh) functions, and rectified linear unit (ReLU) functions. After an input is provided to the neural network and passed forward through the neural network, the results may be compared with training labels or other values in the training set to determine the neural network's performance. The prediction process may be repeated for other images in the training set to calculate the value of the objective function for a particular training round. The neural network then performs backpropagation using a gradient descent method, such as stocgastic gradient descent (SGD), to adjust the coefficients of the various functions and improve the value of the objective function.
[0089] Multiple rounds of forward and backpropagation may be performed. Training may cease when the objective function is sufficiently stable (e.g., the machine learning model has converged) or after a predetermined number of rounds on a particular set of training examples. The trained machine learning model may be used for prediction, object detection, image segmentation, or any other suitable task for which the model was trained.
[0090] Computer Architecture 9 is a block diagram illustrating components of an exemplary computing machine capable of reading instructions from a computer-readable medium and executing the instructions on a processor (or controller). The computer described herein may include a single computing machine as shown in FIG. 9, or a virtual machine, or a distributed computing system including multiple nodes of the computing machine as shown in FIG. 9, or any other suitable arrangement of computing devices.
[0091] 9 shows a diagrammatic representation of a computing machine in the exemplary form of a computer system 900, within which instructions 924 (e.g., software, program code, or machine code) stored on a computer-readable medium may be executed to cause the machine to perform one or more processes described herein. In some embodiments, the computing machine may operate as a standalone device or may be connected (e.g., networked) to other machines. In a networked deployment, the machine may operate in the capacity of a server machine or a client machine in a server-client network environment, or may operate as a peer machine in a peer-to-peer (or distributed) network environment.
[0092] The computing machine configuration described in Figure 9 may correspond to any software or hardware components or combinations of components shown in Figures 1 and 2, including, but not limited to, inventory management system 140, computer server 150, data store 160, user device 170, and the various engines, modules, interfaces, terminals, and machines shown in Figure 2. Although Figure 9 shows various hardware and software elements, each of the components described in Figures 1 and 2 may include additional or fewer elements.
[0093] By way of example, the computing machine may be a personal computer (PC), or a tablet PC, or a set-top box, or a personal digital assistant (PDA), or a mobile phone, or a smartphone, or a web appliance, or a network router, or an internet of things (IoT) appliance, or a switch or bridge, or any machine capable of executing instructions 924 that specify actions to be taken by the machine. Additionally, while only one machine is illustrated, the term "machine" may be taken to include a collection of machines that individually or collectively execute instructions 924 to perform one or more methodologies described herein.
[0094] The exemplary computer system 900 includes one or more processors (generally, processor 902) (e.g., a central processing unit (CPU), or a graphics processing unit (GPU), or a digital signal processor (DSP), or one or more application-specific integrated circuits (ASICs), or one or more radio-frequency integrated circuits (RFICs), or a combination thereof), a main memory 904, and a non-volatile memory 906, which are configured to communicate with each other via a bus 908. The computer system 900 may further include a graphics display device 910 (e.g., a plasma display panel (PDP), or a liquid crystal display (LCD), or a projector, or a cathode ray tube (CRT)). The computer system 900 may also include an alphanumeric input device 912 (e.g., a keyboard), a cursor control device 914 (e.g., a mouse, or a trackball, or a joystick, or a motion sensor, or other pointing means), a storage device 916, a signal generating device 918 (e.g., a speaker), and a network interface device 920, which are also configured to communicate via the bus 908.
[0095] The storage device 916 includes a computer-readable medium 922 having stored thereon instructions 924 embodying one or more methodologies or functions described herein. The instructions 924 may reside, completely or at least partially, within the main memory 904 or within the processor 902 (e.g., within the processor's cache memory) during execution by the computer system 900, with the main memory 904 and the processor 902 also constituting computer-readable media. The instructions 924 may be transmitted or received over a network 926 via a network interface device 920.
[0096] Although computer-readable medium 922 is shown to be a single medium in the illustrated embodiment, the term "computer-readable medium" should be interpreted to include one medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) capable of storing instructions (e.g., instructions 924). Computer-readable media may include any medium capable of storing instructions (e.g., instructions 924) for execution by a machine, causing the machine to perform one or more of the methodologies described herein. Computer-readable media include, but are not limited to, data repositories in the form of solid-state memory, optical media, and magnetic media. Computer-readable media does not include transitory media such as signals or carrier waves.
[0097] Further compositional considerations Advantageously, by recognizing regularly shaped structures within a storage site and using the count of structures to navigate the storage site, a robotic system automates the storage site without adding significant load to the storage site. Conventional robots often require code markings, such as QR codes, on each rack to recognize the racks as they navigate the storage site. However, adding QR codes to a storage site imposes significant upfront costs for automation on conventional storage sites. The use of QR codes also incurs ongoing maintenance costs for the storage site. Use of the robotic system of some embodiments allows the robot to navigate the storage site without prior mapping.
[0098] Particular embodiments are described herein as including logic circuits, or several components, engines, modules, or mechanisms. An engine may comprise a software module (e.g., code embodied on a computer-readable medium) or a hardware module. A hardware engine is a tangible device capable of performing specific tasks, and may be configured or arranged in a particular way. In example embodiments, one or more hardware engines of one or more computer systems (e.g., standalone, client, or server computer systems) or computer systems (e.g., processors or groups of processors) may be configured with software (e.g., applications or application portions) as hardware engines that operate to perform specific tasks as described herein.
[0099] In various embodiments, a hardware engine may be implemented mechanically or electronically. For example, a hardware engine may include dedicated circuitry or logic circuitry that is permanently configured to perform specific tasks (e.g., as a dedicated processor such as a field programmable gate array (FPGA) or application-specific integrated circuit (ASIC)). A hardware engine may also include programmable logic circuitry or circuitry that is temporarily configured by software to perform specific tasks (e.g., contained within a general-purpose processor or other programmable processor). The decision to implement a hardware engine mechanically, dedicated, permanently configured circuitry, or temporarily configured circuitry (e.g., configured in software) is, of course, based on cost and time considerations.
[0100] Various operations of the example methods described herein may be performed, at least in part, by one or more processors, such as processor 902, that are temporarily or permanently configured (e.g., by software) to perform the associated operations. Whether temporarily or permanently configured, such a processor may constitute a processor-implemented engine that operates to perform one or more operations or functions. An engine as referred to herein may, in some example embodiments, include a processor-implemented engine.
[0101] Performance of a particular task may be distributed across one or more processors located across several machines, rather than residing solely within one machine. In some example embodiments, one or more processors or processor-implemented modules may reside in one geographic location (e.g., in a home environment, or in an office environment, or in a server farm). In other example embodiments, one or more processors or processor-implemented modules may be distributed across several geographic locations.
[0102] Upon reading this disclosure, those skilled in the art will recognize still additional alternative structural and functional designs of similar systems or processes through the principles disclosed herein. Thus, while specific embodiments and applications are shown and described, it should be understood that the disclosed embodiments are not limited to the exact structures and components disclosed herein. It will be apparent to those skilled in the art that various modifications, changes, and variations may be made in the arrangements and details of the operations and methods disclosed herein without departing from the spirit and scope as defined in the appended claims.
[0103] (Addendum) (Appendix 1) 1. A method of operating an inventory robot, comprising: controlling the inventory robot to move along a path to a target location; capturing images using one or more image sensors as the inventory robot moves along the path; analyzing the images captured by the image sensor to identify a current location of the inventory robot on the path by tracking the number of regularly shaped structures within the storage site that the inventory robot has passed; A method comprising:
[0104] (Appendix 2) Analyzing the image further includes generating an image segmentation from the captured image; the image segmentation in one of the images comprises segments of frames of the regularly shaped structure appearing in the one of the images; analyzing the images further includes identifying the regularly shaped structures using the image segmentation of the one of the images. The method described in Appendix 1.
[0105] (Appendix 3) 2. The method of claim 1, wherein the image segmentation is generated by using at least a convolutional neural network.
[0106] (Appendix 4) Analyzing the image may further comprise: determining a reference point of a portion of the regularly shaped structure appearing in the first image; identifying the portion of one of the regularly shaped structures in a second image subsequent to the first image; tracking movement of the reference point across the first image and the second image; In response to a reference point passing a particular location with respect to the first image and the subsequent image, increasing the number of regularly shaped structures passed by the inventory robot; 2. The method of claim 1, comprising:
[0107] (Appendix 5) It also involves capturing point cloud data using a depth camera. the point cloud data includes depth data; further identifying one of the regularly shaped structures appearing in one of the images; estimating a plane aligned with the surface of said one of said regularly shaped structures; and using the depth data and the plane of the point cloud data to determine the current position of the inventory robot relative to the one of the regularly shaped structures; 2. The method of claim 1, comprising:
[0108] (Appendix 6) 2. The method of claim 1, further comprising determining an estimated yaw angle of the inventory robot relative to the one of the regularly shaped structures and an estimated distance between the inventory robot and the one of the regularly shaped structures.
[0109] (Appendix 7) 2. The method of claim 1, further comprising taking a photograph of the target location with the image sensor in response to determining that the offset of the center of the inventory robot is within a threshold distance from the target location.
[0110] (Appendix 8) 2. The method of claim 1, wherein the starting location is a base station configured to charge the inventory robot.
[0111] (Appendix 9) 2. The method of claim 1, wherein the regularly shaped structure comprises a rack having rows and columns.
[0112] (Appendix 10) 2. The method of claim 1, wherein determining the current location of the inventory robot is based on the number of racks passed by the inventory robot and the number of rows and columns passed by the inventory robot.
[0113] (Appendix 11) further receiving the target location defined by coordinates specifying a target aisle, a target rack, a row number of the rack, and a column number of the rack; In response to receiving the coordinates, counting the number of racks passed by the inventory robot based on the rack number; turning into an aisle in which the target rack is located; counting the number of rows passed by the inventory robot based on the row number to find the target location; moving to the line corresponding to said line number; 2. The method of claim 1, comprising:
[0114] (Appendix 12) 2. The method of claim 1, wherein the inventory robot is an aerial unmanned aerial vehicle.
[0115] (Appendix 13) An inventory robot, an image sensor configured to capture an image of the storage site; one or more processors; one or more memories coupled to the one or more processors; Including, the one or more memories storing one or more sets of instructions; The one or more sets of instructions, when executed by the one or more processors, cause the one or more processors to: receiving a target location of the storage site; controlling the inventory robot to move along a path to the target location; capturing images using one or more image sensors as the inventory robot moves along the path; analyzing the images captured by the image sensor to identify a current location of the inventory robot on the path by tracking the number of regularly shaped structures within the storage site that the inventory robot has passed; Inventory robot.
[0116] (Appendix 14) 14. The inventory robot of claim 13, wherein the regularly shaped structure includes a rack having rows and columns.
[0117] (Appendix 15) 14. The inventory robot of claim 13, wherein determining the current location of the inventory robot is based on a number of racks passed by the inventory robot and the number of rows and columns passed by the inventory robot.
[0118] (Appendix 16) The instructions for analyzing the image may include: generating an image segmentation from said captured image; the image segmentation in one of the images comprises segments of frames of the regularly shaped structure appearing in the one of the images; The instructions for analyzing the image may include: further comprising: identifying said regularly shaped structures using said image segmentation of said one of said images; The inventory robot described in Appendix 13.
[0119] (Appendix 17) 14. The inventory robot of claim 13, wherein the image sensor includes a depth camera configured to generate depth data for objects within its field of view.
[0120] (Appendix 18) 14. The inventory robot of claim 13, further comprising a visual inertial odometry (VIO) unit configured to access the images stored in the one or more memories to determine the position and orientation of the inventory robot.
[0121] (Appendix 19) 14. The inventory robot of claim 13, wherein the one or more memories further store a total number of the regularly shaped structures in the storage site and dimensional information for the regularly shaped structures.
[0122] (Appendix 20) 1. A system for managing a storage site, the system comprising: (1) an inventory robot including one or more processors and one or more memories coupled to the one or more processors; the one or more memories storing one or more sets of instructions; The one or more sets of instructions, when executed by the one or more processors, cause the one or more processors to: receiving a target location of the storage site; controlling the inventory robot to move along a path to the target location; capturing images using one or more image sensors as the inventory robot moves along the path; analyzing the images captured by the image sensor to identify a current location of the inventory robot along its path by tracking the number of regularly shaped structures within the storage site that the inventory robot has passed; The system further comprises: (2) a base station configured to repower the inventory robot; the base station is configured to download the photograph of the target location in response to the inventory robot returning to the base station; system.
Claims
1. 1. A method of operating an inventory robot, comprising: controlling the inventory robot to move along a path to a target location; capturing images using one or more image sensors as the inventory robot moves along the path; analyzing the image captured by the image sensor; Including, Analyzing the image captured by the image sensor includes: determining a reference point of a portion of a regularly shaped structure appearing in the first image; identifying the portion of one of the regularly shaped structures in a second image subsequent to the first image; tracking movement of the reference point across the first image and the second image; incrementing a number of regularly shaped structures passed by the inventory robot in response to a reference point passing a particular location with respect to the first image and the subsequent second image; determining a current location of the inventory robot along the path by counting and incrementing the number of regularly shaped structures within the storage site that have been passed by the inventory robot along the path; Including, method.
2. Analyzing the image further includes generating an image segmentation from the captured image; the image segmentation in one of the images comprises segments of frames of the regularly shaped structure appearing in said one of the images; analyzing the images further includes identifying the regularly shaped structures using the image segmentation of the one of the images. The method of claim 1.
3. The method described in claim 1, wherein the image segmentation of the captured image is created by using at least a convolutional neural network.
4. Further, it includes capturing point cloud data using a depth camera; the point cloud data includes depth data; further identifying one of the regularly shaped structures appearing in one of the images; estimating a plane aligned with one surface of said regularly shaped structure; using the depth data and the plane of the point cloud data to determine the current position of the inventory robot relative to the one of the regularly shaped structures; 10. The method of claim 1, comprising:
5. 2. The method of claim 1, further comprising determining an estimated yaw angle of the inventory robot relative to the one of the regularly shaped structures and an estimated distance between the inventory robot and the one of the regularly shaped structures.
6. 2. The method of claim 1, further comprising: taking a photograph of the target location with the image sensor in response to determining that the offset of the center of the inventory robot is within a threshold distance from the target location.
7. The method of claim 1, wherein the starting location is a base station configured to charge the inventory robot.
8. The method of claim 1 , wherein the regularly shaped structure comprises a rack having rows and columns.
9. The method of claim 1 , wherein determining the current location of the inventory robot is based on the number of racks passed by the inventory robot and the number of rows and columns passed by the inventory robot.
10. further receiving the target location defined by coordinates specifying a target aisle, a target rack, a row number of the rack, and a column number of the rack; In response to receiving the coordinates, counting the number of racks passed by the inventory robot based on rack numbers; turning into an aisle in which the target rack is located; counting the number of rows passed by the inventory robot based on the row number to find the target location; moving to the line corresponding to said line number; 10. The method of claim 1, comprising:
11. The method of claim 1 , wherein the inventory robot is an aerial unmanned aerial vehicle.
12. An inventory robot, an image sensor configured to capture an image of the storage site; one or more processors; one or more memories coupled to the one or more processors; Including, the one or more memories storing one or more sets of instructions; The set of one or more instructions, when executed by the one or more processors, cause the one or more processors to: receiving a target location of the storage site; controlling the inventory robot to move along a path to the target location; capturing images using one or more image sensors as the inventory robot moves along the path; analyzing the image captured by the image sensor; analyzing the image captured by the image sensor includes: determining a reference point of a portion of a regularly shaped structure appearing in the first image; identifying the portion of one of the regularly shaped structures in a second image subsequent to the first image; tracking movement of the reference point across the first image and the second image; incrementing a number of regularly shaped structures passed by the inventory robot in response to a reference point passing a particular location with respect to the first image and the subsequent second image; determining a current location of the inventory robot along the path by counting and incrementing the number of regularly shaped structures within a storage site that have been passed by the inventory robot along the path; Including, Inventory robot.
13. 13. The inventory robot of claim 12, wherein the regularly shaped structure comprises a rack having rows and columns.
14. 13. The inventory robot of claim 12, wherein determining the current location of the inventory robot is based on a number of racks passed by the inventory robot and the number of rows and columns passed by the inventory robot.
15. The instructions for analyzing the image may include: generating an image segmentation from said captured image; the image segmentation in one of the images comprises segments of frames of the regularly shaped structure appearing in said one of the images; The instructions for analyzing the image may include: further comprising: identifying said regularly shaped structures using said image segmentation of said one of said images; 13. The inventory robot of claim 12.
16. The inventory robot of claim 12 , wherein the image sensor comprises a depth camera configured to generate depth data for objects within its field of view.
17. 13. The inventory robot of claim 12, further comprising a visual inertial odometry (VIO) unit configured to access the images stored in the one or more memories to determine a position and pose of the inventory robot.
18. 13. The inventory robot of claim 12, wherein the one or more memories further store a total number of the regularly shaped structures within the storage site and dimensional information for the regularly shaped structures.
19. 1. A system for managing a storage site, the system comprising: (1) an inventory robot including one or more processors and one or more memories coupled to the one or more processors; the one or more memories storing one or more sets of instructions; The set of one or more instructions, when executed by the one or more processors, cause the one or more processors to: receiving a target location of the storage site; controlling the inventory robot to move along a path to the target location; capturing images using one or more image sensors as the inventory robot moves along the path; analyzing the image captured by the image sensor; analyzing the image captured by the image sensor includes: determining a reference point of a portion of a regularly shaped structure appearing in the first image; identifying the portion of one of the regularly shaped structures in a second image subsequent to the first image; tracking movement of the reference point across the first image and the second image; incrementing a number of regularly shaped structures passed by the inventory robot in response to a reference point passing a particular location with respect to the first image and the subsequent second image; determining a current location of the inventory robot along the path by counting and incrementing the number of regularly shaped structures within a storage site that have been passed by the inventory robot along the path; Including, The system further comprises: (2) a base station configured to repower the inventory robot; the base station is configured to download a photograph of the target location in response to the inventory robot returning to the base station. system.
Citation Information
Patent Citations
Article management device and article management method
JP2018131331A
Systems and methods for locating, identifying and counting items
JP2019513274A
Inventory Management
JP2019527172A
Robotic drone
WO2018035482A1