Analyzing sensor data to identify events

By using sensor data and machine vision algorithms to identify item handover events within facilities, this technology solves the problem of virtual shopping cart updates in existing technologies, and achieves efficient user item handover identification and payment processing.

CN116615743BActive Publication Date: 2026-07-31AMAZON TECH INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AMAZON TECH INC
Filing Date
2021-11-30
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently identify and update events within facilities, particularly user virtual shopping cart information during item handover, especially for customized or variable-weight items.

Method used

By using data generated by sensors, combined with machine vision algorithms and a trained classifier, user gestures within the area of ​​interest of the scanning device are identified, item handover events are determined, and the corresponding user's virtual shopping cart information is updated.

Benefits of technology

It enables efficient identification of item handover events within facilities and accurate updating of virtual shopping carts, supporting payment processing in the grab-and-go retail model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116615743B_ABST
    Figure CN116615743B_ABST
Patent Text Reader

Abstract

This disclosure relates to a technique in which a first user (102) in an environment (106) first scans a visual marker associated with an object (104) (such as a barcode) and then hands the object over to a second user (110). One or more computing devices (118) may receive instructions to scan, retrieve image data of the interactive activity from a camera (116) within the environment (106), identify the user (110) who received the object, and then update a virtual shopping cart (138) associated with the second user (110) to indicate the addition of the object (104).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Patent Application No. 17 / 106,755, filed November 30, 2020, entitled “ANALYZING SENSOR DATA TOIDENTIFY EVENTS,” the entire contents of which are incorporated herein by reference. Background Technology

[0003] Retailers, wholesalers, and other product distributors typically maintain inventory of various items that customers or clients can order, buy, rent, borrow, lease, or view. For example, an e-commerce website might maintain inventory in an order fulfillment center. When a customer orders an item, it is picked from the inventory, sent to a packing station, packaged, and shipped to the customer. Similarly, physical stores maintain inventory in areas accessible to customers (such as shopping areas) where they can pick items and take them to the checkout to buy, rent, etc. Many physical stores also maintain inventory in storage areas, order fulfillment centers, or other facilities that can be used to replenish inventory in shopping areas or fulfill orders for items placed through other channels (such as e-commerce). Other examples of entities maintaining facilities that hold inventory include libraries, museums, rental centers, etc. In each case, users move around within the facility to move items from one location to another, pick items from their current location, and transfer them to a new location. It is generally desirable to generate information about events occurring within the facility. Attached Figure Description

[0004] The accompanying drawings will be described in detail. In these drawings, the leftmost digit of the reference numeral indicates the drawing in which that reference numeral first appears. The same reference numerals are used in different drawings to denote similar or identical objects or features.

[0005] Figure 1 An exemplary architecture is shown in which a first user in the environment scans a visual marker (such as a barcode) associated with an object before handing the object to a second user. The architecture also includes a server computing device configured to receive the scan instruction, retrieve image data of the interactive activity from cameras within the environment, identify the user who received the object, and then update a virtual shopping cart associated with the second user to indicate the addition of the object.

[0006] Figures 2A to 2CAn exemplary sequence of operations is illustrated, in which a first user scans an object and hands it to a second user, while one or more computing devices analyze image data of the environment near the interactive activity during the scan to determine a user identifier associated with the second user who receives the object.

[0007] Figure 3A An image analysis component is shown that uses machine vision algorithms to generate segmentation maps of image data frames. As shown, the segmentation map can indicate pixels associated with different objects (such as background, user's hand, objects, etc.).

[0008] Figure 3B An image analysis component is shown that uses a segmentation map and one or more trained classifiers to determine, for each frame of image data, whether the frame includes a hand, and if so, the location of the hand and whether it is empty or full. This image analysis component can use this information to identify the user who acquires the object after scanning it with a scanning device.

[0009] Figures 4A to 4B A flowchart of an exemplary process is shown, which is used to update the virtual shopping cart data of a user who receives an item scanned by another user using a scanning device.

[0010] Figure 5 A flowchart of another exemplary process is shown, which is used to update the virtual shopping cart data of a user who receives an item.

[0011] Figure 6 This is a block diagram of an exemplary material handling facility site that includes sensors and an inventory management system configured to use sensor data to generate outputs about events occurring in the facility site.

[0012] Figure 7 A block diagram is shown of one or more servers configured to support the operation of a facility site. Detailed Implementation

[0013] This disclosure relates to systems and techniques for identifying events occurring within a facility using sensor data generated by sensors in that facility. In one example, the techniques and systems may identify an event in response to a first user scanning a visual marker (e.g., a barcode, QR code, etc.) of an object using a scanning device and a second user receiving the object. After recognizing that the second user has received the object, the systems and techniques may update the virtual shopping cart associated with the second user to indicate the addition of the item. In some examples, these systems and techniques may be implemented in a "grab-and-go" retail environment, where virtual shopping carts for individual users can be maintained so that users can pick up and otherwise receive items and then "leave directly" from the facility, and then, in response to the user leaving the facility, the user's corresponding pre-stored payment method can be charged for their virtual shopping cart.

[0014] In some of the examples below, an employee at a facility may scan a visual tag for an item before delivering it to a customer within the facility. In some cases, items include custom-made items (e.g., custom salads or sandwiches), items of variable weight (e.g., a certain quantity of seafood or meat), items of variable quantity (e.g., a certain length of quilting material), or any other type of item whose cost varies depending on the actual quantity received by the customer or other parameters (e.g., garnishes). Therefore, after a customer places an order (e.g., ordering a pound of shrimp, a yard of fabric, etc.), an employee may pack the requested items in the required quantity and may use a printing device to print a visual tag associated with the item. For example, the employee may use a scale to measure the quantity of the item and a printer to print a barcode, QR code, etc., where the visual tag encodes information about the item, such as the item identifier, the weight / quantity of the item, the cost of the item, the scanning time, etc. In some cases, the employee may then affix the visual tag to the item or its packaging. For example, a printer can print self-adhesive labels that include visual markings, which employees can then affix to objects or packaging.

[0015] After a visual tag is affixed to an object, an employee can then scan the visual tag using a scanning device. For example, an employee can use a barcode scanner, a tablet computer, or any other device including a camera or other imaging equipment to identify or otherwise capture information about the visual tag. The scanning device can generate scan data that indicates the aforementioned information, such as the object's identification, weight / quantity, cost, and scan time. This scan data can then be sent via one or more networks to one or more computing devices configured to verify the scan data and attempt to determine events associated with the object linked to it. The scan data sent to the computing devices may include, or may be accompanied by, the identifier of the employee logged into the scanning device, the identifier of the scanning device, and / or similar data.

[0016] The computing device can receive scan data and, in response or at a later time, attempt to determine any events related to that scan data. For example, the computing device can identify the scanning device that generated the scan data and determine the object of interest (VOI) associated with that scanning device within the facility premises. That is, each scanning device in the facility premises can be associated with a corresponding VOI within that facility premises, which may include a portion of the environment (e.g., an environment defined in XYZ coordinates) in which a customer might interact with an object associated with the scan data (e.g., receive an object). For example, the VOI of each scanning device may include an XYZ “frame” defined relative to the corresponding scanning device, such as a bounding box that spans some or all of the countertops where the scanning device is located and extends upwards by a predefined length (e.g., to the ceiling). In some cases, these VOIs are manually configured for each scanning device in the environment, while in others, VOIs can be determined by analyzing image data of the scanning devices and the area surrounding them using cameras within the facility premises (e.g., overhead cameras).

[0017] In either case, upon receiving scan data, or at some point thereafter, the computing device can determine the identifier of the scanning device used to generate the scan data, which can then be used to determine the corresponding VOI (Void of View) of the scanning device. After identifying the VOI, the computing device can determine which cameras within the facility have a field of view (FOV) that includes the VOI. Furthermore, the computing device can determine the time associated with the scanning of the object from the scan data. After identifying one or more cameras with the current VOI in their FOV, the computing device can retrieve image data generated by that camera (or cameras) before and after the scanning time. That is, the computing device can retrieve image data spanning the object scanning (e.g., exactly before and after the scanning) at a time such as when the scanning begins or immediately after the scanning.

[0018] Upon receiving the image data, the computing device can use one or more trained classifiers to determine any events involving objects associated with the scan. For example, the computing device can first perform segmentation techniques on individual frames of the image data to identify the content represented within the frames. For instance, the computing device can be configured to identify users, user body parts (e.g., hands, head, body, arms, etc.), backgrounds, tabletops, etc., within the image data frames. These devices can utilize classifiers trained using supervised learning or other techniques to identify predefined objects. In some cases, these classifiers output indications corresponding to pixel values ​​for different objects, such as an indication that a first pixel at a first location corresponds to the background, a second pixel corresponds to a hand, etc.

[0019] The computing device can use this segmentation map to identify whether one or more hands are present within a spatiotemporal window surrounding the object scan. That is, the device can determine whether any frame of image data within a scanning threshold time period includes a hand within the VOI of the scanning device. If included, the computing device can determine the user identifier of the user whose hand is within the VOI within the scanning threshold time period. In response to identifying the user identifier, the computing device can update the virtual shopping cart associated with that user.

[0020] In some cases, the computing device can make this determination over time, rather than solely based on the identification of a hand within an image data frame. For example, in addition to being trained to recognize a user's hand, a classifier can be trained to output a score indicating whether the hand is likely full or empty. Furthermore, the computing device can store corresponding indications of a first score and a second score, where the first score indicates whether the corresponding frame includes a hand, and the second score indicates whether the identified hand (if included) is empty or full. The computing device can also store the position of the identified hand. Some or all of this information can be stored over time; for example, the device can identify motion vectors indicating how the identified hand moves within the VOI over time.

[0021] For example, after determining the spatiotemporal window associated with the scanned data, the computing device can analyze the image within that spatiotemporal space by generating feature data associated with the image data frame and inputting that feature data into one or more trained classifiers. These classifiers can use the image data to indicate the position of a hand, whether the hand is empty or full (if any), and the hand's position (if any). In response to recognizing an empty hand "entering" the VOI and a full hand "leaving" the VOI, the computing device can determine that the user associated with that hand has received an object.

[0022] Therefore, in response to determining that a user has received an object, the system can invoke or otherwise interact with a positioning component configured to maintain the location of each user identifier within the facility at a given time. That is, after a user enters the facility, the system can assign a user identifier to that user (in some cases, this user identifier may not contain personally identifiable data) and can use image data and / or similar data to maintain the position of that user identifier within the facility over time. Thus, when the computing device determines that a hand has received an object within a spatiotemporal window associated with scanned data, the computing device can use the positioning component to determine which user identifier was present at the object's location at the time of handover. Upon receiving an indication from a user identifier, the system can update the virtual shopping cart associated with that user identifier to indicate the addition of the object. For example, the user's virtual shopping cart can be updated to indicate that she has received an object. A 1.2-pound shrimp costs $9.89.

[0023] Therefore, these technologies enable customers to request customized or weight / size-variable items from employees, who can then prepare the item, print a visual label indicating the item's cost, affix the label to the item, scan the label, and then deliver the item to the requesting user. In response or at some point after the interaction, the systems and techniques described herein can analyze image data representing the interaction to determine which user actually received the item. For example, the systems and techniques can analyze a predefined VOI associated with a scanning device within a threshold time period to identify the presence of a hand within the VOI and potentially identify information about the hand, such as whether it is full or empty, its position over time, and / or similar information. This information can be used to determine that the hand (and thus the user) has indeed received the item. After making this determination, the systems and techniques can then determine which user is associated with the hand and, following this determination, update the corresponding user's virtual shopping cart.

[0024] In some cases, the systems and techniques described herein can be executed in response to an instruction from a computing device to receive scanned data. That is, in response to an employee (or other user) scanning an object, these techniques can be executed to determine a user identifier associated with the user who received the object. Meanwhile, in other cases, different triggers may lead to this determination. For example, these techniques can be executed within a system for determining the contents of a user's virtual shopping cart in response to a user leaving the facility. In this example, the system may identify a set of candidate users for each potential event within the facility and may address each event for that particular user in response to the user leaving the environment. For example, when scanning an object at a meat counter in the facility, if a particular user is nearby (e.g., within a threshold distance), that user may be flagged as a candidate user for the specific event involving the scanned object. When the user leaves the store, the system can use the techniques described above and below to determine whether the user actually received the object, and if so, update the user's virtual shopping cart at that time. In summary, different triggers may lead to the execution of the techniques described herein.

[0025] Furthermore, while the above examples describe using a vision-based classifier to determine the presence (and possible state and orientation) of a hand in a VOI to determine if a user has received an object, in other cases, one or more other factors may be used to make that determination. For example, a vision algorithm may be used to track an object scanned by a scanning device until that object is received by another user in the facility. That is, when an object is scanned, one or more computer vision algorithms may be used to identify the scanned object and track its position in the image data over time (e.g., within the VOI or elsewhere) at least until a user different from the user who scanned the object receives it. At this point, the receiving user's virtual shopping cart can be updated.

[0026] Alternatively or additionally, the location of an employee (or other user scanning an item) can be maintained over time. For example, while an employee is scanning an item, one or more computer vision algorithms can be used to identify the employee performing the scan and continue to locate that employee in the image data over time until the item is passed to another user. Similarly, this may result in updates to the virtual shopping cart associated with the receiving user.

[0027] Furthermore, while the discussion above and below includes examples of scanning devices associated with fixed locations, it should be understood that these techniques can also be applied to mobile scanning devices, such as those used by store employees to scan items and hand them over to the appropriate users. For example, when an employee uses a mobile scanning device to scan an item, the device can provide the scan data to one or more computing devices mentioned above. These computing devices can use the identifier of the mobile scanning device to determine the current location of the mobile device and / or the employee within the facility premises. That is, the computing devices can use the system's tracking component to determine the location of the device and / or employee, which maintains the current location of customers and employees within the store. This information can then be used to determine one or more cameras (e.g., overhead cameras) with the current location of the mobile scanning device in the field of view. The computing devices can then acquire image data from these cameras to analyze the VOI (Voice of Interest) around the scanning device. As mentioned above, the VOI can be defined relative to the location of the scanning device. The computing devices can then use the techniques described above and further detailed below to identify the hands of customers within the VOI, such as the hand when a customer enters the VOI and the hand when leaving the VOI. The user identifier associated with that hand can then be used to update the virtual shopping cart of the corresponding user.

[0028] Finally, while the examples included herein are described with reference to a single object, it should be understood that these techniques can also be applied to multiple objects. In these cases, an employee may scan multiple objects sequentially before handing the group of objects or a container (e.g., a bag or box) containing the multiple objects to a customer. Here, the computing device can first determine, through the corresponding scan data, that these objects were scanned within a threshold amount of time relative to each other. For example, the computing device can determine that the time elapsed between scans of an object and a subsequent object is less than the threshold amount of time, therefore, the items were scanned sequentially. In response to making this determination, the computing device can analyze the VOI of the hand using the techniques described above and can associate each of the sequentially scanned objects with a determined user identifier associated with that hand. Thus, if an employee scans, for example, five objects before handing them to a customer (e.g., in a bag), these techniques can associate each of the five objects with the same customer whose hand was identified within the VOI. Furthermore, while the example above describes a computing device determining that objects are related based on a set of objects being scanned within a threshold time period, in another example, the scanning device may include controls (e.g., icons) that an employee can select to indicate that multiple objects will be scanned, and these objects are for a single customer. Therefore, when the computing device receives scan data for multiple objects, it can determine, based on the described hand recognition and hand tracking techniques, that each object is associated with the same user identifier.

[0029] The following description introduces the use of these technologies in material handling facilities. The facilities described herein may include, but are not limited to, warehouses, distribution centers, docking and distribution facilities, order fulfillment center facilities, packaging facilities, transportation facilities, leasing facilities, libraries, retail stores, wholesale stores, museums, or other facilities or combinations of facilities performing one or more functions of material (inventory) handling. In other embodiments, the technologies described herein may be implemented in other facilities or other settings. Certain specific embodiments and implementations of this disclosure will now be described more fully below with reference to the accompanying drawings, in which various aspects are illustrated. However, the aspects may be implemented in many different forms and should not be construed as limited to the specific embodiments set forth herein. This disclosure includes variations of the embodiments as described herein. The same reference numerals refer to the same elements throughout.

[0030] Figure 1 An exemplary architecture 100 is illustrated, in which a first user 102 in environment 106 scans object 104 using scanning device 108. For example, the first user 102 may include an employee of a retail facility (e.g., staff, etc.) and may use any type of device to scan visual markings (such as barcodes, QR codes, text, etc.) associated with object 104, whichever type of device is capable of generating scan data that encodes or otherwise indicates object details. For example, scanning device 108 may include a device configured to read barcodes or QR codes affixed to object 104. Meanwhile, in some cases, object 104 may include custom-made or variable-weight objects; therefore, the first user 102 may use scales, printing equipment, and / or similar devices to generate visual markings, such as barcodes. After physically printing the barcode, etc., the first user may affix (e.g., adhere) the visual markings to object 104 before or after scanning object 104.

[0031] After scanning object 104 using scanning device 108, first user 102 may hand object 104 to second user 110, or place object 104 on a counter for second user 110 to retrieve, etc. In either case, second user 110 may reach their hand 112 into a body of interest (VOI) 114 associated with scanning device 108 to receive object 104. VOI 114 may include a three-dimensional area within environment 106 associated with scanning device 108. Although not shown, environment 106 may include multiple scanning devices, each of which may be associated with a corresponding VOI.

[0032] In some cases, each VOI 114 within the environment 106 can be defined relative to a corresponding scanning device (such as the scanning device 108 shown). For example, when the environment 106 is initially configured with sensors (such as scanning devices and cameras), each scanning device can store an association with a corresponding three-dimensional space of the environment 106. This three-dimensional space may include VOI 114, which may be adjacent to the corresponding scanning device 108, and may include scanning device 108, etc. In one example, VOI 114 corresponds to the three-dimensional space above the worktable surface where scanning device 108 is located. In another example, VOI 114 includes an area defined by the radius around the scanning device. Therefore, VOI 114 may include a three-dimensional region of arbitrary shape (such as a sphere, cube, etc.). Furthermore, after storing the initial VOI associated with the scanner, the VOI can be continuously adjusted over time based on interactions occurring within the environment. In addition, while the examples above describe pre-associating VOIs with scanners, in other cases, VOIs can be dynamically determined based on interactions occurring within the environment 106.

[0033] As shown, environment 106 may also include one or more cameras (such as camera 116) for generating image data that can be used to identify the user receiving the object, such as identifying a second user 110 receiving the object 104. Environment 106 may include multiple cameras (such as overhead cameras, shelf cameras, and / or such cameras) configured to acquire image data of different and / or overlapping portions of environment 106. Additionally, the association between each VOI and one or more cameras having a corresponding field of view (FOV) can be stored, which FOVs include some or all of the VOIs. For example, given that the FOV of camera 116 includes VOI 114 in environment 106, the association between the illustrated VOI 114 and the illustrated camera 116 can be stored.

[0034] In response to scanning data generated by scanning device 108, scanning device 108 may transmit the scanning data to one or more server computing devices 118 via one or more networks 120. Network 102 may represent any combination of one or more wired networks and / or wireless networks. Meanwhile, server computing devices 118 may reside in the environment, be remote from the environment, and / or a combination thereof. As shown, server computing devices 118 may include one or more processors 122 and memory 124, which may partially store positioning components 126, image analysis components 128, event determination components 130, and virtual shopping cart components 132. In addition, memory 124 may store sensor data 134 received from one or more sensors in the environment (e.g., scanning devices, cameras, etc.), user data 136 indicating the location of a user identifier in the environment 106, virtual shopping cart data (or "cart data") 138 indicating the contents of the virtual shopping cart of the corresponding user, and environmental data 140 indicating information about sensors and other information (such as location) within the environment, in one or more data storage areas.

[0035] In response to receiving scan data from scanning device 108, server computing device 118 may store the scan data in sensor data storage area 134. In response to receiving the scan data, or in response to another triggering event (such as detecting a second user 110 or another user leaving environment 106), event determination component 130 may attempt to determine the outcome of any event involving the scanned object 104. For example, the event determination component may initiate a process for determining the identity of the user who received the object in order to update shopping cart data associated with the identified user, thereby indicating the addition of the object.

[0036] To determine the outcome of an event involving the scanned object 104, the event determination unit 130 may instruct the image analysis unit 128 to analyze image data generated by one or more cameras in the environment that have a corresponding VOI in their FOV to identify the user receiving the object 104. Furthermore, the event determination unit may utilize the positioning unit 126 (which stores the current and past positions of the user identifier in the environment 106 over time) and combine it with the output from the image analysis unit 128 to determine the outcome of the event involving the object 104. After determining the outcome, the event determination unit 130 may instruct the virtual shopping cart unit 132 to update the corresponding virtual shopping cart accordingly.

[0037] First, the image analysis unit 128 may receive instructions to analyze image data of a specific VOI within a specific time range. For example, the event determination unit 130 may determine the location and time of a scan from scan data. For example, scan data received from the scanning device 108 may include the scan time and the identifier of the scanning device 108. The event determination unit may provide this information to the image analysis unit 128, or may otherwise use this information to enable the image analysis unit to analyze a corresponding spatiotemporal window, as described below.

[0038] Upon receiving a request from the event determination unit 130, the image analysis unit 128 can determine, via the environmental data 140, which camera(s) includes the FOV of VOI 114. That is, the environmental data 140 may store corresponding indications about which cameras have views of which VOIs, or may otherwise store indications about which camera will be used to determine an event occurring within a particular VOI. In this case, the image analysis unit 128 can determine that the illustrated camera 116 has an FOV including VOI 114. Therefore, the image analysis unit 128 can retrieve image data from the sensor data storage area 134 to run one or more computer vision algorithms on the image data. In some cases, the image analysis unit 128 analyzes image data of VOI 114 within a time range at least partially based on the scan time. For example, the image analysis unit 128 can analyze image data starting at the scan time and continuing for thirty seconds thereafter, starting fifteen seconds before the scan and ending one minute later, and / or such time ranges.

[0039] After retrieving image data representing the current VOI 114 for a defined time range, the image analysis unit 128 can use one or more trained classifiers to determine events occurring within VOI 114. For example, the trained classifiers may first be configured to receive feature data generated for each frame of the image data, and then output a segmentation map frame-by-frame indicating predefined objects represented within the corresponding frame. For example, the image analysis unit 128 may utilize a classifier that has been trained (e.g., via supervised learning) to identify background, user, specific parts of the user (e.g., hands, head, arms, body, etc.), one or more objects, and / or similar objects. The following... Figure 3A An exemplary segmentation graph is shown.

[0040] In addition to generating segmentation maps for each frame of the image data, the image analysis unit 128 may also utilize one or more trained classifiers configured to identify events occurring within VOI 114 using at least the segmentation maps. For example, the classifiers may be configured to determine whether each frame of the image data includes a hand, and if so, the state of the hand, such as "empty" (not holding an object) or "full" (holding an object). In some cases, the trained classifier may receive feature data generated from each frame of the image data and output a score indicating whether each frame includes a hand, the location of any such hand, and a score indicating whether the hand is empty or full. One or more thresholds may be applied to these scores to determine whether each individual frame includes a hand, and if so, whether the hand is empty or full.

[0041] In addition to storing this information for each frame, the image analysis unit 128 can also determine the motion vector of the identified hand within the VOI 114 over time. For example, if the image analysis unit 128 determines that an empty hand is detected at a first position in a first frame, and an empty hand is detected at a second position in a second and subsequent frame, the image analysis unit can use this information to determine the motion vector associated with that hand. Furthermore, the image analysis unit 128 may include one or more trained classifiers configured to determine, at least in part, whether a user has received or returned an object based on these motion vectors and associated information about the hand's state. For example, the image analysis unit 128 may have been trained (e.g., using supervised learning) to determine that an empty hand “entering” the VOI 114 and a full hand “leaving” the VOI 114 indicate that the user associated with that hand has taken an object. Therefore, in this example, the image analysis unit 114 may output an instruction to “take” or “pick up”. In the opposite example, a classifier can be trained to determine that a full hand entering VOI 114 and an empty hand leaving VOI 114 can represent a return.

[0042] In the example shown, image analysis component 128 can determine (e.g., using one or more trained classifiers) that an empty hand enters VOI 114 after the scan time of object 104, and a full hand leaves VOI 114. Image analysis component 128 can provide this information and / or related information (e.g., an indication of "taking") to event determination component 130. Furthermore, positioning component 126 can locate the user identifier of the corresponding user as they move throughout the environment, and can store these locations as user data 136 over time. For example, when the second user 110 shown enters environment 106, positioning component 126 may have created an identifier associated with that user, and may have stored the user's location associated with that user identifier over time. In some cases, the user identifier may not contain personally identifiable information, making it impossible to track the actual identifier of user 110, but instead tracking an identifier that has no other identifiable connection to user 110.

[0043] In addition to the data received from the image analysis unit 128, the event determination unit 130 may also use the user data 136 to determine the user identifier associated with the user 110 who acquired the object 104 (i.e., to determine the user 110 associated with the hand identified in VOI 114). For example, the image analysis unit 128 (or the event determination unit 128) may have determined that the "acquisition" of the object 104 occurred at a specific time (e.g., 10:23:55). The positioning unit 126 (or the event determination unit 130) may determine which user identifier was at the location of VOI 114 at that specific time, and may use this information to determine that the user associated with that user identifier acquired the object 104. In response to this determination, the event determination unit 130 may instruct the virtual shopping cart unit 132 to update the virtual shopping cart data 138 associated with the specific user identifier (e.g., the user identifier associated with user 110). By using the above technology, in response to user 110 receiving item 104 directly from user 102 (e.g., in the form of handover), in response to user 110 taking item 104 from the counter after user 102 places item 104 on the counter, and / or similar events, the event determining component may thereby instruct the virtual shopping cart component to update the corresponding virtual shopping cart data 138.

[0044] Figures 2A to 2C An exemplary operation sequence 200 is illustrated together, wherein the above Figure 1 The first user 102 scans the object 104 and hands it over to the second user 110. At the same time, the server computing device 118 analyzes the image data of the environment near the interaction activity during the scan to determine the user identifier associated with the second user 110 who received the object.

[0045] First, a second user 110 or another user may request a specific item, such as a certain quantity of food, a certain length of fabric, a salad with specific garnishes, and / or similar items. In response, the first user may prepare the customized item and may print or otherwise generate a physical or digital visual tag associated with the item, such as a barcode, QR code, etc. In some cases, the visual tag may be coded information about the item, such as an item identifier, the item's weight, the item's length, the item's quantity, the item's cost, the time the item was ordered, the time the visual tag was created, and / or similar information. When the visual tag is a physical visual tag, the first user 102 may affix the visual tag to the item, or when the visual tag is a digital or physical visual tag, the first user may otherwise associate the visual tag with the item.

[0046] After affixing a visual marker or otherwise associating it with object 104, operation 202 indicates that object 104 should be scanned using a scanning device to generate scan data. For example, the first user 102 can use any type of scanning device to scan the visual marker to generate scan data. As described above, a system that is to receive instructions for scanning data can store the association between a specific scanning device and its location within the environment. For example, the system can store the association between a specific VOI and each specific scanning device.

[0047] Operation 204 indicates that scan data is sent to one or more computing devices, such as the server computing device 118 mentioned above. In some cases, the scan data includes or is accompanied by additional information, such as the identifier of the scanning device, the scan time, etc.

[0048] Operation 206 indicates the generation of image data using a camera in the environment. It can be understood that, in some cases, the camera may continuously generate this image data for the purpose of identifying events including those involving objects scanned at operation 202, as well as other events.

[0049] Figure 2B The illustration continues with operation sequence 200 and includes sending the generated image data to one or more computing devices, such as server computing device 118, at operation 208. In some cases, the camera continuously sends the image data to the computing device, which can analyze the image data to identify events occurring within the environment.

[0050] Operation 210 indicates that the computing device receiving scan data determines the time associated with the scan and the VOI associated with the scanning device using the scan data. For example, this operation may include reading the timestamp of the scan from the scan data and determining the VOI associated with the scanning device stored in facility location data by using the identifier of the scanning device as a key. The computing device can then use the scan time and VOI to define a spatiotemporal window. For example, as described above, the spatiotemporal window may include the VOI in relation to the spatial portion of the window and a time range for a predefined amount of time (e.g., 10 seconds, 30 seconds, 2 minutes, etc.) in relation to the temporal portion of the window. This spatiotemporal window can be used to analyze image data to identify the hand of a user who acquired the object after the scan.

[0051] Operation 212 represents the analysis of VOI during a threshold amount of time following the scan time through image data analysis. For example, this operation may include generating feature data associated with the image data and feeding the feature data into a trained classifier configured to recognize the user's hand and the state of the user's hand (e.g., full or empty).

[0052] Figure 2C The illustration of operation sequence 200 is summarized, and includes, at operation 214, identifying a user's hand within a threshold time interval from the scan time within the VOI. In some cases, this operation includes an indication of hand recognition in a time-stamped image data frame, the timestamp falling within a defined time range of a spatiotemporal window. In other cases, the operation may include identifying at least one frame within the spatiotemporal window (where the classifier identifies an empty hand) and at least one subsequent frame within the spatiotemporal window (where the classifier identifies a full hand). In other examples, the operation may include generating motion vectors for the identified hands across frames, and identifying an empty hand moving towards and entering the VOI, and a full hand moving away from and leaving the VOI.

[0053] Operation 216 represents determining a user identifier associated with a user, which is associated with the identified hand. In some cases, this operation may include accessing user data generated by the positioning component to determine which user identifier is near or located at the scanning device and / or VOI during scanning or within a time range defined by the scan.

[0054] Operation 218 indicates updating the virtual cart data associated with the user identifier to indicate the addition of an item. For example, in the example shown, the user's virtual cart is updated to include "1.2 pound shrimp" added for $12.34. While this example describes adding an item identifier to the virtual cart, the item identifier can be removed from the virtual cart accordingly if the user returns a scanned item. In these examples, the spatiotemporal window triggered by the item scan can be defined before and after the scan, and image data generated before the scan can be analyzed to identify returns, such as identifying a full hand entering the VOI and an empty hand leaving the VOI.

[0055] Figure 3A Image analysis component 128 is shown, which uses machine vision algorithms to generate a segmentation map 300 of an image data frame. As shown, the segmentation map 300 can indicate pixels associated with different objects (such as background, user's hand, objects, etc.). For example, the segmentation map 300 shown indicates that different regions of the image data frame have been associated with exemplary semantic labels (e.g., "tags") using one or more trained classifiers 302, such as those of image analysis component 128. In this example, semantic labels include background 304, head 306, body 308, arm 310, hand 312, object (or object in hand) 314, and door 316. Of course, it should be understood that these are merely examples, and any other type of semantic label can be used. It should also be noted that the classifier used to generate this exemplary segmentation map 302 can be trained by employing a human user to assign the corresponding semantic labels 304 to 316 to different regions of the frame using computer graphics tools. After one or more human users have assigned these semantic labels to a threshold number of image data, the trainable classifier applies the semantic labels to more additional image data.

[0056] In some cases, the first trained classifier of classifier 302 outputs segmentation maps 300 frame-by-frame indicating certain body parts (e.g., hands, forearms, upper arms, heads, etc.) and the locations of these corresponding parts. In some cases, in addition to the outline of candidate hands, the first trained classifier may also output a score indicating the probability that a specific portion of an image data frame represents a hand. The first classifier may also associate each identified hand in the image data with an identified head. This head can be used to determine the user's user identifier, thus each hand can be associated with a user identifier.

[0057] Figure 3BAn image analysis component 128 is shown that uses a segmentation map 300 and one or more trained classifiers 302 to determine, for each frame of image data, whether the frame includes a hand, and if so, the location of the hand and whether it is empty or full. The image analysis component 128 can use this information to identify the user who acquired the object after scanning it with a scanning device. For example, the frame shown in this example may correspond to a timestamped image data frame of VOI 114 within a predefined time range of a spatiotemporal window.

[0058] As shown in the figure, the image analysis component analyzes the first frame 318(1). In this example, the first classifier did not identify a hand. However, the first classifier did identify an empty hand in the subsequent frame 318(2). For example, as described above, the classifier can output a score indicating the presence of a hand, a score indicating whether the identified hand is empty or full, and the position of the identified hand. Furthermore, the classifier can also associate the hand with a head, and can also associate it with a user identifier associated with that head.

[0059] The classifier determines that the third exemplary frame 318(3) represents another empty hand, while the classifier determines that the fourth frame 318(4) and the fifth frame 318(5) represent full hands, respectively. Furthermore, the image analysis unit 128 can use the corresponding user identifier and the position of the identified hand across frames to generate one or more motion vectors associated with the hand. For example, the image analysis unit can use the user identifier associated with the hand to identify the movement of the same hand over time. As shown, in this example, the image analysis unit 128 can identify motion vectors indicating that an empty hand moves into the VOI and a full hand moves away from and / or out of the VOI. Classifier 302 or another classifier can use this information to determine that the user associated with the hand actually picked up the scanned object.

[0060] Figures 4A to 4BA flowchart of an exemplary process 400 is shown, which is used to update virtual shopping cart data of a user who receives an object scanned by another user using a scanning device. Process 400 and other processes described herein may be implemented in hardware, software, or a combination thereof. In a software context, the described operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more hardware processors, perform the operations. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular abstract data type. Those skilled in the art will readily recognize that some steps or operations shown in the above figures may be eliminated, combined, or performed in an alternative order. Any steps or operations may be performed sequentially or in parallel. Furthermore, the order in which operations are described is not intended to be construed as limiting. Additionally, these processes may be performed by baskets (e.g., shopping carts, baskets, bags, etc.), servers, other computing devices, or a combination thereof.

[0061] Operation 402 indicates that an indication has been received that the scanner has generated scan data. For example, the event determination component 130 may receive an indication that a specific scanning device in the environment has scanned a visual mark (such as a barcode, QR code, etc.) associated with an object.

[0062] At operation 404, event determining component 130 or another component may determine an object identifier associated with the scanned object. For example, event determining component 130 or another component may use the scan data to identify the barcode of the scanned object, etc. In some cases, this operation may occur in response to receiving scan data, while in other cases, it may occur in response to different triggering factors, such as a user designated as a candidate user for an event associated with the scan data leaving the environment. Therefore, in some cases, this operation may occur after the subsequent operations described below.

[0063] Operation 406 indicates determining a first time associated with the scanner that generated the scan data. This first time may include the time when the scanner generated the scan data, the time when the scanner sent the scan data, the time when the computing device received the scan data, and / or such times.

[0064] At operation 408, image analysis component 128 or another component may determine the volume of interest (VOI) associated with the scanning device. For example, as described above, the system may store the association between each scanning device within the facility and the corresponding VOI (e.g., the three-dimensional space of the facility). The image analysis component may determine an identifier for the scanning device, which may be included in or accompany the scan data, and this identifier may be used to determine the corresponding VOI.

[0065] Operation 410 represents the analysis of the first frame of image data including the VOI. For example, this operation may include image analysis unit 128 or another unit identifying a camera whose FOV includes the VOI (e.g., by accessing a data storage area storing the association between the respective camera and the VOI), and receiving image data generated from that camera and close to the first time determined at operation 406. For example, image analysis unit 128 may receive image data generated by the camera during a time range that begins at, before, or immediately after the first time.

[0066] Additionally, this operation may include analyzing a portion of image data corresponding to the VOI to determine whether that portion of the image data includes a hand. For example, operation 410 may include sub-operation 410(1), where feature data generated from the image data is input to a classifier. At sub-operation 410(2), the classifier may output a first score indicating whether the image data includes a hand and a second score indicating whether the hand is empty or full. In some cases, analyzing the image data includes the two-step process described above: first, segmenting each frame of the image data into different predefined objects, including the user's hand and head; second, tracking the movement of any identified hand across frames. For example, a first classifier may generate a segmentation map identifying at least one hand and a corresponding head, where the latter can be used to determine the user's user identifier. This segmentation information may be output by the first classifier and input into, for example, a hand tracking component for tracking each hand across frames.

[0067] At operation 412, in this example, image analysis component 128 or another component identifies an empty hand of the user within the VOI. For example, image analysis component 128 or another component uses one or more classifiers to determine that an empty hand exists within the VOI after a first time associated with the scan data.

[0068] Operation 414 indicates that image analysis component 128 or another component analyzes a second frame of image data including VOI, wherein the second frame corresponds to a time following the time associated with the first frame. Similarly, this operation may include sub-operation 414(1), wherein feature data generated from the second frame of image data is input to a classifier. At sub-operation 414(2), the classifier may output a third score indicating whether the image data includes a hand and a fourth score indicating whether the hand is empty or full.

[0069] Figure 4BThe illustration continues with process 400, and includes identifying a user's full hand at operation 416 based on analysis of the second frame at operation 414. For example, this operation may include receiving an indication from the classifier of a full hand within the VOI in the second frame of image data. In some cases, this indication may also indicate that the full hand and the empty hand identified at operation 412 are associated with the same user identifier.

[0070] Finally, operation 418 indicates that the item identifier associated with the item is stored in the virtual shopping cart data associated with the user, who is linked to the user identifier. For example, virtual shopping cart component 132 or another component can update the corresponding user's shopping cart to indicate the addition of the item.

[0071] In some cases, operation 418 may occur in response to determining that a user's hand enters the VOI empty and leaves the VOI full. Meanwhile, in other cases, operation 418 may occur in response to recognizing a user's hand in the VOI within a time range defined by a first time, or in response to recognizing a user's full hand in the VOI within that time range. For example, in some cases, if a single hand is recognized within that time range, the user's virtual shopping cart may be updated in response to recognizing the user's hand. However, if multiple hands (corresponding to different user identifiers) are recognized within the VOI within that time range, the virtual shopping cart for that specific user may be updated in response to recognizing a user's empty hand, and subsequently recognizing that user's full hand.

[0072] Figure 5 A flowchart of another exemplary process 500 is shown, which is used to update the virtual shopping cart data of a user receiving an object. Operation 502 represents receiving sensor data generated by sensors in the environment that identify the object. In some cases, the sensors may include scanning devices that generate scan data in response to scanning visual markers associated with the object.

[0073] Operation 504 indicates determining a portion of the environment associated with the sensor. As described above, determining this portion of the environment may include determining a volume of interest (VOI) within the environment relative to the sensor. In some cases, the data storage area of ​​the described system may store associations between the scanning device and the VOI, associations between the scanning device and a camera with the FOV of the VOI, and / or such associations, etc.

[0074] Operation 506 represents receiving image data generated by a camera within the environment, representing a portion of the environment associated with the sensor. For example, this operation may include receiving image data from a camera with a field of view (FOV) having a VOI. In some cases, scan data may indicate a first time associated with sensor data (e.g., the time when the sensor generated the sensor data); therefore, receiving image data may include receiving image data from the corresponding camera within a time range based on the first time.

[0075] Operation 508 represents analyzing image data to identify a user receiving an object. In some cases, this operation may include analyzing image data generated by the camera after a first time and within a threshold time amount of the first time. Additionally, the analysis may include analyzing at least a portion of the image data corresponding to a VOI to identify the user's hand within the VOI. This may further include analyzing the image data to identify an empty hand of the user within the VOI at least through a first frame of the image data, and a full hand of the user within the VOI through a second frame of the image data. In other cases, the analysis may include: analyzing at least a portion of the first frame of the image data corresponding to the VOI to identify an empty hand of the user at a first position within the VOI; analyzing at least a portion of the second frame of the image data corresponding to the VOI to identify an empty hand of the user at a second position within the VOI; determining a first direction vector at least partially based on the first and second positions; analyzing at least a portion of the third frame of the image data corresponding to the VOI to identify a full hand of the user at a third position within the VOI; analyzing at least a portion of the fourth frame of the image data corresponding to the VOI to identify a full hand of the user at a fourth position within the VOI; and determining a second direction vector at least partially based on the third and fourth positions. In other words, in some cases, analyzing image data to identify the user receiving the object may include determining that the user's empty hand enters the VOI and that the user's full hand leaves the VOI.

[0076] In addition, as described above, analyzing image data may include generating a segmentation map for identifying one or more hands within the VOI, and using this information to determine the user who received the object. For example, this operation may include: generating a segmentation map using a first frame of image data, the segmentation map identifying at least a first set of pixels in the first frame corresponding to the user's hand; inputting first data indicating the first set of pixels in the first frame corresponding to the user's hand into a trained classifier; and receiving second data, as the output of the trained classifier, indicating whether the user received the object. Furthermore, in some cases, objects may be identified and tracked within the VOI in addition to the user's hand. In these cases, objects may be identified within image data frames and tracked across frames to identify objects placed in the user's hand. In each of the examples described herein, the user's hand may receive the object from another user (e.g., a facility employee), from a counter where another user placed the object, and / or in any other way.

[0077] Finally, operation 510 indicates updating the virtual shopping cart data associated with the user to indicate the item identifier associated with the item. This operation may include adding information about the item to the corresponding virtual shopping cart, such as the item identifier, the item's cost, the item's description, the time the user received the item, and / or similar information.

[0078] Figure 6 This is a block diagram of an exemplary material handling facility site 602, including sensors and an inventory management system configured to generate outputs about events occurring within the facility site using sensor data. In some cases, facility site 602 corresponds to the architecture 100 and / or environment 106 described above.

[0079] However, the following description is merely an illustrative example of the industries and environments in which the techniques described herein may be used. A material handling facility site 602 (or “facility site”) includes one or more physical structures or areas that can accommodate one or more objects 604(1), 604(2), ..., 604(Q) (generally denoted as 604). As used in this disclosure, letters in parentheses such as “(Q)” indicate integer results. Objects 604 include physical goods, such as books, pharmaceuticals, repair parts, electronic equipment, groceries, etc.

[0080] Facility location 602 may include one or more areas designated for different functions of inventory handling. In this illustration, facility location 602 includes a receiving area 606, a storage area 608, and a transition area 610. Receiving area 606 may be configured to receive items 604 from suppliers for inclusion in facility location 602. For example, receiving area 606 may include loading docks where trucks or other cargo transport vehicles unload items 604.

[0081] Storage area 608 is configured to store items 604. Storage area 608 can be arranged in various physical configurations. In one embodiment, storage area 608 may include one or more aisles 612. Aisles 612 may be configured to have storage locations 614 on one or both sides of the aisle, or the aisle may be defined by storage locations. Storage locations 614 may include one or more of shelves, racks, boxes, cabinets, crates, floor locations, or other suitable storage structures for storing or preserving items 604. Storage locations 614 may be fixed to the floor or another part of the facility site structure, or they may be movable, allowing the arrangement of aisles 612 to be reconfigured. In some embodiments, storage locations 614 may be configured to move independently of external operators. For example, storage locations 614 may include racks with power supplies and motors that can be operated by computing devices to allow the racks to be moved from one location within the facility site 602 to another.

[0082] One or more users 616(1), 616(2), ..., 616(U), baskets 618(1), 618(2), ..., 618(T) (generally referred to as 618) or other material handling devices may move within facility site 602. For example, user 616 may move around within facility site 602 to pick or place items 604 at various inventory locations 614, placing these items on baskets 618 for transport. A single basket 618 is configured to carry or otherwise transport one or more items 604. For example, basket 618 may include baskets, shopping carts, bags, etc. In other embodiments, other mechanisms (such as robots, forklifts, cranes, drones, etc.) may move around within facility site 602 to pick, place, or otherwise move items 604.

[0083] One or more sensors 620 may be configured to acquire information within facility premises 602. Sensors 620 in facility premises 602 may include sensors fixed to the environment (e.g., a camera mounted on the ceiling) or other forms of sensors, such as user-owned sensors (e.g., mobile phones, tablets, etc.). Sensors 620 may include, but are not limited to, cameras 620(1), weight sensors, radio frequency (RF) receivers, temperature sensors, humidity sensors, vibration sensors, etc. Sensors 620 may be fixed or mobile relative to facility premises 602. For example, inventory location 614 may include a camera 620(1) configured to acquire images of items 604 being picked up or placed on shelves, images of users 616(1) and 616(2) in facility premises 602, etc. In another example, the floor of facility premises 602 may include weight sensors configured to determine the weight of users 616 or other objects on the floor.

[0084] During operation of facility 602, sensor 620 may be configured to provide information suitable for tracking the movement of objects or other events within facility 602. For example, a series of images acquired by camera 620(1) may instruct one of users 616 to remove item 604 from a specific inventory location 614 and place item 604 on or at least partially inside one of baskets 618.

[0085] Although storage area 608 is depicted as having one or more aisles 612, storage locations 614 for stored items 604, sensors 620, etc., it should be understood that receiving area 606, transition area 610, or other areas of facility site 602 may be similarly configured. Furthermore, the arrangement of the various areas within facility site 602 is depicted functionally rather than graphically. For example, multiple different receiving areas 606, storage areas 608, and transition areas 610 may be interspersed within facility site 602 rather than isolated from each other.

[0086] Facility 602 may include or be coupled to inventory management system 622, which can perform the above references Figures 1 to 5 Some or all of the technologies described herein. As described below, the inventory management system 622 may include those shown in claim 1 and referenced above. Figures 1 to 5 The components of server 118. For example, the inventory management system can maintain a virtual shopping cart for each user within the facility. The inventory management system can also store records associated with each user, indicating the user's identifier, location, and whether the user is eligible to leave the facility with items without manual checkout. The inventory management system can also generate notification data and output it to the user to indicate whether the user is eligible.

[0087] As shown in the figure, the inventory management system 622 may reside at facility location 602 (e.g., as part of a local server), on server 118 located remotely from facility location 602, or a combination thereof. In each case, the inventory management system 622 is configured to recognize interactions and events with user 616 and devices (such as sensors 620, robots, material handling equipment, computing devices, etc.) in one or more of the receiving area 606, storage area 608, or transition area 610, as well as interactions and events between user and devices. As described above, some interactions may also indicate the presence of one or more events 624, or predefined related activities. For example, events 624 may include user 616 entering facility location 602, storing item 604 in inventory location 614, picking item 604 from inventory location 614, returning item 604 to inventory location 614, placing item 604 in basket 618, movement of user 616 relative to each other, gestures of user 616, etc. Other events 624 involving user 616 may include user 616 providing authentication information in facility location 602, using computing devices at facility location 602 to authenticate an identifier to inventory management system 622, etc. Some events 624 may involve one or more other objects within facility location 602. For example, event 624 may include the movement of inventory location 614 (such as a counter mounted on wheels) within facility location 602. Event 624 may involve one or more sensors 620. For example, a change in the operation of sensor 620 (such as sensor malfunction, calibration change, etc.) may be designated as event 624. Continuing with this example, a change in the orientation of field of view 628 caused by movement of camera 620(1) (such as due to someone or something colliding with camera 620(1)) (e.g., camera 104) may be designated as event 624.

[0088] By determining that one or more of events 624 have occurred, the inventory management system 622 can generate output data 626. Output data 626 includes information about the events 624. For example, if event 624 includes the removal of item 604 from inventory location 614, output data 626 may include an item identifier indicating the specific item 604 removed from inventory location 614 and a user identifier of the user who removed the item.

[0089] The inventory management system 622 may use one or more automated systems to generate output data 626. For example, artificial neural networks, one or more classifiers, or other automated machine learning techniques may be used to process sensor data from one or more sensors 620 to generate output data 626. For instance, the inventory management system may perform some or all of these techniques to generate and utilize a classifier to identify user activity in image data, as described in detail above. The automated systems may operate using probabilistic or non-probabilistic techniques. For example, the automated system may use a Bayesian network. In another example, the automated system may use a support vector machine to generate output data 626 or a provisional result. The automated system may generate confidence data that provides information indicating the accuracy or confidence level of the output data 626 or provisional data in corresponding to the physical world.

[0090] Various techniques can be used to generate confidence data, at least in part, depending on the type of automated system in use. For example, a probabilistic system using a Bayesian network can use the probability assigned to the output as a confidence level. Continuing this example, a Bayesian network can indicate that there is a 95% probability that an object depicted in image data corresponds to an object previously stored in memory. This probability can be used as a confidence level for the object depicted in the image data.

[0091] In another example, the output from a non-probabilistic technique (such as a support vector machine) may have a confidence level based on distances in a mathematical space in which image data of the object has been classified against previously stored images of the object. In this space, the greater the distance between a reference point (such as a previously stored image) and the image data acquired during the event, the lower the confidence level.

[0092] In another example, image data of an object (such as object 604, user 616, etc.) can be compared with a set of previously stored images. Differences between the image data and the previously stored images can be evaluated. These differences could include, for example, differences in shape, color, relative proportions between features in the images, etc. These differences can be represented in mathematical space as distances. For example, the color of an object depicted in the image data and the color of an object depicted in a previously stored image can be represented as coordinates in a color space.

[0093] The confidence level can be determined at least in part based on these differences. For example, user 616 may pick up item 604(1) from inventory location 614, such as a perfume bottle, which is typically cubic in shape. Other items 604 in nearby inventory locations 614 may be primarily spherical. Based on the shape differences between adjacent items (cube versus sphere) and the shape correspondence with previously stored images of perfume bottle items 604(1) (cube versus cube), user 106 has a high confidence level in picking up perfume bottle item 604(1).

[0094] In some cases, automation technology may fail to generate output data 626 with a confidence level higher than a threshold. For example, automation technology may fail to distinguish which user 616 in a group of users 616 picked up item 604 from inventory location 614. In other cases, manual verification of the accuracy of event 624 or output data 626 may be required. For example, some items 604 may be considered age-restricted, meaning these items can only be handled by users 616 who are over a minimum age threshold.

[0095] In cases requiring manual verification, sensor data associated with event 624 may be processed to generate query data. Query data may include a subset of sensor data associated with event 624. Query data may also include one or more provisional results determined by automation techniques, or supplementary data. A subset of sensor data may be determined using information about one or more sensors 620. For example, camera data (such as the location of camera 620(1) within facility site 602, the orientation of camera 620(1), and the field of view 628 of camera 620(1)) may be used to determine whether a particular location within facility site 602 is within field of view 628. The subset of sensor data may include images that can display inventory location 614 or stored item 604. The subset of sensor data may also ignore images from other cameras 620(1) that do not have inventory location 614 within field of view 628. Field of view 628 may include a portion of the scene in facility site 602 about which sensors 620 are able to generate sensor data.

[0096] Continuing this example, a subset of the sensor data may include video clips acquired by one or more cameras 620(1) having a field of view 628 that includes object 604. Provisional results may include “best guesses” about which objects 604 are likely involved in event 624. For example, provisional results may include results determined by an automated system with a confidence level above a minimum threshold.

[0097] Facility location 602 can be configured to receive different kinds of items 604 from various suppliers and store these items until a customer orders or picks up one or more items 604. The arrows in Figure 2 indicate the general flow of items 604 throughout facility location 602. Specifically, as shown in this example, items 604 can be received at receiving area 606 from one or more suppliers (such as manufacturers, distributors, wholesalers, etc.). In various specific implementations, items 604 may include merchandise, daily necessities, perishable food, or any suitable type of item 604, depending on the nature of the business operating facility location 602. Receiving items 604 may include one or more events 624 for which inventory management system 622 can generate output data 626.

[0098] After receiving item 604 from the supplier at receiving area 606, these items can be prepared for storage. For example, item 604 can be unpacked or otherwise rearranged. Inventory management system 622 may include one or more software applications that execute on a computer system to provide inventory management functions based on events 624 associated with unpacking or rearrangement. These inventory management functions may include maintaining information on indication type, quantity, condition, cost, location, weight, or any other appropriate parameters about item 604. Item 604 may be stored, managed, or allocated in count, individual unit, or multiple (such as parcels, cartons, crates, pallets, or other suitable aggregates). Alternatively, some items 604 (such as bulk products, daily necessities, etc.) may be stored in continuous quantities or arbitrarily divisible quantities that may not themselves be organized into countable units. Such objects 604 can be managed according to measurable quantities, such as units of length, area, volume, weight, time, shelf life, or other dimensional attributes characterized by units of measurement. Generally speaking, the quantity of object 604 can refer to the countable number of individual or aggregate objects 604, or the measurable quantity of object 604, depending on the specific circumstances.

[0099] After item 604 arrives through receiving area 606, it can be stored in storage area 608. In some embodiments, similar items 604 may be stored or displayed together in storage location 614, such as in boxes, on shelves, or hanging on mounting boards. In this embodiment, all items 604 of a given type are stored in one storage location 614. In other embodiments, similar items 604 may be stored in different storage locations 614. For example, to optimize the retrieval of certain frequently turning items 604 within a large physical facility 602, these items 604 may be stored in several different storage locations 614 to reduce congestion that may occur in a single storage location 614. The storage of items 604 and their corresponding storage locations 614 may include one or more events 624.

[0100] When a customer order specifying one or more items 604 is received, or when a user 616 moves through facility 602, the corresponding items 604 can be selected or “picked” from inventory locations 614 containing those items 604. In various embodiments, item picking can range from manual picking to fully automated picking. For example, in one embodiment, a user 616 may have a list of desired items 604 and can walk through facility 602 to pick items 604 from inventory locations 614 within storage area 608, then place these items 604 into baskets 618. In other embodiments, employees of facility 602 may use a written or electronic picking list derived from a customer order to pick items 604. As employees move through facility 602, the selected items 604 can be placed into baskets 618. The selection may include one or more events 624, such as user 616 moving to inventory location 614, taking an item 604 from inventory location 614, etc.

[0101] After item 604 has been selected, it can be processed in transition area 610. Transition area 610 can be any designated area within facility site 602, where item 604 is transitioned from one location to another or from one entity to another. For example, transition area 610 can be a packaging station within facility site 602. When item 604 arrives at transition area 610, item 604 can be transitioned from storage area 608 to packaging station. Transitions may include one or more events 624. Inventory management system 622 can use output data 626 associated with these events 624 to maintain information about the transitions.

[0102] In another example, if item 604 leaves facility 602, inventory management system 622 can obtain and use a list of item 604 to transfer responsibility or custody of item 604 from facility 602 to another entity. For example, a carrier can accept item 604 for transport, and then assume responsibility for item 604 indicated in the list. In another example, a customer can purchase or rent item 604 and remove item 604 from facility 602. Purchase or rental may include one or more events 624.

[0103] The inventory management system 622 can access or generate sensor data about facility location 602 and its contents, including objects 604, users 616, baskets 618, etc. Sensor data may be acquired by one or more sensors 620, provided by other systems, etc. For example, sensor 620 may include a camera 620 (1) configured to acquire image data of a scene in facility location 602. Image data may include still images, video, or a combination thereof. The inventory management system 622 can process the image data to determine the location of user 616, basket 618, identifier of user 616, etc. As used herein, a user identifier may represent a unique identifier of the user (e.g., name, number associated with the user, username, etc.), an identifier that distinguishes the user from other users in the environment, etc.

[0104] An inventory management system 622, or a system coupled thereto, may be configured to identify user 616 and determine other candidate users. In one embodiment, this determination may include comparing sensor data with previously stored identification data. For example, user 616 may be identified by presenting their face to a facial recognition system, by presenting a token carrying authentication credentials, by providing a fingerprint, by scanning a barcode or other type of unique identifier upon entering the facility premises. The identification of user 616 may be determined before, during, or after entering the facility premises 602. Determining the identification of user 616 may include comparing sensor data associated with user 616 in the facility premises 602 with previously stored user data.

[0105] In some cases, the inventory management system groups users within a facility into corresponding sessions. That is, the inventory management system 622 can use sensor data to identify groups of users who are effectively "together" (e.g., shopping together). In some cases, a particular session may include multiple users who enter the facility 602 together and are likely to explore the facility together. For example, when a family consisting of two adults and two children enters the facility together, the inventory management system can associate each user with a specific session. Given that users in a session may not only individually pick up or return or otherwise interact with items, but may also pass items back and forth between each other, identifying sessions beyond individual users can help determine the outcome of individual events. For example, in the example above, a child might first pick up a box of cereal and then hand it to her mother, who might then place the box of cereal in her basket 618. Noting that the child and mother belong to the same session increases the chances of successfully adding the box of cereal to the mother's virtual shopping cart.

[0106] By identifying the occurrence of one or more events 624 and their associated output data 626, the inventory management system 622 can provide one or more services to users 616 of facility 602. The overall accuracy of the system is improved by utilizing one or more human employees to process query data and generate response data, which can then be used to produce output data 626. This improved accuracy enhances the user experience for one or more users 616 of facility 602. In some examples, output data 626 may be transmitted over a network 630 to one or more servers 118.

[0107] Figure 7 A block diagram is shown of one or more servers 118 configured to support the operation of a facility site. Servers 118 may physically reside at facility site 602, be accessible via network 630, or a combination of both. Servers 118 do not require end users to know the actual location and configuration of the system providing the service. Common expressions associated with servers 118 may include “on-demand computing,” “Software as a Service (SaaS),” “cloud service,” “data center,” etc. The services provided by servers 118 may be distributed across one or more physical or virtual devices.

[0108] Server 118 may include one or more hardware processors 702 (processors) configured to execute one or more stored instructions. Processor 702 may include one or more cores. Server 118 may include one or more input / output (I / O) interfaces 704 to allow processor 702 or other parts of server 118 to communicate with other devices. I / O interface 704 may include internal integrated circuits (I2C), a Serial Peripheral Interface bus (SPI), or a Universal Serial Bus (USB) as published by the USB Developer Forum.

[0109] Server 118 may also include one or more communication interfaces 706. Communication interfaces 706 are configured to provide communication between server 118 and other devices, such as sensor 620, interface devices, routers, etc. Communication interfaces 706 may include devices configured to couple to personal area networks (PANs), wired LANs and wireless LANs (LANs), wired wide area networks and wireless wide area networks (WANs), etc. For example, communication interfaces 706 may include interfaces with Ethernet, Wi-Fi, etc. TM Compatible devices. Server 118 may also include one or more buses or other internal communication hardware or software that allow data transfer between the various modules and components of server 118.

[0110] Server 118 may also include power supply 740. Power supply 740 is configured to provide power suitable for operating the components in server 118.

[0111] Server 118 may also include one or more memories 710. Memory 710 includes one or more computer-readable storage media (CRSMs). CRSMs may be any one or more of electronic storage media, magnetic storage media, optical storage media, quantum storage media, mechanical computer storage media, etc. Memory 710 provides storage for computer-readable instructions, data structures, program modules, and other data for the operation of server 118. Some exemplary functional modules of storage are shown in memory 710, but the same functionality may alternatively be implemented using hardware, firmware, or a system-on-a-chip (SoC).

[0112] The memory 710 may include at least one operating system (OS) component 712. The OS component 712 is configured to manage hardware resource devices (such as I / O interface 704, communication interface 708) and provide various services to applications or components executing on the processor 702. The OS component 712 may implement FreeBSD, such as that released by the FreeBSD project. TM Operating system variants; other UNIX TM Or UNIX-like variants; such as Linux released by Linus Torvalds. TMVariants of the operating system; Microsoft Corporation of Redmond, Washington, USA Server operating system; etc.

[0113] One or more of the following components may also be stored in memory 710. These components can be executed as foreground applications, background tasks, daemons, etc. Communication component 714 can be configured to establish communication with one or more sensors 620, one or more devices used by employees, other servers 118, or other devices. Authentication, encryption, etc., can be performed on the communication.

[0114] The memory 710 can store the inventory management system 622. The inventory management system 622 is configured to provide the above reference. Figures 1 to 5 Some or all of the technologies described herein. For example, inventory management system 622 may include components for receiving scan data, determining events occurring for scanned items, updating a user's virtual shopping cart, etc.

[0115] The inventory management system 622 can access information stored in one or more data repositories 718 in the memory 710. The data repository 718 can use flat files, databases, linked lists, trees, executable code, scripts, or other data structures to store information. In some implementations, the data repository 718, or a portion thereof, can be distributed across one or more other devices (including other servers 118, network-attached storage devices, etc.). The data repository 718 may include the aforementioned data storage areas, such as user data 136, environmental data 140, sensor data 134, and shopping cart data 136.

[0116] The data repository 718 may also include physical layout data 720. Physical layout data 720 provides a mapping of the physical locations of devices and objects (such as sensor 620, inventory location 614, etc.) within the physical layout. Physical layout data 720 may indicate the coordinates of inventory location 614 within facility site 602, the sensor 620 within the field of view of inventory location 614, and so on. For example, physical layout data 720 may include camera data including one or more of the following: the location of camera 620(1) within facility site 602, the orientation of camera 620(1), operating status, etc. Continuing this example, physical layout data 720 may indicate the coordinates of camera 620(1), panning and tilt information indicating the direction of field of view 628 along its orientation, whether camera 620(1) is operational or malfunctioning, etc.

[0117] In some implementations, the inventory management system 622 may access physical layout data 720 to determine whether the location associated with event 624 is within the field of view 628 of one or more sensors 620. Continuing the example above, given the location of event 624 within facility site 602 and camera data, the inventory management system 622 may determine the camera 620 that may generate an image of event 624 (1).

[0118] Item data 722 includes information associated with item 604. This information may include information indicating one or more inventory locations 614 where one or more items of item 604 are stored. Item data 722 may also include order data, SKU or other product identifiers, price, current quantity, weight, expiration date, image of item 604, detailed description information, rating, ranking, etc. Inventory management system 622 may store information associated with inventory management functions in item data 722.

[0119] Data repository 718 may also include sensor data 134. Sensor data 134 includes information acquired from one or more sensors 620, or information based on those sensors. For example, sensor data 134 may include 3D information about objects in facility site 602. As described above, sensor 620 may include camera 620(1) configured to acquire one or more images. These images may be stored as image data 726. Image data 726 may include information describing multiple image elements or pixels. Non-image data 728 may include information from other sensors 620, such as input from microphones, weight sensors, etc.

[0120] User data 730 may also be stored in data repository 718. User data 730 may include identification data, information indicating a personal profile, purchase history, location data, an image of user 616, demographic data, etc. Individual user 616 or user group 616 may selectively provide user data 730 for use by inventory management system 622. Individual user 616 or user group 616 may also authorize the collection of user data 730 during use of facility location 602, or authorize access to user data 730 obtained from other systems. For example, user 616 may opt in to participate in the collection of user data 730 to receive enhanced services while using facility location 602.

[0121] In some implementations, user data 730 may include information specifying special handling for user 616. For example, user data 730 may indicate that a particular user 616 has been associated with an increasing number of errors in the output data 626. The inventory management system 622 may be configured to use this information to apply additional review to events 624 associated with that user 616. For example, an event 624 including an item 604 with a cost or result exceeding a threshold may be presented to an employee for processing, regardless of the confidence level determined in the output data 626 generated by the automated system.

[0122] The inventory management system 622 may include one or more components such as a positioning component 124, an identification component 734, an image analysis component 128, an event determination component 130, a virtual shopping cart component 132, a query component 738, and possibly other components 756.

[0123] The positioning component 124 is used to locate objects or users within the facility premises environment, allowing the inventory management system 622 to assign certain events to the correct users. Specifically, the positioning component 124 can assign a unique identifier to a user upon entering the facility premises and, with the user's consent, can locate the user throughout the facility premises 602 while the user is within the facility premises 602. The positioning component 124 can perform this positioning using sensor data 134 (such as image data 726). For example, the positioning component 124 can receive image data 726 and can use facial recognition technology to identify the user from the image. After identifying a specific user within the facility premises, the positioning component 124 can subsequently locate that user within the image as the user moves throughout the facility premises 602. Additionally, if the positioning component 124 temporarily "loses" a specific user, it can re-attempt to identify the user within the facility premises based on facial recognition and / or using other technologies (such as voice recognition, etc.).

[0124] Therefore, upon receiving indication of the time and location of the event in question, the positioning unit 124 can query the data repository 718 to determine which user(s) are at the event location or within a threshold distance of the event location at a specific time. Additionally, the positioning unit 124 can assign different confidence levels to different users, where these confidence levels indicate the likelihood that each corresponding user is indeed the user associated with the relevant event.

[0125] Positioning component 124 can access sensor data 134 to determine location data of a user and / or object. The location data provides information indicating the location of an object (such as object 604, user 616, basket 618, etc.). This location can be an absolute location relative to facility location 602 or a relative location relative to another object or reference point. Absolute items may include width, length, and height relative to a geodetic reference point. Relative items may include a location as specified in a plan of facility location 602, such as 25.4 meters (m) along the x-axis and 75.2 meters along the y-axis, 5.2 meters from inventory location 614 in a 169° direction, etc. For example, the location data may indicate that user 616(1) is 25.2 meters along aisle 612(1) and standing in front of inventory location 614. In contrast, the relative location may indicate that user 616(1) is 32 cm from basket 618 in a 73° direction relative to basket 618. The location data may include orientation information, such as which direction user 616 is facing. Orientation can be determined by the relative direction the user's body is facing. In some implementations, orientation can be relative to the interface device. Continuing with this example, the location data could indicate that the user 616(1) is oriented at 0° or looking north. In another example, the location data could indicate that the user 616 is facing the interface device.

[0126] The identification component 734 is configured to identify an object. In one embodiment, the identification component 734 may be configured to identify an object 604. In another embodiment, the identification component 734 may be configured to identify a user 616. For example, the identification component 734 may use facial recognition technology to process image data 726 and determine the identification data of the user 616 depicted in the image by comparing features in the image data 726 with previously stored results. The identification component 734 may also access data from other sensors 620, such as from RFID readers, RF receivers, fingerprint sensors, etc.

[0127] Event determination component 130 is configured to process sensor data 134 using the techniques described above and others, and generate output data 726. Event determination component 130 has access to information stored in data repository 718, including but not limited to event description data 742, confidence levels 744, or thresholds 746. In some cases, event determination component 130 may be configured to perform some or all of the techniques described above with respect to event determination component 106. For example, event determination component 130 may be configured to create and utilize an event classifier to identify events (e.g., predefined activities) within image data, potentially without using additional sensor data acquired by other sensors in the environment.

[0128] Event description data 742 includes information indicating one or more events 624. For example, event description data 742 may include predefined profiles that specify events 624 related to the movement of item 604 from inventory location 614 and "picking". Event description data 742 may be generated manually or automatically. Event description data 742 may include data indicating triggers associated with events occurring in facility location 602. An event can be determined to have occurred upon detection of a trigger. For example, sensor data 134 at inventory location 614 (such as weight changes from weight sensor 620(6)) may trigger the detection of an event in which item 604 is added to or removed from inventory location 614. In another example, a trigger may include an image of user 616 reaching for inventory location 614. In yet another example, a trigger may include two or more users 616 approaching each other within a threshold distance.

[0129] The event determination component 130 may use one or more techniques to process the sensor data 134, including but not limited to artificial neural networks, classifiers, decision trees, support vector machines, Bayesian networks, etc. For example, the event determination component 130 may use a decision tree to determine the occurrence of the “picking” event 624 based on the sensor data 134. The event determination component 130 may also use the sensor data 134 to determine one or more provisional results 748. One or more provisional results 748 include data associated with event 624. For example, in the case where event 624 includes disambiguation of user 616, provisional result 748 may include a list of possible user 616 identifiers. In another example, when event 624 includes disambiguation between objects 104, provisional result 748 may include a list of possible object identifiers. In some specific implementations, provisional result 748 may indicate possible actions. For example, actions may include user 616 picking, placing, moving, damaging, or providing gesture input.

[0130] In some implementations, the provisional result 748 may be generated by other components. For example, the provisional result 748 (such as one or more possible identifiers or locations of user 616 involved in event 624) may be generated by the positioning component 124. In another example, the provisional result 748 (such as possible objects 604 that may be involved in event 624) may be generated by the identification component 734.

[0131] The event determination component 130 can be configured to provide a confidence level 744 associated with the determination of the provisional result 748. The confidence level 744 provides a marker regarding the expected level of accuracy of the provisional result 748. For example, a low confidence level 744 may indicate a low probability that the provisional result 748 corresponds to the actual occurrence of event 624. In contrast, a high confidence level 744 may indicate a high probability that the provisional result 748 corresponds to the actual occurrence of event 624.

[0132] In some specific implementations, a provisional result 748 with a confidence level 744 exceeding a threshold can be considered sufficiently accurate and thus usable as output data 626. For example, the event determination component 130 can provide provisional results 748 indicating three possible objects 604(1), 604(2), and 604(3) corresponding to the “selection” event 624. The confidence levels 744 associated with the possible objects 604(1), 604(2), and 604(3) can be 25%, 70%, and 92%, respectively. Continuing with this example, a threshold result can be set such that a 90% confidence level 744 is considered sufficiently accurate. Thus, the event determination component 130 can designate the “selection” event 624 as involving object 604(3).

[0133] The query component 738 can be configured to generate query data 750 using at least a portion of the sensor data 134 associated with event 624. In some embodiments, the query data 750 may include one or more provisional results 748 or supplementary data 752. The query component 738 can be configured to provide the query data 750 to one or more devices associated with one or more human employees.

[0134] The employee user interface is displayed on the employee's corresponding device. Employees can generate response data 754 by selecting a specific provisional result 748, entering new information, or indicating that they cannot answer a query.

[0135] Supplementary data 752 includes information associated with event 624 or information that can be used to interpret sensor data 134. For example, supplementary data 752 may include previously stored images of object 604. In another example, supplementary data 752 may include one or more graphical overlays. For example, graphical overlays may include graphical user interface elements, such as overlays of markers depicting related objects. These markers may include highlights, bounding boxes, arrows, etc., which are overlaid or placed on top of image data 626 during presentation to employees.

[0136] The query component 738 processes response data 754 provided by one or more employees. This processing may include calculating one or more statistical results associated with the response data 754. For example, the statistical results may include a count 748 of the number of times an employee selected a particular provisional result, a determination 748 of the percentage of employees who selected a particular provisional result, etc.

[0137] The query component 738 is configured to generate output data 626 based at least in part on the response data 754. For example, given that the response data 754 returned by most employees indicates that the object 604 associated with the “pick” event 624 is object 604(5), the output data 626 may indicate that object 604(5) was selected.

[0138] The query component 738 can be configured to selectively assign queries to specific employees. For example, some employees may be better suited to answering specific types of inquiries. Performance data (such as statistics on employee performance) can be determined by the query component 738 based on the response data 754 provided by the employees. For example, information indicating the percentage of different queries can be maintained, where a particular employee selected response data 754 that is inconsistent with the majority of employees. In some implementations, employees may be provided with test or practice query data 750 with previously known correct answers for training or quality assurance purposes. Determining the group of employees to be used can be based at least in part on performance data.

[0139] By using query component 738, event determination component 130 may be able to provide highly reliable output data 626 that accurately represents event 624. The output data 626 generated by query component 738 based on response data 754 can also be used to further train the automated system used by inventory management system 622. For example, sensor data 134 and output data 626 based on response data 754 can be provided to one or more components of inventory management system 622 for process improvement training. Continuing this example, this information can be provided to artificial neural networks, Bayesian networks, etc., to further train these systems so that the confidence level 744 and provisional results 748 generated for the same or similar inputs in the future are improved. Finally, as... Figure 7 As shown, server 118 may store and / or utilize other data 758.

[0140] The implementation may be provided as a software program or computer program product, including a non-transitory computer-readable storage medium on which instructions (in compressed or uncompressed form) are stored, which can be used to program a computer (or other electronic device) to perform the processes or methods described herein. The computer-readable storage medium may be one or more of electronic storage media, magnetic storage media, optical storage media, quantum storage media, etc. For example, computer-readable storage media may include, but is not limited to, hard disk drives, floppy disks, optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic or optical cards, solid-state storage devices, or other types of physical media suitable for storing electronic instructions. Additionally, the implementation may also be provided as a computer program product including transient machine-readable signals (in compressed or uncompressed form). Examples of machine-readable signals (whether modulated or unmodulated using a carrier wave) include, but are not limited to, signals that a computer system or machine hosting or running a computer program can be configured to access, including signals transmitted by one or more networks. For example, transient machine-readable signals may include software transmissions of the Internet.

[0141] Individual instances of these programs can be executed or distributed on any number of separate computer systems. Therefore, although some steps are described as being performed by certain devices, software programs, processes, or entities, this is not necessarily the case, and various alternative implementations will be understood by those skilled in the art.

[0142] Furthermore, those skilled in the art will readily recognize that the above-described technology can be used in a variety of devices, environments, and situations. Although the subject matter has been described in language specific to structural features or methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described. Rather, specific features and actions are disclosed as exemplary forms for implementing the claims.

[0143] While the foregoing invention has been described with reference to specific examples, it should be understood that the scope of the invention is not limited to these specific examples. Since other modifications and variations adapted to specific operational requirements and environments will be readily apparent to those skilled in the art, the invention should not be considered limited to the examples chosen for disclosure purposes, and the invention encompasses all modifications and variations that do not constitute a departure from the true spirit and scope of the invention.

[0144] The following terms can be used to describe the implementation scheme of this disclosure:

[0145] 1. A system comprising: a camera for generating image data of an environment; a scanner for generating scan data indicating a barcode associated with an object in the environment; and one or more computing devices, the one or more computing devices comprising: one or more processors; and one or more computer-readable media storing computer-executable instructions, which, when executed, cause the one or more processors to perform actions, the actions including: receiving an indication that the scanner has generated the scan data; determining a first time associated with the scanner generating the scan data; and determining a time associated with the scanner and in the environment. A Entity of Interest (VOI) within the territory, wherein the VOI is represented in the image data; analyzing a first frame of the image data, the first frame being associated with a second time following the first time; identifying an empty hand of a user within the VOI based at least in part on the analysis of the first frame; analyzing a second frame of the image data, the second frame being associated with a third time following the second time; identifying a full hand of the user within the VOI based at least in part on the analysis of the second frame; determining an object identifier associated with the object based on the scan data; and storing the object identifier of the object in association with the user's virtual shopping cart.

[0146] 2. The system according to Clause 1, wherein: the analysis of the first frame of the image data includes: inputting first feature data associated with the first frame into a trained classifier; and receiving a first score indicating whether the first frame includes a hand and a second score indicating whether the hand is empty or full, as output of the trained classifier; and the analysis of the second frame of the image data includes: inputting second feature data associated with the second frame into the trained classifier; and receiving a third score indicating whether the second frame includes a hand and a fourth score indicating whether the hand is empty or full, as output of the trained classifier.

[0147] 3. One or more computing devices, comprising: one or more processors; and one or more computer-readable media storing computer-executable instructions, which, when executed, cause the one or more processors to perform actions, the actions including: receiving sensor data generated by a sensor in an environment, the sensor data identifying an object; determining a portion of the environment associated with the sensor; receiving image data generated by a camera within the environment, the image data representing the portion of the environment associated with the sensor; analyzing the image data; identifying a hand represented in the image data based at least in part on the analysis; determining a user identifier associated with the hand; and updating virtual shopping cart data associated with the user identifier to indicate an object identifier associated with the object.

[0148] 4. One or more computing devices as described in Clause 3, wherein receiving the sensor data generated by the sensor includes receiving scan data generated by a scanning device that scans visual markers associated with the object.

[0149] 5. One or more computing devices according to Clause 3, wherein the one or more computer-readable media further stores computer-executable instructions, which, when executed, cause the one or more processors to perform actions, the actions including receiving data at a first time instructing the sensor to generate the sensor data, and wherein the analysis includes analyzing image data generated by the camera after the first time and within a threshold time amount of the first time.

[0150] 6. One or more computing devices according to Clause 3, wherein: determining the portion includes determining a volume of interest (VOI) relative to the sensor within the environment; and the analysis includes analyzing at least a portion of the image data corresponding to the VOI.

[0151] 7. One or more computing devices according to Clause 3, wherein: determining the portion includes determining a volume of interest (VOI) relative to the sensor within the environment; the analysis includes analyzing at least a portion of the image data corresponding to the VOI; and the identification includes identifying the hand after the hand enters the VOI.

[0152] 8. One or more computing devices according to Clause 3, wherein: the determination portion includes determining a volume of interest (VOI) relative to the sensor within the environment; the analysis includes: analyzing at least a portion of a first frame of image data corresponding to the VOI; and analyzing at least a portion of a second frame of image data corresponding to the VOI; the identification includes: identifying an empty hand within the VOI based at least in part on the analysis of the at least a portion of the first frame; and identifying a full hand within the VOI based at least in part on the analysis of the at least a portion of the second frame.

[0153] 9. One or more computing devices according to Clause 3, wherein: the determining portion includes determining a volume of interest (VOI) relative to the sensor within the environment; the analysis includes: analyzing at least a portion of a first frame of image data corresponding to the VOI; analyzing at least a portion of a second frame of image data corresponding to the VOI; analyzing at least a portion of a third frame of image data corresponding to the VOI; and analyzing at least a portion of a fourth frame of image data corresponding to the VOI; the identification includes: identifying an empty hand at a first location within the VOI based at least in part on the analysis of at least a portion of the first frame; and identifying an empty hand at a second location within the VOI based at least in part on the analysis of at least a portion of the second frame. The one or more computer-readable media further store computer-executable instructions that, when executed, cause the one or more processors to perform actions including: determining a first direction vector based at least partially on the first and second positions; and determining a second direction vector based at least partially on the third and fourth positions; the update including updating the virtual shopping cart data associated with the user identifier based at least partially on the first and second direction vectors.

[0154] 10. One or more computing devices according to Clause 3, wherein: the analysis comprises: generating a segmentation map using a first frame of the image data, the segmentation map identifying at least a first set of pixels in the first frame corresponding to a user's hand; and inputting first data indicating the first set of pixels in the first frame corresponding to the user's hand into a trained classifier; the identification comprises receiving second data, as the output of the trained classifier, indicating that the hand has received the object.

[0155] 11. One or more computing devices according to Clause 3, wherein: the determination of the portion includes determining a volume of interest (VOI) relative to the sensor within the environment; the analysis includes analyzing at least a portion of the image data corresponding to the VOI; the identification includes identifying the hand within the VOI; and the one or more computer-readable media further stores computer-executable instructions that, when executed, cause the one or more processors to perform an action, the action including identifying the object within the VOI.

[0156] 12. A method comprising: receiving sensor data generated by a sensor in an environment, the sensor data identifying an object; determining a portion of the environment associated with the sensor; receiving image data generated by a camera within the environment, the image data representing the portion of the environment associated with the sensor; analyzing the image data; identifying a hand represented in the image data based at least in part on the analysis; determining a user identifier associated with the hand; and updating virtual shopping cart data associated with the user identifier to indicate an object identifier associated with the object.

[0157] 13. The method according to Clause 12, wherein receiving the sensor data generated by the sensor includes receiving scan data generated by a scanning device that scans visual markers associated with the object.

[0158] 14. The method according to Clause 12 further includes receiving data at a first time instructing the sensor to generate the sensor data, and wherein the analysis includes analyzing image data generated by the camera after the first time and within a threshold time amount of the first time.

[0159] 15. The method according to Clause 12, wherein: determining the portion includes determining a volume of interest (VOI) relative to the sensor within the environment; and the analysis includes analyzing at least a portion of the image data corresponding to the VOI.

[0160] 16. The method according to Clause 12, wherein: determining the portion includes determining a volume of interest (VOI) relative to the sensor within the environment; the analysis includes analyzing at least a portion of the image data corresponding to the VOI; and the identification includes identifying the hand after the hand enters the VOI.

[0161] 17. The method according to Clause 12, wherein: the determination of the portion includes determining a volume of interest (VOI) relative to the sensor within the environment; the analysis includes: analyzing at least a portion of a first frame of image data corresponding to the VOI; and analyzing at least a portion of a second frame of image data corresponding to the VOI; the identification includes: identifying an empty hand within the VOI based at least in part on the analysis of the at least a portion of the first frame; and identifying a full hand within the VOI based at least in part on the analysis of the at least a portion of the second frame.

[0162] 18. The method according to Clause 12, wherein: the determination of the portion includes determining a volume of interest (VOI) relative to the sensor within the environment; the analysis includes: analyzing at least a portion of a first frame of image data corresponding to the VOI; analyzing at least a portion of a second frame of image data corresponding to the VOI; analyzing at least a portion of a third frame of image data corresponding to the VOI; and analyzing at least a portion of a fourth frame of image data corresponding to the VOI; the identification includes: identifying an empty hand at a first location within the VOI based at least in part on the analysis of at least a portion of the first frame; and identifying an empty hand at a second location within the VOI based at least in part on the analysis of at least a portion of the second frame; to The full hand at a third position within the VOI is identified, in part, based on the analysis of at least a portion of the third frame; and the full hand at a fourth position within the VOI is identified, in part, based on the analysis of at least a portion of the fourth frame; the one or more computer-readable media also store computer-executable instructions that, when executed, cause the one or more processors to perform actions including: determining a first direction vector based at least partially on the first position and the second position; and determining a second direction vector based at least partially on the third position and the fourth position; the update includes updating the virtual shopping cart data associated with the user identifier based at least partially on the first direction vector and the second direction vector.

[0163] 19. The method according to Clause 12, wherein the analysis comprises: generating a segmentation map using a first frame of the image data, the segmentation map identifying at least a first set of pixels in the first frame corresponding to a user's hand; and inputting first data indicating the first set of pixels in the first frame corresponding to the user's hand into a trained classifier; the identification comprising receiving second data, as the output of the trained classifier, indicating that the hand has received the object.

[0164] 20. The method according to Clause 12, wherein: determining the portion includes determining a volume of interest (VOI) relative to the sensor within the environment; the analysis includes analyzing at least a portion of the image data corresponding to the VOI; the identification includes identifying the hand within the VOI; and the one or more computer-readable media further stores computer-executable instructions that, when executed, cause the one or more processors to perform an action, the action including identifying the object within the VOI.

Claims

1. One or more computing devices, comprising: One or more processors; and One or more computer-readable media storing computer-executable instructions, which, when executed, cause the one or more processors to perform actions, including: Receive sensor data generated by sensors in the environment, the sensor data identifying objects; Determine a portion of the environment associated with the sensor; A first camera having a field of view (FOV) including a portion of the environment associated with the sensor is determined from multiple cameras within the environment; Receive image data generated by the first camera within the environment, the image data representing a portion of the environment associated with the sensor; Analyze the image data; The hand represented in the image data is identified, at least in part, based on the analysis. Determine the user identifier associated with the hand; and Update the virtual shopping cart data associated with the user identifier to indicate the item identifier associated with the item.

2. The computing device of claim 1 or more, wherein receiving the sensor data generated by the sensor includes receiving scan data generated by a scanning device that scans visual markers associated with the object.

3. The computing device of claim 1 or more, wherein the one or more computer-readable media further stores computer-executable instructions, which, when executed, cause the one or more processors to perform actions, the actions including receiving data at a first time instructing the sensor to generate the sensor data, and wherein the analysis includes analyzing image data generated by the first camera after the first time and within a threshold time amount of the first time.

4. The computing device according to claim 1, wherein: Determining the portion includes determining the volume of interest (VOI) relative to the sensor within the environment; and The analysis includes analyzing at least a portion of the image data corresponding to the VOI.

5. The computing device according to claim 1, wherein: Determining the portion includes determining the volume of interest (VOI) within the environment relative to the sensor; The analysis includes analyzing at least a portion of the image data corresponding to the VOI; The identification includes identifying the hand after it enters the VOI.

6. The computing device according to claim 1, wherein: Determining the portion includes determining the volume of interest (VOI) within the environment relative to the sensor; The analysis includes: Analyze at least a portion of the first frame of the image data corresponding to the VOI; and Analyze at least a portion of the second frame of the image data corresponding to the VOI; The identification includes: The empty hand within the VOI is identified, at least in part, based on the analysis of at least a portion of the first frame; and The full hand within the VOI is identified at least in part based on the analysis of at least a portion of the second frame.

7. The computing device according to claim 1, wherein: Determining the portion includes determining the volume of interest (VOI) within the environment relative to the sensor; The analysis includes: Analyze at least a portion of the first frame of the image data corresponding to the VOI; Analyze at least a portion of the second frame of the image data corresponding to the VOI; Analyze at least a portion of the third frame of the image data corresponding to the VOI; and Analyze at least a portion of the fourth frame of the image data corresponding to the VOI; The identification includes: The empty hand at a first location within the VOI is identified, at least in part, based on the analysis of at least a portion of the first frame; and The empty hand at the second location within the VOI is identified, at least in part, based on the analysis of at least a portion of the second frame; The full hand at the third position within the VOI is identified, at least in part, based on the analysis of at least a portion of the third frame; and The full hand at the fourth position within the VOI is identified at least in part based on the analysis of at least a portion of the fourth frame; The one or more computer-readable media further store computer-executable instructions, which, when executed, cause the one or more processors to perform actions, including: The first direction vector is determined at least in part based on the first position and the second position; and The second direction vector is determined at least in part based on the third position and the fourth position; The update includes updating the virtual shopping cart data associated with the user identifier based at least in part on the first direction vector and the second direction vector.

8. One or more computing devices according to claim 1, wherein: The analysis includes: A segmentation map is generated using a first frame of the image data, the segmentation map identifying at least a first group of pixels in the first frame corresponding to the user's hand; and First data indicating the first set of pixels corresponding to the user's hand in the first frame is input into a trained classifier; The identification includes receiving second data, which is the output of the trained classifier, indicating that the hand has received the object.

9. The computing device according to claim 1, wherein: Determining the portion includes determining the volume of interest (VOI) within the environment relative to the sensor; The analysis includes analyzing at least a portion of the image data corresponding to the VOI; The identification includes identifying the hand within the VOI; and The one or more computer-readable media also store computer-executable instructions that, when executed, cause the one or more processors to perform actions, including identifying the object within the VOI.

10. A method comprising: Receive sensor data generated by sensors in the environment, the sensor data identifying objects; Determine a portion of the environment associated with the sensor; A first camera having a field of view (FOV) including a portion of the environment associated with the sensor is determined from multiple cameras within the environment; Receive image data generated by the first camera within the environment, the image data representing a portion of the environment associated with the sensor; Analyze the image data; The hand represented in the image data is identified, at least in part, based on the analysis. Determine the user identifier associated with the hand; and Update the virtual shopping cart data associated with the user identifier to indicate the item identifier associated with the item.

11. The method of claim 10, wherein receiving the sensor data generated by the sensor includes receiving scan data generated by a scanning device that scans visual markers associated with the object.

12. The method of claim 10, further comprising: The system receives data at a first time instructing the sensor to generate the sensor data, and the analysis includes analyzing image data generated by the first camera after the first time and within a threshold time amount of the first time.

13. The method of claim 10, wherein: Determining the portion includes determining the volume of interest (VOI) relative to the sensor within the environment; and The analysis includes analyzing at least a portion of the image data corresponding to the VOI.

14. The method of claim 10, wherein: Determining the portion includes determining the volume of interest (VOI) within the environment relative to the sensor; The analysis includes analyzing at least a portion of the image data corresponding to the VOI; The identification includes identifying the hand after it enters the VOI.

15. The method of claim 10, wherein: Determining the portion includes determining the volume of interest (VOI) within the environment relative to the sensor; The analysis includes: Analyze at least a portion of the first frame of the image data corresponding to the VOI; and Analyze at least a portion of the second frame of the image data corresponding to the VOI; The identification includes: The empty hand within the VOI is identified, at least in part, based on the analysis of at least a portion of the first frame; and The full hand within the VOI is identified at least in part based on the analysis of at least a portion of the second frame.