Managing retail self-checkout workspaces
A central camera system in retail stores addresses scanning errors and shoplifting by locking SCO terminals based on error thresholds, improving customer satisfaction and operational efficiency.
Patent Information
- Application Number
- JP2025513040
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-29
- Filing Date
- 2023-06-01
- Publication Date
- 2025-09-04
AI Technical Summary
Self-checkout (SCO) terminals in retail stores face issues with customers making scanning errors or intentionally avoiding payment, leading to shoplifting and increased wait times due to terminal locking, which reduces customer satisfaction and increases operational costs.
A central camera system captures an overview of the SCO area, identifying non-scanning events and locking terminals based on thresholds of concurrent and consecutive errors, balancing terminal usage to minimize fraud and wait times.
Reduces shoplifting and wait times by intelligently managing SCO terminals, enhancing customer satisfaction and operational efficiency.
Smart Images

Figure 2025529220000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to retail stores, and more particularly to managing self-checkout (SCO) work areas in retail stores. [Background technology]
[0002] SCO terminals provide a mechanism for customers to process their own purchases from a retail store, replacing traditional cashier-assisted checkout. A typical retail store contains an SCO work area, an area where multiple SCO terminals are installed. At the SCO terminals, customers are required to personally scan each item with a scanner, where they then perform the necessary payment.
[0003] Customers may have little or no training in operating SCO terminals and may make mistakes when self-checking out items. Customers may unintentionally miss some items during scanning and leave the store without making the required payment. Furthermore, shoplifting is a major drawback with SCO terminals. For example, customers may intentionally not scan some items, place the unscanned items in their shopping cart, and leave the store without paying in full. As a result, retailers may suffer significant losses. Systems exist to detect discrepancies between the products in a user's shopping basket and the list of scanned items generated by the scanner. If a discrepancy occurs, such systems alert a store associate and lock the corresponding SCO terminal, meaning the customer cannot continue scanning products.
[0004] However, locking SCO terminals leads to an increase in the overall time spent by users attending to them, which increases queues and reduces customer satisfaction in the SCO workspace. Also, because the number of store associates in a SCO workspace is limited, overall wait times increase when more SCO terminals require attendant attention. In view of the above, there is a need for a system and method for managing a retail store that reduces queues and improves customer satisfaction in the SCO workspace. Summary of the Invention
[0005] In one aspect of the present disclosure, a system for managing multiple SCO terminals in a SCO work area of a retail store is provided. The system includes a central camera that captures an overview image of the SCO work area and a central control unit communicatively coupled to a processor at each SCO terminal. The central control unit includes a memory that stores one or more instructions and a central processing unit communicatively coupled to the memory that executes the one or more instructions. The central processing unit is configured to identify a non-scanning event at the SCO terminal, check whether the number of other SCO terminals already locked is less than a first threshold, lock the SCO terminal if the number of other locked SCO terminals is less than the first threshold, and if the number of other locked SCO terminals reaches the first threshold, perform a check to determine whether the number of consecutive non-scanning events at the SCO terminal has reached a second threshold, and lock the SCO terminal if the number of consecutive non-scanning events detected at the SCO terminal reaches the second threshold.
[0006] In another aspect of the present disclosure, a method for managing multiple SCO terminals in a SCO work area of a retail store is provided, the method including: capturing an overview image of the SCO work area with a central camera; identifying a non-scanning event at the SCO terminal; checking whether a number of other SCO terminals already locked is less than a first threshold; locking the SCO terminal if the number of other locked SCO terminals is less than the first threshold; if the number of other locked SCO terminals reaches the first threshold, checking whether a number of consecutive non-scanning events at the SCO terminal reaches a second threshold; and locking the SCO terminal if the number of consecutive non-scanning events detected at the SCO terminal reaches the second threshold.
[0007] In yet another aspect of the present disclosure, there is provided a computer programmable product for managing a plurality of SCO terminals in a SCO work area of a retail store, the computer programmable product including a set of instructions that, when executed by a processor, cause the processor to capture an overview image of the SCO work area with a central camera, identify a non-scanning event at the SCO terminal, check whether a number of other SCO terminals already locked is less than a first threshold, lock the SCO terminal if the number of other locked SCO terminals is less than the first threshold, perform a check to determine whether a number of consecutive non-scanning events at the SCO terminal has reached a second threshold if the number of consecutive non-scanning events detected at the SCO terminal has reached the second threshold, and lock the SCO terminal if the number of consecutive non-scanning events detected at the SCO terminal has reached the second threshold.
[0008] It will be understood that features of the present disclosure can be combined in various combinations without departing from the scope of the present disclosure as defined by the appended claims. [Brief explanation of the drawings]
[0009] The foregoing summary, as well as the following detailed description of exemplary embodiments, will be better understood when read in conjunction with the accompanying drawings. For the purpose of illustrating the disclosure, there are shown in the drawings exemplary configurations of the disclosure. However, the disclosure is not limited to the particular methods and instrumentalities disclosed herein. Moreover, those skilled in the art will appreciate that the drawings are not to scale. Wherever possible, like elements will be designated by like numerals.
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
[0011] In the accompanying drawings, underlined numbers are used to identify the item on which the underlined number is located or to which the underlined number is adjacent. Numbers without underlines relate to the item identified by the line linking the ununderlined number to the item. When a number is not underlined and has an associated arrow, the ununderlined number is used to identify the general item to which the arrow is pointing. DETAILED DESCRIPTION OF THE INVENTION
[0012] The following detailed description sets forth embodiments of the present disclosure and methods of practicing them. Although best modes for carrying out the disclosure are disclosed, those skilled in the art will recognize that there are other possible embodiments for making or practicing the disclosure.
[0013] Referring to FIG. 1, a retail store environment 100 in which various embodiments of the present disclosure can be practiced is shown. The retail store environment 100 includes first through third shelves 102a through 102c for storing and displaying one or more items. The retail store environment 100 further includes first through third cashier terminals 104a through 104c, which are located at first through third cashiers 106a through 106c, respectively, for scanning and claiming items in corresponding shopping carts of customers. The retail store environment 100 further includes an SCO work area 108, which includes one or more SCO terminals and enables customers to scan and claim items in their respective shopping carts. The SCO work area is described in further detail with reference to FIG. 2.
[0014] 2 illustrates a central control unit 200 managing a retail store's SCO work area 108, according to one embodiment of the present disclosure. The SCO work area 108 includes first through fifth SCO terminals 202a through 202e (hereafter referred to as SCO terminals 202), corresponding first through fifth customers 204a through 204e with first through fifth shopping carts 206a through 206e, a central camera 208, and one or more store associates 210.
[0015] In one embodiment of the present disclosure, various components of the SCO workspace 108 may be communicatively coupled to the central control unit 200 via a communications network. The communications network may be any suitable wired network, wireless network, combination thereof, or any other conventional network without limiting the scope of the present disclosure. In some examples, the network may include a local area network (LAN), a wireless LAN connection, an Internet connection, a point-to-point connection, or other network connections and combinations thereof. In one example, the network may include, for example, a 2G, 3G, 4G, or 5G mobile communications network. The communications network may be connected to one or more other networks, thereby providing connectivity between multiple devices. This may be the case when the networks are coupled together via the Internet.
[0016] Each SCO terminal 202a-202e includes a scanner that allows customers to scan one or more items themselves and a user display that allows the user to select and pay for one or more items. In one example, the scanner may be a barcode scanner that scans item barcodes to identify the items. Preferably, the scanner is a wall-mounted scanner or table-mounted scanner designed for checkout counters in supermarkets and other retail stores that scans items placed in a scan zone. In the context of the present disclosure, a scan zone is the area in front of the scanner where a user brings items to be scanned for purchase. Each SCO terminal 202a-202e may include a processor (not shown) that records the scanning of one or more items and provides instructions corresponding to the user display for paying for one or more scanned items. In one embodiment of the present disclosure, the processor of each SCO terminal 202a-202e may be communicatively coupled to the central control unit 200 to enable the central control unit 200 to control the operation of the SCO terminal 202 and also to process information captured by the central camera 208.
[0017] In one embodiment of the present disclosure, each SCO terminal 202a-202e is equipped with one or more overhead cameras 207a-207e that continuously capture each corresponding scan zone, facilitating the detection of fraudulent scans due to discrepancies between items picked up for scanning by a user and the items actually scanned at each SCO terminal 202a-202e. Fraudulent scans occur when an item identified by scanning during a predetermined time interval is not present in the list of scanned items generated by the scanner during the corresponding time interval. In one example, a user may place an item in a scanner's scan zone, but the user may hold the item so that the item is not visible to the barcode scanner. In such a case, the user may place the item in their shopping bag after performing the scanning action, but it is not actually scanned by the scanner, and the user may not receive a claim for the item. In one embodiment of the present disclosure, overhead cameras 207a to 207e are communicatively coupled to central control unit 200, and central control unit 200 is configured to control the operation of overhead cameras 207a to 207e and to process information captured by camera 208.
[0018] The central camera 208 is configured to generate an overview image of the entire SCO work area 108. Examples of the central camera 208 include an overhead 360° camera, a 180° camera, etc. In one embodiment of the present disclosure, the central camera 208 may be communicatively coupled to the central control unit 200 to enable the central control unit 200 to control the operation of the central camera 208 and to process information captured by the central camera 208. The central camera 208 may facilitate an enhanced customer experience at the SCO work area 108; for example, a customer with a child or an overflowing shopping cart may be detected by the central camera 208 at the entrance to the SCO work area, and a store associate 210 may be alerted to offer assistance with the checkout process. If an attendant is unavailable, an attendant may be given priority to provide assistance when one becomes available.
[0019] Although not shown, the central control unit 200 is communicatively coupled to the computing devices of the store clerks 210 and may issue alerts / warnings or instructions thereto.
[0020] In one embodiment of the present disclosure, various components of the SCO workspace 108 may be communicatively coupled to the central control unit 200 via a communications network. The communications network may be any suitable wired network, wireless network, combination thereof, or any other conventional network without limiting the scope of the present disclosure. In some examples, the communications network may include a local area network (LAN), a wireless LAN connection, an Internet connection, a point-to-point connection, or other network connections and combinations thereof. In one example, the network may include, for example, a 2G, 3G, 4G, or 5G mobile communications network. The communications network may be connected to one or more other networks, thereby providing connectivity between multiple devices. This may be the case when the networks are coupled together via the Internet.
[0021] In one embodiment of the present disclosure, central control unit 200 includes a central processing unit 214, a memory 216, and an operation panel 218. Central processing unit 214 includes a processor, computer, microcontroller, or other circuitry that controls the operation of various components, such as operation panel 218 and memory 216. Central processing unit 214 may execute software, firmware, and / or other instructions stored in, or otherwise provided to, volatile or non-volatile memory, such as memory 216. Central processing unit 214 may be connected to operation panel 218 and memory 216 via wired or wireless connections, such as one or more system buses, cables, or other interfaces.
[0022] The operation panel 218 may be a user interface, which may take the form of a physical keypad or a touchscreen. The operation panel 218 may accept input from one or more users regarding selected functions, preferences, and / or authorizations, and may provide and / or accept visual and / or audio inputs.
[0023] In addition to storing instructions and / or data used by central processing unit 214, memory 216 may also contain user information associated with one or more operators of SCO workspace 108. For example, user information may include authentication information (e.g., username / password pairs), user preferences, and other user-specific information. Central processing unit 214 may access this data to assist in providing control functions (e.g., sending and / or receiving one or more control signals) related to the operation of operation panel 218 and memory 216.
[0024] In one embodiment of the present disclosure, the central processing unit 214 is configured to detect one or more fraudulent scans based on information received from the overhead cameras 207a-207e and the scanners of the SCO terminals 202a-202e, and to lock the corresponding one or more SCO terminals 202a-202e based on the detected fraudulent scans, i.e., to prevent customers from continuing to scan products. Upon locking, the central processing unit 214 may alert the store associate 210, as appropriate. In the context of the present disclosure, the SCO associate 210 may manually verify whether the reported fraudulent scans are valid.
[0025] In one embodiment of the present disclosure, the central processing unit 214 is configured to automatically lock a SCO terminal, such as the first SCO terminal 202a, based on the lock status of other SCO terminals. In one example, the central processing unit 214 is configured to lock the first SCO terminal 202a if an unauthorized scan is detected and if the number of other SCO terminals already locked, such as the second and third SCO terminals 202b and 202c, is less than a first threshold. If the number of other SCO terminals already locked is greater than the first threshold, the central processing unit 214 disables the lock on the first SCO terminal 202a unless the number of unauthorized scans detected at the first SCO terminal 202a reaches a second threshold.
[0026] In another embodiment of the present disclosure, the central processing unit 214 is configured to automatically lock a SCO terminal, such as the first SCO terminal 202a, based on the location of the store clerk or SCO work area manager 210 and their status, i.e., whether they are free or busy. Location refers to actual physical location, and the physical location of the store clerk and the location of the SCO terminal are used to determine the distance between them. A smaller distance would mean a shorter response time from the store clerk 210. To take advantage of this, the central processing unit 214 would have the ability to lock the first SCO terminal 202a only if the store clerk is within a predetermined distance from the SCO terminal. If the distance is greater than the predetermined distance, the central processing unit 214 would not lock the SCO terminal 202a.
[0027] In yet another embodiment of the present disclosure, the central processing unit 214 is configured to automatically lock SCO terminals, such as the first SCO terminal 202a, based on the length of each SCO terminal's sequence of non-scanning events since the last lock. Although a non-scanning event occurs at the first SCO terminal 202a, the central processing unit 214 may not lock the terminal to reduce customer friction. In one example, during Black Friday, the central processing unit 214 may be configured to ignore the first three non-scanning events at the first SCO terminal 202a. However, if a fourth non-scanning event occurs at the first SCO terminal 202a, the first SCO terminal 202a may be locked.
[0028] In yet another embodiment of the present disclosure, the central processing unit 214 is configured to automatically lock a SCO terminal, such as the first SCO terminal 202a, based on the status of the corresponding cart, such as a full cart (a cart with many products), a bulk-loaded cart (a cart with few but numerous items), or a cart with a large object (i.e., a TV). A large object is an object whose size is greater than a predetermined size threshold. Scanning a bulk-loaded cart is also much faster because it only requires scanning a few items and manually entering the number of occurrences of the items. In one example, the central processing unit 204 may be configured to lock the first SCO terminal 202a when a full cart or bulk-loaded cart is detected and a store associate is immediately available, so that the corresponding customer will receive assistance from the store associate 210. The central processing unit 214 is further configured to notify the store associate 210 for proactive assistance when a full cart or bulk-loaded cart is detected at the entrance to the SCO work area 108.
[0029] In yet another embodiment of the present disclosure, the central processing unit 214 is configured to automatically lock the exit gate of the SCO work area 108 and notify the store associate 210 if a large product moves through the exit of the SCO work area 108 without appearing to have been scanned in the list of scanned products. The exit gate is a gate in the retail store through which products may be removed after self-checkout is complete.
[0030] In yet another embodiment of the present disclosure, the central processing unit 214 is configured to notify the store associate 210 to investigate when ownership of a product changes from one customer to another.
[0031] In yet another embodiment of the present disclosure, the central processing unit 214 is configured to send an alert to the computing device of the store associate 210 if the queue size at the entrance to the SCO work area 108 is greater than a predetermined third threshold, thereby allowing more potentially available store associates to be assigned to the area. The alert may be an audio signal, a visual display, a tactile alert, an instant message, or the like. The entrance to the SCO work area 108 may be an entry point where a customer enters the SCO work area 108 to begin the self-checkout process. The central processing unit 214 may be configured to change the first and second thresholds if the queue length at the entrance to the SCO work area 108 exceeds the third threshold. In the context of the present disclosure, the queue length may be determined automatically using a 360-degree camera.
[0032] In yet another embodiment of the present disclosure, the central processing unit 214 is configured to automatically lock the SCO terminal based on an emergency situation, for example, someone having a gun. In an embodiment of the present disclosure, the emergency situation can be detected using the video camera and the central camera 208. In one example, someone having (actually waving) a gun can be detected using the central camera 208.
[0033] In an embodiment of the present disclosure, the above parameters may be pre-configured by the store manager of the corresponding retail store or by a person in charge of the overall security system. Based on the pre-configured parameters, real-time information captured by the central camera 208 and overhead cameras 207a-207e, the status of the SCO terminal 202, and the status of the SCO clerk 210, the central processing unit 214 automatically controls the locking of the SCO terminal 202 and sends messages to the clerk 210 and the store manager. In one embodiment of the present disclosure, the central processing unit 214 is configured to dynamically generate and adapt interactions between customers of the retail store and each SCO terminal 202a-202e in the SCO workspace 108 to optimize customer flow at the SCO terminal 202. The SCO terminal may be unlocked by the intervention of the clerk / attendant / SCO workspace manager 210.
[0034] In various embodiments of the present disclosure, the central processing unit 214 is configured to reduce overall waiting lines and increase customer satisfaction in the SCO work area 108 by balancing the loss of attendant intervention at the SCOs 202a-202e against the loss of potential product leakage (product that may leave the retail store without being claimed). In one embodiment of the present disclosure, the central processing unit 214 may be configured to calculate the loss value of customer wait time and leakage items per minute. This loss may be weighted relative to other leakage items. The central processing unit 214 may further be configured to predict total wait time by taking into account the number of locked terminals, the number of attendants in the area, and the length of the queues at the entrances to the SCO area, and to build a model that suggests how many additional minutes of wait time will be added in response to a new alert.
[0035] 2 is merely an example. Those skilled in the art would recognize many variations, alternatives, and modifications of the embodiments herein.
[0036] 3 illustrates steps in a method 300 for managing a SCO workspace 108 by the central control unit 200 of the central processing unit 214 in accordance with the present disclosure. The method is depicted as a collection of steps in a logical flow diagram, which represents a sequence of steps that can be implemented in hardware, software, or a combination thereof.
[0037] In step 302, a non-scan event is identified at a SCO terminal in a SCO workspace. A non-scan event is referred to as an event when a user picks up an item for scanning with a corresponding scanner in the scan area, which may or may not be successfully scanned by the scanner. In one example, a user may place an item within the scan area of a scanner, but the user may hold the item so that the item's barcode is not visible to the barcode scanner. Actions corresponding to a non-scan event are not captured by the scanner, but may be captured by an overhead camera located there.
[0038] In step 304, a check is performed to determine if the number of other locked SCO terminals is less than a first threshold, and if the number is less than the first threshold, the SCO terminal is automatically locked in step 306. The value of the first threshold may be set based on the number of SCO terminals and store associates in the corresponding SCO workspace.
[0039] If the number of other locked SCO terminals reaches a first threshold, then a check is performed in step 308 to determine if the number of consecutive non-scan events at the SCO terminal has reached a second threshold.
[0040] If the number of consecutive non-scanning events detected reaches a second threshold, the SCO terminal is automatically locked in step 310. In one example, the value of the first threshold may be 2, and the value of the second threshold may be 3. Thus, when at least two terminals are already locked, a third terminal is locked only if the current non-scanning event is the third non-scanning event in the current procedure.
[0041] 3 is merely an example. Those skilled in the art would recognize many variations, alternatives, and modifications of the embodiments herein.
[0042] 4 illustrates a block diagram of software 400 for managing a retail store's SCO workspace, according to one embodiment of the present disclosure. The software 400 monitors multiple video cameras C1 through C2 installed at various locations around the retail store (not shown). n a video unit 402 communicatively coupled to a plurality of video sensors, comprising video cameras C1 to C n Each of at least some of the SCO terminals SCO1 to SCO2 in a retail store (not shown) n Specifically, the video cameras C1 to C n At least some of the SCO terminals SCO1 to SCO n It is installed directly above the station to obtain a bird's-eye view of the station.
[0043] In one embodiment, video cameras C1 to C n is the video camera C1 to C n The video cameras C1 to C are configured to capture video footage of the environment within the field of view of the n The video footage from (not shown) includes multiple consecutively captured video frames, where p is the number of video frames in the captured video footage.
[0044]
number
[0045] is the video camera C1 to C n At some point
[0046]
number
[0047] (also called sampling time), where τ is the time when the video footage capture begins,
[0048]
number
[0049] is the time interval (also called the sampling interval) between the capture of the first video frame and the capture of the next video frame. Using this notation, the video footage captured by the first camera 26 is represented by the video cameras C1 through C2. n The video footage captured by
[0050]
number
[0051] It can be described as:
[0052] In one embodiment, the software 400 further includes a plurality of SCO terminals SCO1 to SCO2 within the retail store. n In particular, the SCO unit 404 includes an SCO terminal SCO1 to an SCO n The sales till data is received from the SCO terminal SCO1 through the SCO terminal SCO2. n During a product scan performed on SCO terminal SCO1, n The sales register data includes the Universal Product Code (UPC) of the product as detected by a scanner device (not shown). The sales register data also includes the quantity of the same product.
[0053] In one embodiment, the SCO unit 404 receives SCO terminals SCO1 through SCO2. n The status signal is further configured to receive a status signal from the SCO terminal SCO1. n The status signal may further include an indicator showing whether the SCO terminal SCO1 is locked or active. nThe status signal may include a timestamp of when the device was locked. In one embodiment, the status signal may be obtained from the NCR Remote Access Program (RAP) Application Program Interface (API). However, those skilled in the art will recognize that the above status signal sources are provided for illustrative purposes only. In particular, those skilled in the art will appreciate that the software of the preferred embodiment is not limited to the above status signal sources. To the contrary, the software of the preferred embodiment is operable with any status signal source, including any manufacturer's API for SCO terminals.
[0054] In one embodiment, the SCO unit 404 sends a message to each SCO terminal SCO1 through SCO2 to lock the SCO terminal. n In one embodiment, a particular SCO terminal SCO1 is further configured to issue a control signal to the SCO n The issuance of control signals to the SCO terminals, the receipt of control signals by the associated SCO terminals, and the execution of locking operations in response to the received control signals are accomplished through the NCR Remote Access Program (RAP) Application Program Interface (API). However, those skilled in the art will recognize that the above mechanisms for issuance of control signals to the SCO terminals, the receipt of control signals by the associated SCO terminals, and the execution of locking of the SCO terminals are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the software of the preferred embodiment is not limited to the above mechanisms. Rather, the software of the preferred embodiment is operable with any mechanism for issuing control signals to, receiving control signals by, and executing the locking of the SCO terminals, e.g., the API of the SCO terminal manufacturer.
[0055] In one embodiment, the software 400 further comprises a control unit 406 communicatively coupled to the video unit 402 and the SCO unit 404. The control unit 406 controls the video cameras C1 through C2 from the video unit 402. nThe control unit 406 receives from the SCO unit 404 the video footage captured by each SCO terminal SCO1 to SCO2. n The control unit 406 is further configured to receive a status signal from the SCO unit 404. The status signal may include an indicator indicating whether the corresponding SCO is locked or operational. If the SCO is locked, the status signal from the SCO terminal may include a timestamp indicating the time the SCO terminal was locked. Similarly, the control unit 406 is configured to issue a control signal to the SCO unit 404, which is configured to lock the designated SCO terminal.
[0056] In one embodiment, the control unit 406 is further communicatively coupled to a Human Classification Module 408, a Human Tracking module 410, a Motion Detection module 412, and an Object Recognition module 414, each of which and their operation are described in more detail below. The control unit 406 itself comprises a processing unit 416 communicatively coupled to a logic unit 418, which is communicatively coupled to the SCO unit 404, each of which and their operation are also described in more detail below.
[0057] In one embodiment, the person classification module 408 monitors video cameras C1 to C2 installed at various locations around a retail store (not shown). n In another embodiment, the human classification module 408 is configured to receive from the control unit 406 video frames from video footage captured by the
[0058]
number
[0059] to detect the presence of persons therein and classify each detected person as one of a child, an adult customer, and a staff member.
[0060] In one embodiment, the human classification module 408 may be implemented by an object detection machine learning (ML) algorithm such as EfficientDet (described in M. Tan, R. Pang and Q.V. Le, EfficientDet: Scalable and Efficient Object Detection, 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2020, pp. 10778-10787). Alternatively, the human classification module (HCM) 408 may be implemented by a panoptic segmentation algorithm such as a bidirectional aggregation network (BANet) (Y. Chen, G. Lin, S. Li, O. Bourahla, Y. Wu, F. Wang, J. Feng, M. Xu, X. Li, Banet: Bidirectional aggregation network with occlusion handling for panoptic segmentation, in: Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 3793-3802).
[0061] Those skilled in the art will recognize that the above examples of algorithms used for object detection and panoptic segmentation are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the preferred embodiment is not limited to the above algorithms. On the contrary, the preferred embodiment can operate with any algorithm suitable for detecting objects in video frames or for combining instance and semantic segmentation of video frames, such as YOLOv4 (described in A. Bochkovskiy, C.Y. Wang and H.Y. M. Liao, 2020 arXiv: 2004.10934) and AuNet (described in Y. Li, X. Chen, Z. Zhu, L. Xie, G. Huang, D. Du, X. Wang, Attention-guided unified network for panoptic segmentation, in: Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 7026-7035).
[0062] The objectives of object detection or panoptic segmentation algorithms are to: Detect one or more people appearing in a video frame. Establishing localization information for the detected person (e.g., by a frame or a frame of interest (hereafter also called a bounding box) established by an object detection algorithm around the detected person in the video frame). Determine whether a detected person is an adult customer, a child, or a staff member.
[0063] FIG. 5 illustrates video cameras C1 to C2 installed in a retail store (not shown) according to one embodiment of the present disclosure. n Video frames captured by
[0064]
number
[0065] 4 and the processing of video frames by the human classification module 408 of the software 400 of FIG. 4.
[0066] In one embodiment, referring to FIG. 5 , an object detection algorithm detects a primary object of interest consisting of a person in a received video frame 500 and ignores secondary objects of interest such as cash registers, shopping carts, and stacks of merchandise that appear in the video frame 500. Detected people are indicated by bounding boxes that substantially enclose each person. The bounding boxes facilitate subsequent tracking of individual people. The object detection algorithm then distinguishes between staff members 502, children 504, and adult customers 506. Distinguishing between staff members 502 and adult customers 506 may be based on the staff members 502 wearing uniforms with distinctive colors or patterns that include a prominent logo.
[0067] In one embodiment, an object detection or panoptic segmentation algorithm is trained using video frames selected from video footage captured by multiple video cameras installed at various locations within a retail store. The video frames are hereinafter referred to as the training dataset. The individual video frames of the training dataset are selected and compiled to provide robust, class-balanced information about staff members, children, and adult customers obtained from views acquired at different positions and orientations relative to the video cameras. Furthermore, the video frames of the training dataset are selected from video footage acquired from various locations within the retail store. Similarly, the video frames of the training dataset include individuals wearing various types and colors of clothing. Members of the training dataset may be further subjected to data augmentation techniques (e.g., rotation, flipping, brightness modification) to generate more video frames, thereby increasing the size of the training dataset, preventing overfitting, normalizing the deep neural network model, balancing the classes within the training dataset, and synthetically generating new video frames that better represent the current task. Thus, the video frames of the training dataset are balanced with respect to gender, age, and skin color.
[0068] Video frames of the training dataset for the object detection algorithm are manually tagged with bounding boxes positioned to substantially enclose each individual appearing in the video frame and corresponding class labels for adult customers / staff / children, as appropriate. Members of the training dataset are organized in pairs, with each data pair including a video frame and a corresponding XML file. The XML file includes the coordinates of the bounding boxes relative to the coordinate system of the video frame and the corresponding label for each bounding box.
[0069] In contrast, individual pixels in each video frame of a training dataset for a panoptic segmentation algorithm are manually tagged with a class label—adult customer, staff member, or child—as needed. Individual pixels are also tagged with an instance number that indicates which instance of a given class the pixel corresponds to. For example, the instance number may indicate whether the pixel corresponds to the second adult customer appearing in the video frame or the third child appearing in the video frame. The members of the training dataset are organized in a pairwise manner, with each data pair including a video frame and a corresponding XML file. The XML file includes the class label and instance number for each pixel in the corresponding video frame.
[0070] In one embodiment, returning to FIG. 4, once the object detection algorithm is trained, the video frames subsequently received by the human classification module 408
[0071]
number
[0072] Its output in response to is a set of bounding boxes (each bounding box is defined by two opposite corners) (Bnd_Bx i (t)) and a corresponding set of class labels.
[0073] In contrast, once the panoptic segmentation algorithm is trained, the video frames subsequently presented to the human classification module 408
[0074]
number
[0075] Its output responds to video frames
[0076]
number
[0077] The human classification module 408 is configured to communicate this output to the control unit 406.
[0078] In one embodiment, the human tracking module 410 includes video cameras C1 to C2 installed at various locations. n The video frame capturing unit 406 is configured to receive from the control unit 406 video frames from the video footage captured by the video frame capturing unit 406.
[0079] Typical person re-identification algorithms assume that a person's physical appearance does not change significantly between video frames. Therefore, physical appearance is an important piece of information that can be used to re-identify a person. Therefore, the human tracking module 410 represents a person through a variety of rich semantic features related to visual appearance, body movements, and interactions with the surroundings. These semantic features essentially form a biometric signature of the person, which is used to re-identify the person in different video frames.
[0080] In one embodiment, the human tracking module 410 builds an internal repository of semantic features of people in the store. For simplicity, this internal repository is hereafter referred to as the Gallery Feature Set. The Gallery Feature Set is populated with feature representations of each person extracted by a trained person re-identification neural network model. Since the specific identities of these people are largely unknown, each person's semantic features are associated with them through person identification data. Person identification data essentially includes a person identifier (PID). In other words, the human tracking module 410 maps a person's biometric signature to that person's PID. i Link to the PID i and the corresponding biometric information in the gallery feature set is deleted at the end of each day, or more frequently at the request of the operator.
[0081] In one embodiment, a further video frame in which a person appears (i.e., a query image of the query person) is selected, and the trained person re-identification network extracts a feature representation of the person and establishes its associated semantic features. The feature representation of the person in the query image may correspond to the query identification data. The extracted feature representation is compared with the feature representations in the gallery feature set. If a match is found, the person is identified by a PID corresponding to the matching feature representation in the gallery feature store. i If the query person's feature representation from the query image does not match any in the gallery feature set, a new unique PID i is assigned to the person, and the corresponding feature representation of the person is added to the gallery feature set, and PID i It is associated with.
[0082] In one embodiment, the person re-identification network uses a standard ResNet architecture. However, those skilled in the art will recognize that this architecture is provided for illustrative purposes only. In particular, those skilled in the art will recognize that the preferred embodiment is not limited to use with this architecture. Rather, the preferred embodiment is operable with any neural network architecture capable of forming an internal representation of a person's semantic features. For example, the person re-identification network may use a Batch Normalization (BN)-Inception architecture to normalize layer inputs by recentering and rescaling, thereby making the training of machine learning algorithms faster and more stable. In use, the person re-identification network is trained using a dataset including: A video frame in which a person appears. Annotated bounding boxes that essentially surround each person appearing in each video frame.
[0083] In one embodiment, the annotation for each bounding box includes the PID of the person enclosed in the bounding box. iThis allows the same person to be identified across multiple video frames collected from a set of video cameras, and therefore the same PID i A set of bounding boxes annotated with encapsulates the appearance information of the same person extracted from different views. Therefore, the training data consists of the frame number, the PID of the person represented in the video frame, i , and a set of video frames each described by the corresponding bounding box details. Because multiple people may appear in a video frame, the training data for such a video frame contains multiple entries, one for each person represented in the video frame.
[0084] The output from the human tracking module 410 is video frames captured by video cameras placed at various locations within the retail store.
[0085]
number
[0086] 4 is a dataset detailing the time and location where a person was detected in a video frame. The location where the person was detected is established from the coordinates of a bounding box established in each video frame in which the person appears and the ID of the video camera that captured the video frame. The output from the human tracking module 410 may also include extracted feature representations of the person.
[0087] In one embodiment, the human tracking module 410 is configured to communicate an output from the human tracking module 410 to the control unit 406 .
[0088] In one embodiment, the motion detection unit 412 is configured to receive video footage from a video camera (not shown) located directly above a SCO terminal (not shown) in a retail store (not shown) and provide a bird's-eye view of the SCO terminal (not shown). The motion detection unit 412 detects successively captured video frames within the received video footage.
[0089]
number
[0090] and,
[0091]
number
[0092] to detect movement within a predetermined distance of the SCO terminal (not shown), where the predetermined distance is determined by intrinsic parameters of the video camera (not shown), which in combination with a position directly above the SCO terminal (not shown) establish a field of view for the video camera (not shown).
[0093] In one embodiment, video frames in the received video footage are encoded using the H.264 video compression standard. The H.264 video format uses motion vectors as a key factor in compressing video footage. The motion detection unit 412 uses the motion vectors obtained from decoding the H.264 encoded video frames to detect motion within a predetermined distance of the SCO terminal (not shown). In another embodiment, successive samples from the video footage are
[0094]
number
[0095] and compare them to detect differences between them. A difference exceeding a predetermined threshold is taken to indicate that motion has occurred during the intervening period between successive samples. The threshold is set to avoid temporary changes such as light flicker being mistaken for motion. When the motion detection unit 412 detects motion within a predetermined distance of the SCO terminal (not shown), a "motion trigger" signal is sent from the motion detection unit 412 to the control unit 406.
[0096] In one embodiment, the object recognition module 414 is configured to receive video footage from a video camera mounted directly above the SCO terminal (not shown) in a retail store (not shown) to provide a bird's-eye view of the SCO terminal (not shown). The object recognition module 414 is also configured to receive video footage from a video camera (not shown) mounted within a predetermined distance of the SCO terminal (not shown), with its field of view positioned to encompass the area where customers approach the SCO terminal (not shown) with products to purchase. The object recognition module 414 is configured to recognize and identify specific objects appearing within the received video frames. In one embodiment, the objects may include inventory items in the retail store. In another embodiment, the objects may include a full shopping cart.
[0097] Reference is now made to FIG. 6, which illustrates an object recognition module of the software of FIG. 4 that manages the SCO workspace of a retail store, according to one embodiment of the present disclosure.
[0098] In one embodiment, the object recognition module 414 includes an object detection module 422 communicatively coupled to a cropping module 424, which is communicatively coupled to an embedding module 426 and a cart evaluation module 428. The embedding module 426 is further communicatively coupled to an expert system 430, which is also communicatively coupled to an embedded database 432 and a product database 434. Each of these and their operations are described in more detail below.
[0099] The input of the object detection module 422 is video frames from video footage captured by a video camera located within a predetermined distance of a SCO terminal (not shown) in a retail store (not shown).
[0100]
number
[0101] The predetermined distance is empirically determined depending on the layout of the retail store (not shown) and the SCO terminals (not shown) present therein, and allows for detection of products being scanned at the SCO terminal (not shown) or shopping carts approaching the SCO terminal (not shown). The output from the object detection module 422 is a video frame
[0102]
number
[0103] Each object (Obj) that appears in the i ) position (Loc(Obj i )) and their corresponding class labels. Thus, the output of the object detection module 422 is
[0104]
number
[0105] Contains the positions and labels of all objects that appear in the
[0106] Therefore, for a given video frame
[0107]
number
[0108] For a given video frame, the object detection module 422 is configured to determine the coordinates of a bounding box that substantially encloses the detected object within the video frame. The bounding box coordinates are established relative to the coordinate system of the received video frame. Specifically, for a given video frame,
[0109]
number
[0110] , the object detection module 422 calculates a bounding box
[0111]
number
[0112] configured to output one or more details of the set of
[0113]
number
[0114] is the video frame
[0115]
number
[0116] is the number of objects detected in
[0117]
number
[0118] is the bounding box surrounding the j-th detected product.
[0119] Each bounding box
[0120]
number
[0121] The details of include four variables, namely [x,y], h and w, where [x,y] are the time domains of the video frame.
[0122]
number
[0123] are the coordinates of the upper left corner of the bounding box relative to the upper left corner of the
[0124]
number
[0125] The details of are hereafter referred to as bounding box coordinates.
[0126] In one embodiment, the object detection module 422 includes a deep neural network whose architecture is substantially based on EfficientDet (described in M. Tan, R. Pang and Q.V. Le, EfficientDet: Scalable and Efficient Object Detection, 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2020, pp. 10778-10787). The deep neural network architecture may also be based on YOLOv4 (described in A. Bochkovskiy, C.Y. Wang and H.Y. M. Liao, 2020 arXiv: 2004.10934). However, those skilled in the art will recognize that these neural network architectures are provided for illustrative purposes only. In particular, those skilled in the art will understand that the preferred embodiment is not limited to these neural network architectures. Conversely, the preferred embodiment is operable with any deep neural network architecture and / or training algorithm suitable for detecting and localizing objects in video frames. For example, preferred embodiments are operable with region-based convolutional neural networks (RCNN), Faster-RCNN, or single-shot detectors (SSD).
[0127] The goal of training the deep neural network of the object detection module 422 is to establish an internal representation of the object so that the deep neural network can recognize the presence of the object in subsequently received video footage. To this end, the dataset used to train the deep neural network of the object detection module 422 includes multiple video frames captured by a video camera positioned within a predetermined distance from a SCO terminal (not shown) within a retail store. The predetermined distance is empirically determined depending on the layout of the retail store and the SCO terminals (not shown) located therein, and enables detection of products being scanned at the SCO terminal (not shown) and shopping carts approaching the SCO terminal (not shown). The video frames are selected and compiled to provide robust, class-balanced information about the object of interest obtained from views of the object acquired at various positions and orientations relative to the video camera (not shown). For clarity, this dataset is hereinafter referred to as the training dataset.
[0128] Before being used in the training dataset, video frames with similar appearances are removed from it. Members of the training dataset are subjected to further data augmentation techniques (e.g., rotation, flipping, brightness changes) to generate more video frames, thereby increasing the size of the training dataset, preventing overfitting, normalizing the deep neural network model, balancing the classes in the training dataset, and synthetically generating new video frames that better represent the task at hand. In a further preprocessing step, each video frame of the training dataset is further provided with a bounding box, with each such bounding box positioned to substantially enclose an object appearing in the video frame. Each video frame is also appropriately provided with a class label of "product," "shopping cart," or "other" corresponding to each bounding box in the respective video frame. The class label "product" indicates that the detected object is a product in a retail store's inventory, rather than a customer's personal belongings, which may also appear in the video frame. The class label "shopping cart" indicates that the detected object is a shopping cart, which may have various degrees of fullness.
[0129] In one embodiment, the object detection module 422 is further configured to concatenate the bounding box coordinates of each object detected in the video frame and the corresponding label classification of the detected object to form a detected object vector. Specifically, the output from the object detection module 422 is one or more detected object vectors
[0130]
number
[0131] where the object detection module 422 is further configured to communicate this output to the cropping module 424.
[0132] In one embodiment, the cropping module 424 is communicatively coupled to the object detection module 422 and receives the detected object vectors therefrom.
[0133]
number
[0134] The cropping module 424 receives the video frame received by the object detection module 422.
[0135]
number
[0136] The cropping module 424 is further configured to receive from each received video frame a corresponding detected object vector with a class label of “product”.
[0137]
number
[0138] The cropping module 424 is configured to crop product crop regions whose perimeters are established by the bounding box coordinates of the product image. The cropping module 424 is further configured to resize each product crop region to the same predetermined size. The predetermined size, hereafter referred to as the "processed product image size," is empirically established as the size that provides optimal product recognition by the embedding module 426. The cropping module 424 is further configured to send the resulting product crop regions to the embedding module 426.
[0139] In one embodiment, the cropping module 424 extracts from each video frame of the video footage received from the video camera a corresponding detected object vector with a class label of “shopping card.”
[0140]
number
[0141] The cropping module 424 is further configured to crop each cart crop region whose perimeter is established by the bounding box coordinates of the product image. The cropping module 424 is further configured to resize each cart crop region to the same predetermined size. The predetermined size, hereafter referred to as the "processed product image size," is empirically established as the size that allows the shopping cart fullness to be optimally assessed by the cart evaluation module 428. The cropping module 424 is further configured to send the resulting cart crop regions to the cart evaluation module 428.
[0142] In one embodiment, the embedding module 426 has two distinct operational phases: an initial configuration phase and a runtime phase, as described below. The embedding module 120 uses a deep metric learning module, as reviewed in K. Musgrave, S Belongie, and S.-N. Li, A Metric Learning Reality Check (retrieved August 19, 2020, from https: / / arxiv.org / abs / 2003.08505), to learn a unique representation of each product in a retail store's inventory, in the form of an embedding vector, from video frames in which the products appear. This allows for the identification of products appearing in subsequently captured video frames. For simplicity, a video frame in which a product appears, or a portion thereof, will hereinafter be referred to as an "image." Thus, the deep metric learning module is configured to generate embedding data including embedding vectors in response to images in which products appear; if the images contain the same product, the embedding vectors will be close to each other (in the embedding space); if the images contain different products, the embedding vectors will be far apart, as measured by a similarity or distance function (e.g., dot-product similarity or Euclidean distance). The query image can then be validated based on a similarity or distance threshold in the embedding space.
[0143] Initial configuration stage of the embedded module 426
[0144] In one embodiment, during the initial configuration stage, the embedding module 426 is configured to identify the products p i One or more embedding vectors that form a unique representation of E i The network is trained to learn the following: (a) a training data preparation phase; (b) a network training phase. The initial construction stage therefore includes several distinct phases: a training data preparation phase; and a network training phase. These phases are implemented sequentially in a cyclical, iterative manner to train the embedding module 426. Each of these phases is described in more detail below.
[0145] Training data preparation phase
[0146] The dataset used to train the embedding module 426 includes multiple video frames depicting each product in a retail store's inventory. These video frames are captured by a video camera mounted above an SCO terminal (not shown) in the retail store (not shown). The video frames, hereafter referred to as the training dataset, are compiled with the goal of providing robust, class-balanced information about the target products from various views of the products acquired at various positions and orientations of the products relative to the video camera. The members of the training dataset are selected to create sufficient diversity to overcome subsequent product recognition challenges posed by changing lighting conditions, viewpoint variations, and, most importantly, intra-class variation.
[0147] Before being used in the training dataset, video frames with similar appearances are removed. Members of the training dataset may also be further subjected to data augmentation techniques (such as rotation, flipping, and brightness changes) to increase diversity, thereby increasing the robustness of the trained deep neural network in the embedding module 426. A polygonal region encompassing each product appearing in the video frame is cropped therefrom. The cropped region is resized to match the size of the processed product image, generating a cropped product image. Each cropped product image is also provided with a class label that identifies the corresponding product.
[0148] Model training phase
[0149] For simplicity, the deep neural network (not shown) in embedding module 426 is hereafter referred to as an "embedded neural network (ENN)." An ENN includes a deep neural network (e.g., ResNet, Inception, EfficientNet) in which the last layer or layers (which typically output a classification vector) are replaced with a linear normalization layer that outputs a unit-norm (embedding) vector of a desired dimension. The dimension is a parameter set when the ENN is created.
[0150] During the model training phase, positive and negative pairs of cropped product images are constructed from the training dataset. A positive pair contains two cropped product images with the same class label, and a negative pair contains two cropped product images with different class labels. For brevity, the resulting cropped product images are hereafter referred to as "paired cropped images." Paired cropped images are sampled according to a pair mining strategy (e.g., MultiSimilarity or ArcFace as outlined in R. Manmatha, C.-Y. Wu, A.J. Smoia, and P. Krahenbuhl, "Sampling Matters in Deep Embedded Learning," 2017 IEEE International Conference on Computer Vision (CCV2017) Venice, 2017, pp. 2859-2867, doi: 10.1109 / ICCV.2017.309). Next, a pairwise metric learning loss is calculated from the sampled paired video frames (as described in K. Musgrave, S. Belongie and S.-N. Lim, A Metric Learning Reality Check, 2020, https: / / arxiv.org / abs / 2003.08505). The weights of the ENN are then optimized using a backpropagation approach that minimizes the pairwise metric learning loss value.
[0151] Every pair of cropped images is processed by the ENN to generate a corresponding embedding vector. The resulting embedding vectors are organized in a pairwise fashion similar to the pair of cropped images. The resulting embedding vectors are stored in the embedding database 432. Thus, given an image of each product in the retail store's inventory, the trained ENN can generate the embedding vectors computed for each product. E iThe embedding data including the embedded vectors ( E i , Id i ) and all products in stock at retailers p i The corresponding identifier Idi of the
[0152] Runtime Phase of the Embedding Module 426
[0153] For clarity, execution time is defined as the normal business hours of the associated store. During execution time, the ENN (not shown) generates embedding vectors for each product appearing in video frames captured by a video camera located within a predetermined distance of an SCO terminal (not shown) in the retail store (not shown). Accordingly, the embedding module 426 is coupled to the cropping module 424 and receives crop regions therefrom. The query embedding data, including the embedding vectors generated by the ENN (not shown) in response to the received crop regions, is hereinafter referred to as the query embedding. QE The embedding module 426 is communicatively coupled to the expert system module 430 and is referred to as a query embedding module. QE to the expert system module 430.
[0154] In one embodiment, the expert system module 430 is coupled to the embedding module 426 and receives query embeddings generated by the ENN during the runtime operation of the embedding module 426. QE Receive.
[0155] The expert system module 430 implements query embedding QE When it receives the embedding vector , it queries the embedding database 432 and obtains the embedding vector . E i The expert system module 430 uses a similarity or distance function (e.g., dot product similarity or Euclidean distance) to obtain the query embedding. QEthe embedding vector E i The expert system module 430 uses a similarity or distance function (e.g., dot product similarity or Euclidean distance) to compare the query embedding. QE The embedding vector obtained is E i Compare with.
[0156] Query Embedding QE and the obtained embedding vector E i If the similarity with, exceeds a preconfigured threshold (Th), the query embedding QE The embedding vector E obtained by i It can be concluded that the value of the threshold (Th) parameter is established using a grid search method.
[0157] In one embodiment, the embedding database 432 is queried to obtain the embedding vector E i , the received query embedding QE The process of comparing with is continued until a match is found or until all embedding vectors from the embedding database 432 are found. E i This is repeated until the query is embedded. QE and the embedding vectors from the embedding database 432 E i If a match is found between E i is hereinafter referred to as a matching embedded ME. The expert system module 430 is further adapted to use the matching embedded ME to obtain a product identifier corresponding to the matching embedded ME from a product database 434, where the product identifier is the identifier of the product represented by the matching embedded ME. For simplicity, this product identifier is hereinafter referred to as a matching class label.
[0158] In one embodiment, the cart assessment module 428 is configured to receive the cart crop region from the crop module 424. The cart assessment module 428 is configured to implement a panoptic segmentation algorithm, such as a bidirectional aggregation network (BANet) (described in Y. Chen, G. Lin, S. Li, O. Bourahla, Y. Wu, F. Wang, J. Feng, M. Xu, X. Li, Banet: Bidirectional aggregation network with occlusion handling for panoptic segmentation, in: Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 3793-3802), to establish a class label and an instance number for each pixel of the cart crop region.
[0159] In one embodiment, those skilled in the art will recognize that the above examples of panoptic segmentation algorithms are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the preferred embodiment is not limited to the above algorithms. On the contrary, the preferred embodiment can operate with any algorithm suitable for combining instance segmentation and semantic segmentation of cart crop regions, such as AuNet (described in Y. Li, X. Chen, Z. Zhu, L. Xie, G. Huang, D. Du, X. Wang, Attention-guided unified network for panoptic segmentation, in: Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 7026-7035) or the EfficientPS network (described in R. Mohan and A. Valada, EfficientPS: Efficient Panoptic Segmentation, International Journal of Computer Vision, 2021, 129(5), 1551-1579).
[0160] More specifically, the objectives of the panoptic segmentation algorithm are to: Identify full, partially full, or empty shopping carts and products from retail store inventory within the cart cutout area. Identify all instances of full, partially full, or empty shopping carts and products within the cart snippet area. Identify all products that appear in the cart crop area. Identify all instances of a product identified in the cart snippet.
[0161] The dataset used to train the panoptic algorithm includes multiple video frames depicting each product in inventory at a retail store. The dataset also includes multiple video frames depicting shopping carts with varying degrees of fullness. Specifically, the dataset includes video frames depicting empty shopping carts, partially full shopping carts, and completely full or overflowing shopping carts. The video frames are captured by a video camera mounted above an SCO terminal (not shown) at a retail store (not shown) and a video camera mounted within a predetermined distance of the SCO terminal (not shown). The predetermined distance is empirically determined according to the layout of the retail store (not shown) and the specific parameters of the video camera, and the field of view of the video camera encompasses the approach area to the SCO terminal (not shown) at the retail store.
[0162] Video frames, hereafter referred to as the training dataset, are compiled with the goal of providing robust, class-balanced information about a target product from different views of the product obtained at various positions and orientations of the product relative to the video camera. Video frames in the training dataset are further compiled to provide robust, class-balanced information about shopping carts at various positions and orientations of the shopping cart relative to the video camera, with various degrees of fullness. Members of the training dataset are selected to create sufficient diversity to overcome subsequent product and shopping cart recognition challenges posed by changing lighting conditions, changing viewpoints, and, most importantly, intra-class variation.
[0163] Before being used in the training dataset, video frames with similar appearances are removed from it. Members of the training dataset are subjected to further data augmentation techniques (rotation, flipping, brightness changes, etc.) to increase diversity and thereby increase the robustness of the trained neural network of the panoptic segmentation algorithm. Polygonal regions surrounding individual shopping carts appearing in the video frames are cropped from them. The cropped regions are manually resized from the size of the processed cart image to the cropped shopping cart image.
[0164] Each pixel in each cropped shopping cart image in the training dataset is manually tagged with a class label that identifies the corresponding product or whether the shopping cart is empty, partially full, or full. Each pixel is also tagged with an instance number that indicates which instance of a given class the pixel corresponds to. For example, the instance number may indicate whether the pixel corresponds to the second ice cream container appearing in the video frame or the third pack of toilet paper appearing in the video frame. The members of the training dataset are organized in a pairwise manner, with each data pair including a video frame and a corresponding XML file. The XML file includes the class label and instance number for each pixel in the corresponding video frame.
[0165] During training of the neural network model of the panoptic segmentation algorithm, individual members of the training dataset and corresponding entries from the XML file are presented to the neural network model, with the goal of building a representation of large-scale and small-scale features and contextual features sufficient to reproduce the presented members of the training dataset and corresponding entries from the XML file.
[0166] At runtime, the trained panoptic segmentation algorithm is presented with the cart clipping region received from the crop module 424. The panoptic segmentation algorithm labels each pixel in the cart clipping region that corresponds to the area of the shopping cart that appears therein as an empty shopping cart, a partially full shopping cart, or an empty shopping cart. If multiple shopping carts are present in the cart clipping region, the panoptic segmentation algorithm labels each pixel in the cart clipping region that corresponds to the area of the shopping cart that appears therein with the shopping cart instance number. The panoptic segmentation algorithm labels pixels in the cart clipping region that correspond to the area of a product that appears therein with the product class label and that product instance number. Output from the cart evaluation module 428 includes the pixels of the cart clipping region and their labels. For simplicity, this output is hereinafter referred to as "cart-related data."
[0167] FIG. 7 illustrates a processing unit 416 of the control unit (not shown) of the software of FIG. 4 that manages the SCO workspace of a retail store, according to one embodiment of the present disclosure.
[0168] In one embodiment, the processing unit 416 includes a non-scan event detector 436, a SCO administrator locator module 438, a queue analyzer module 440, a product movement analyzer 442, a SCO group analyzer module 444, a non-scan sequence analyzer module 446, a cart fullness analyzer module 448, and a customer group analyzer module 450. Each of these and their operation are described in more detail below.
[0169] In one embodiment, the non-scanning event detector 436 is communicatively coupled to the video unit 402, the SCO unit 404, the motion detection module 412, and the object recognition module 414. Specifically, the non-scanning event detector 436 is communicatively coupled to the motion detection unit 412 and receives a motion trigger signal therefrom indicating that motion has been detected within a predetermined distance of the SCO terminal (not shown), where the predetermined distance is determined by the intrinsic parameters of a video camera (not shown) mounted above the SCO terminal (not shown) and an installation height for establishing the field of view of the video camera (not shown). The received motion trigger signal indicates that a customer is approaching the SCO terminal (not shown) and scanning a product with the SCO terminal (not shown).
[0170] Upon receiving the motion trigger signal, the non-scan event detector 436 extracts successive video frames from the video footage captured by a video camera mounted above the SCO terminal (not shown).
[0171]
number
[0172] and,
[0173]
number
[0174] from the video unit 402. The non-scan event detector 436 is configured to receive successive video frames
[0175]
number
[0176] and,
[0177]
number
[0178] to the object recognition module 414 to detect the presence of a product in the retail store's inventory in the video frames. Upon detecting the presence of a product in the retail store's inventory in the received video frames and recognizing the product, the object recognition module 414 is configured to return a corresponding matching class label to the non-scan event detector 436. The matching class label is an identifier for the recognized product. More specifically, the matching class label may be the UPC of the recognized product.
[0179] Upon receiving the motion trigger signal, the non-scan event detector 436 is also configured to receive sales register data from the SCO unit 404. The received sales register data originates from an SCO terminal (not shown) where the motion indicated by the motion trigger signal was detected. Upon receiving the matching class label, the non-scan event detector 436 is configured to compare the matching class label with sales register data received within a time interval of a predetermined period occurring before and after receiving the matching class label. The predetermined period is empirically determined to be long enough to find a match between the matching class label and members of the sales register data associated with products scanned at the SCO terminal (not shown) during the time interval without delaying operation of the SCO terminal (not shown).
[0180] If the received sales register data and the received matching class label do not match, the non-scan event detector 436 is configured to issue a non-scan alert signal that includes an identifier of the SCO terminal (not shown) where the motion indicated by the motion trigger signal was detected. For simplicity, the identifier is hereinafter referred to as the originating SCO identifier, and the SCO terminal (not shown) corresponding to the originating SCO identifier is hereinafter referred to as the "alert-originating SCO."
[0181] In one embodiment, the SCO manager locator module 438 is communicatively coupled to the non-scan event detector 436, the human classification module 408, and the human tracking module 410. Specifically, the SCO manager locator module 438 is configured to receive a non-scan alert signal from the non-scan event detector 436. Upon receiving the non-scan alert signal, the SCO manager locator module 438 is further configured to activate the human classification module 408 and the human tracking module 410 to determine the locations of all SCO managers in the retail store. Using the location information, the SCO manager locator module 438 is further configured to calculate the distance between each SCO manager and the alert-issuing SCO.
[0182] In one embodiment, the SCO administrator locator module 438 is further configured to detect whether there are adult customers or children within a predetermined distance of each SCO administrator, the predetermined distance being empirically determined as the maximum expected distance between the staff member and the customers and / or children when the staff member is assisting the customers and / or children.
[0183] If the SCO administrator locator module 438 determines that the SCO administrator (not shown) is located within a predetermined distance of the adult customer and / or child, the SCO administrator locator module 438 is configured to activate the people tracking module 410 to track the movements of the SCO administrator (not shown) and the adult customer and / or child at predetermined time intervals. The predetermined time interval is empirically determined to be the expected minimum duration of engagement between a staff member and the customer and / or child when the SCO administrator (not shown) is assisting the customer and / or child. The purpose of tracking the movements of the SCO administrator (not shown) and the adult customer and / or child at predetermined time intervals is to exclude situations in which the SCO administrator (not shown) is accidentally near the adult customer and / or child rather than active engagement between the SCO administrator (not shown) and the adult customer and / or child.
[0184] If the SCO administrator locator module 438 determines that the SCO administrator (not shown) is located within a predetermined distance of adult customers and / or children for more than a predetermined time interval, the SCO administrator locator module 438 assigns a "busy" status tag to the SCO administrator (not shown). Assigning this status tag to the SCO administrator (not shown) may cause the SCO administrator locator module 438 to disable tracking of the SCO administrator (not shown) and nearby adult customers and / or children.
[0185] In one embodiment, the SCO administrator locator module 438 is further configured to activate the people tracking module 410 to track the movements of the remaining SCO administrators at predetermined time intervals. The predetermined time intervals are empirically determined to be a time interval sufficient to determine whether the SCO administrator is moving from one part of the retail store to another, but not so long as to unduly slow the operation of the SCO administrator locator module 438. If the SCO administrator locator module 438 determines that the SCO administrator (not shown) is moving toward a stockroom, a checkout room, or the like, the SCO administrator locator module 438 assigns a "busy" status tag to the SCO administrator (not shown).
[0186] In one embodiment, the SCO administrator locator module 438 is further configured to identify a SCO administrator (not shown) that is not tagged as "busy" and is closest to the alert-emitting SCO. If the identified SCO administrator (not shown) is determined to be located less than a predetermined distance from the alert-emitting SCO, the SCO administrator locator module 438 is configured to emit an output signal O1 that includes a "SCO locked" signal. Otherwise, the output signal O1 includes a "void" signal.
[0187] In one embodiment, the predetermined distance from the alert-emitting SCO is empirically determined depending on the retail store layout (not shown) and is the maximum distance that allows a SCO administrator to return to a locked SCO terminal (not shown) to identify the cause of the non-scan alert signal and, if necessary, unlock the SCO terminal (not shown).
[0188] In one embodiment, the queue analyzer module 440 is communicatively coupled to the non-scan event detector 436 and the person classification module 408. Specifically, the SCO queue analyzer module 440 is configured to receive a non-scan alert signal from the non-scan event detector 436. Upon receiving the non-scan alert signal, the queue analyzer module 440 is further configured to activate the person classification module 408 to calculate the number of adult customers and children located within a predetermined distance approaching the alert-emitting SCO. The queue analyzer module 440 is further configured to compare the locations of the adult customers and children located within a predetermined distance approaching the alert-emitting SCO to determine whether at least some of the adult customers and children are arranged in a queue pattern upon approaching the alert-emitting SCO.
[0189] If the queue analyzer module 440 determines that at least some adult customers and children are arranged in a queue pattern upon approaching the alert-emitting SCO, the queue analyzer module 440 is configured to calculate the number of adult customers and children in the queue. If the number of adult customers and children in the queue is less than a predetermined threshold, the queue analyzer module 440 is configured to emit an output signal O2 that includes a "SCO locked" signal. Otherwise, the output signal O2 includes an "empty" signal.
[0190] The predetermined approach distance to the warning SCO and the threshold number of people in line at the warning SCO are empirically determined based on the operator's understanding of the potential for revenue loss due to customer hesitation caused by excessively long lines, while balancing the risk of revenue loss due to non-payment of products at the warning SCO.
[0191] In one embodiment, the product movement analyzer 442 is communicatively coupled to the non-scan event detector 436, the human classification module 408, the human tracking module 410, and the object recognition module 414. Specifically, the product movement analyzer 442 is configured to receive a non-scan alert signal from the non-scan event detector 436. Upon receiving the non-scan alert signal, the product movement analyzer 442 is further configured to activate the human classification module 408 and the human tracking module 410 to receive therefrom extracted features of an adult customer or child that was closest to the alert-emitting SCO immediately upon issuance of the received non-scan alert signal. Upon receiving the non-scan alert signal, the product movement analyzer 442 is further configured to activate the object recognition module 414 to recognize and issue an identifier for a product (not shown) that was closest to the alert-emitting SCO immediately upon issuance of the received non-scan alert signal. The product movement analyzer 442 is configured to temporarily store the extracted features received from the human tracking module 410 and the product identifier received from the object recognition module 414. The product movement analyzer 442 is further configured to activate the object recognition module 414 again after a predetermined time interval to determine the location of a product having an identifier matching the stored product identifier, which for simplicity is hereinafter referred to as the "non-scan query product."
[0192] In one embodiment, the product movement analyzer module 442 is further configured to again activate the person classification module 408 and the person tracking module 410 to receive extracted features from the adult customer or child closest to the non-scanned query product. For simplicity, this adult customer or child is hereinafter referred to as the “non-scanned query person.” The product movement analyzer 442 is further configured to compare the extracted features of the non-scanned query person with stored extracted features. If no match is found between the extracted features of the non-scanned query person and the stored extracted features, the product related to the non-scanned event is considered to have changed hands and is now owned by another person. The movement of the product between people immediately after the non-scanned event suggests intent by the person involved in the non-scanned event. Therefore, the product movement analyzer 442 is configured to emit an output signal O3 including a “SCO locked” signal. Otherwise, the output signal O3 includes an “empty” signal.
[0193] Social engineering of customers at SCO terminals using nudge theory involves two main elements: freezing or locking the SCO terminal when a non-scan event is detected, inconveniencing the customer with the associated delays, and customer interaction with an SCO administrator investigating the non-scan event, which may also inconveniencing the customer. Both of these have a deterrent effect on would-be thieves by shifting the perceived balance between the risks of detection and the rewards of theft. However, the time spent by SCO administrators at the SCO terminal and those involved in the non-scan event also incurs costs to the vendor, time that could be better utilized elsewhere in the retail store. Sales are also lost when customers become frustrated by delays or long lines at the SCO terminal and leave.
[0194] The challenge of managing this balance is exacerbated in retail stores where multiple SCO terminals are operating in parallel, because SCO administrators can only handle locked episodes of SCO terminals sequentially. Using an analogy from fault management, a locked episode of a SCO terminal, even if it is an intentionally generated fault, can be considered a fault in the sequential operation of the SCO terminal. Based on this analogy, the separation between parallel fault generation and sequential fault resolution becomes particularly acute as the number of sources of such faults increases; for example, the number of SCO terminals used during busy periods increases compared to quiet periods.
[0195] The balance between the vendor's two competing objectives and the decoupling impact between parallel fault generation and sequential fault resolution can be addressed by a three-threshold system. The first threshold is based on the number of locked SCO terminals that the SCO administrator can address in a given period of time. The second threshold is based on the observation that people's frustration with waiting in line often increases depending on the amount of time they have already spent in line. Thus, the second and third thresholds address the length of time that individual SCO terminals remain locked. The values of these three thresholds can be adjusted by the vendor depending on their risk tolerance for revenue loss resulting from theft at SCO terminals and their knowledge of customer tolerance for delays, recognizing customer patterns and profiles that may vary over time.
[0196] Accordingly, the SCO group analysis module 444 is communicatively coupled to the non-scan event detector 436 and the SCO unit 404. Specifically, the SCO group analysis module 444 is configured to receive a non-scan alert signal from the non-scan event detector 436. Upon receiving the non-scan alert signal, the SCO group analysis module 444 is also configured to receive a status signal from each SCO terminal (not shown) coupled to the SCO unit 404. The SCO group analysis module 444 is further configured to calculate the number of locked SCO terminals from the received status signals. The SCO group analysis module 444 is further configured to calculate the duration for which each locked SCO terminal (not shown) has been locked. For simplicity, the SCO terminal (not shown) that has been locked for the longest period of time will be referred to as the "senior-locked SCO terminal." Similarly, the duration for which a senior-locked SCO terminal has been locked will hereinafter be referred to as the "senior-locked duration."
[0197] In one embodiment, the SCO group analysis module 444 is further configured to compare the number of locked SCO terminals (not shown) with a first threshold, and if the number of locked SCO terminals (not shown) is less than the first threshold, the SCO group analysis module 444 is configured to emit an output signal O4 including a SCO lock signal. Alternatively or additionally, the SCO group analysis module 444 is further configured to compare a senior lock period with a second threshold, and if the senior lock period is less than the second threshold, the SCO group analysis module 444 is configured to emit an output signal O4 including a SCO lock signal. Alternatively or additionally, the SCO group analysis module 444 is further configured to calculate the number of SCO terminals (not shown) that have been locked for a period exceeding the second threshold, and if the number of SCO terminals (not shown) is less than a third threshold, the SCO group analysis module 444 is configured to emit an output signal O4 including a "SCO" signal. Otherwise, the output signal O4 includes an "empty" signal.
[0198] In one embodiment, the non-scan sequence analyzer module 446 is communicatively coupled to the non-scan event detector 436 and the SCO unit 404. Specifically, the non-scan sequence analyzer module 446 is configured to receive a non-scan alert signal from the non-scan event detector 436. Upon receiving the non-scan alert signal, the non-scan sequence analyzer module 446 is configured to store an outgoing SCO identifier and a timestamp of the non-scan alert signal. The non-scan sequence analyzer module 446 is further configured to compare the outgoing SCO identifier of a subsequently received non-scan alert signal with the stored outgoing SCO identifier to identify a match. If a match is found, the non-scan sequence analyzer module 446 is configured to compare the timestamp of the subsequently received non-scan alert signal with the stored timestamp corresponding to the matching stored outgoing SCO identifier. For simplicity, the elapsed time between the timestamp of the subsequently received non-scan alert signal and the stored timestamp corresponding to the matching stored outgoing SCO identifier is hereinafter referred to as the "time since last non-scan alert." If the time elapsed since the last no-scan warning is less than a predetermined threshold, the no-scan sequence analyzer module 446 is configured to emit an output signal O5 that includes a "SCO locked" signal. Otherwise, the output signal O5 includes an "empty" signal.
[0199] In one embodiment, the cart fullness analyzer module 448 is communicatively coupled to the non-scan event detector 436 and the object recognition module 414. Specifically, the cart fullness analyzer module 448 is configured to receive a non-scan alert signal from the non-scan event detector 436. Upon receiving the non-scan alert signal, the cart fullness analyzer module 448 is further configured to activate the object recognition module 414 to recognize the presence of a shopping cart near the alert-emitting SCO. Specifically, the cart fullness analyzer module 448 is configured to receive cart-related data from a cart evaluation module (not shown) of the object recognition module 414.
[0200] In one embodiment, the cart-related data includes pixels of areas occupied by the shopping cart and products contained therein that appear in video frames received from a video camera mounted above the alert-emitting SCO. The cart-related data also includes a class label and an instance label for each pixel. The cart fullness analyzer module 448 is configured to calculate the percentage of pixels in the cart-related data that are labeled as "full shopping cart," "partially full shopping cart," or "empty shopping cart." The cart fullness analyzer module 448 is configured to set a cart state variable to a value of "full" if a majority of the pixels in the cart-related data are labeled as "full shopping cart." Similarly, the cart fullness analyzer module 448 is configured to set a cart state variable to a value of "partially full" if a majority of the pixels in the cart-related data are labeled as "partially full shopping cart." Similarly, the cart fullness analyzer module 448 is configured to set a cart state variable to a value of "empty" if a majority of the pixels in the cart-related data are labeled as "empty shopping cart."
[0201] If the cart state variable is set to a value of "full," the cart fullness analyzer module 448 is configured to count the number of instances of each displayed product included in the shopping cart. If the number of instances of the displayed product included in the shopping cart exceeds a predetermined threshold, the cart fullness analyzer module 448 is configured to add the label "bulk loaded" to the cart state variable.
[0202] The predetermined threshold is empirically determined according to the operator's insight, experience, and past knowledge regarding situations in which a thief may attempt to conceal the theft of a product by placing the product in a shopping cart containing other products, particularly when the other products are essentially identical.
[0203] In one embodiment, the cart fullness analyzer module 448 is further configured to count the number of pixels in the cart-related data that are labeled with the same product class label and instance number. The number of counted pixels provides an initial approximation of the visible area of the corresponding product. For simplicity, a product whose single instance forms the majority of the number of pixels in the cart-related data that are not labeled as "full shopping cart," "partially full shopping cart," or "empty shopping cart" is hereinafter referred to as the "most visible product."
[0204] In one embodiment, the cart fullness analyzer module 448 is communicatively coupled to a product details database (not shown). The product details database (not shown) contains details of the quantity of each product in the retail store's inventory. The cart fullness analyzer module 448 is configured to query the product details database (not shown) to retrieve a record corresponding to the largest visible product. Thus, the retrieved record contains details of the quantity of the largest visible object contained in the shopping cart located near the alert-emitting SCO. For simplicity, this quantity is hereinafter referred to as the "largest quantity instance." When the largest quantity instance exceeds a predetermined threshold, the cart fullness analyzer module 448 is configured to add the label "large item" to the cart state variable. The predetermined threshold is empirically determined according to the operator's insight, experience, and past knowledge of the typical size range of products sold at the retail store. The cart fullness analyzer module 448 is configured to emit an output signal O6 that includes the cart state variable.
[0205] FIG. 8 illustrates a table of outputs from a processing unit of the control unit of the software of FIG. 4 that manages the SCO workspace of a retail store, according to one embodiment of the present disclosure.
[0206] Thus, the processing unit 416 of the control unit 406 is configured to send processed output signals including O1, O2, O3, O4, O5, and O6 to the logic unit 418 of FIG. 4, as shown in the table of FIG. 8.
[0207] 4, logic unit 418 of control unit 406 is configured to receive the processed output signal from processing unit 416 of control unit 406. Logic unit 418 includes a plurality of Boolean logic units (not shown) operable to verify the contents of one or more of the O1, O2, O3, O4, O5, and O6 components of the processed output signal and issue a SCO lock command along with an outgoing SCO identifier of the SCO terminal (not shown) on which the non-scan event was detected. For simplicity, the SCO lock command and the outgoing SCO identifier are collectively referred to as the "SCO lock control signal."
[0208] The Boolean logic unit (not shown) may be configured according to the requirements of the operator. However, in one example, the Boolean logic unit (not shown) is configured such that if any one of the O1, O2, O3, O4, and O5 components of the processed output signal has a value “SCO Lock,” then the logic unit 418 issues a SCO Lock control signal to the SCO unit 404. Similarly, in another example, the Boolean logic unit (not shown) is configured such that if the O6 component of the processed output signal has a value “Full” or “Full Bulk Loaded,” then the logic unit 418 issues a SCO Lock control signal to the SCO unit 404.
[0209] Those skilled in the art will recognize that the above example configurations of Boolean logic units (not shown) are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the software of the preferred embodiment is not limited to the above example configurations of Boolean logic units (not shown). Rather, the software of the preferred embodiment is operable with any configuration of Boolean logic units (not shown) suitable for processing one or more components of the processed output signal received from the processing unit 416 of the control unit 406.
[0210] In one embodiment, the SCO unit 404 receives a SCO lock control signal from the logic unit 418 of the control unit 406 and locks the SCO from a SCO terminal SCO1 in the retail store indicated by the originating SCO identifier of the SCO lock control signal. n configured to lock the
[0211] 9, a method 900 for managing one or more SCO terminals in a SCO environment is shown. The method includes: n , SCO1 to SCO2. The video sensors may be located at one or more SCO terminals SCO1 to SCO3. n above, or from SCO terminal SCO1 to SCO n One or more video cameras C1 to C2 installed at a predetermined distance from n Includes.
[0212] In one embodiment, the method 900 includes: n The state data includes a next step 904 of acquiring state data for each of the SCO terminals SCO1 to SCO2. n The status signal is sent from SCO terminal SCO1 to SCO n It may further include a timestamp of when it was locked.
[0213] In one embodiment, the method 900 includes a next step 906 of coupling a control unit to one or more video sensors and SCO units. The control unit 406 includes a processing unit and memory. The processing unit includes a processor, computer, microcontroller, or other circuitry that controls the operation of various components, such as memory. The processing unit may execute software, firmware, and / or other instructions stored in, for example, volatile or non-volatile memory, such as memory. The processing unit may be connected to the memory via a wired or wireless connection, such as one or more system buses, cables, or other interfaces.
[0214] In one embodiment, the method 900 includes a next step 908 of determining one or more frames of interest from the plurality of video frames using a machine learning model. In the same embodiment, determining the one or more frames includes detecting a primary subject of interest using the human classification module 408. In the same embodiment, the detected primary subject of interest is classified based on age group, i.e., child or adult. In the same embodiment, the method includes detecting one or more secondary subjects of interest after detecting the primary subject of interest.
[0215] In one embodiment, the method 900 uses the human tracking module 410 to track one or more SCO terminals SCO1 to SCO2. n , the next step 910 of determining the detection locations and times of the primary and secondary features of interest within a predetermined distance from the target.
[0216] In one embodiment, the method 900 uses the motion detection unit 412 to detect the motion of the SCO terminals SCO1 to SCO2. n The method includes a next step 912 of generating a motion trigger based on detecting a change in position of the primary and secondary objects of interest within a predetermined distance of one of the primary and secondary objects of interest.
[0217] In one embodiment, based on the generated motion trigger, the method 900 includes a next step 914 of receiving transaction data from the SCO unit 404, where the transaction data is received from the SCO terminal SCO1 to the SCO n The transaction includes a transaction received by scanning one or more secondary objects of interest with any one of the following:
[0218] In one embodiment, the method 900 includes a next step 916 of comparing transaction data from the SCO unit 404 to the detected secondary objects of interest. Additionally, the method 900 includes a next step 918 of generating a non-scan event alert based on a discrepancy in the comparison of the transaction data to the one or more detected secondary objects of interest.
[0219] Modifications to the embodiments of the present disclosure described above are possible without departing from the scope of the disclosure, which is defined by the appended claims. Words such as "including," "comprising," "incorporating," "consisting of," "having," "is," and the like, as used to describe and claim the present disclosure, are intended to be construed in a non-exclusive manner, i.e., there may be items, components, or elements present that are not expressly described. References to the singular should also be construed to relate to the plural.
Claims
1. 1. A method for managing one or more self-checkout (SCO) terminals in a SCO environment, comprising: capturing a plurality of video frames from one or more of said SCO terminals using one or more video sensors installed at predetermined locations; acquiring status data for each of one or more of said SCO terminals using an SCO unit communicatively coupled to each of said SCO terminals; coupling a control unit to one or more of the video sensors and the SCO unit, the control unit comprising a processing unit connected to a memory, the memory comprising: determining one or more frames of interest from the plurality of video frames using a machine learning model, wherein determining the one or more frames includes: Detecting a primary subject of interest using a human classification module; classifying the detected primary subjects of interest based on an age group of the primary subjects of interest; detecting one or more secondary objects of interest after detecting the primary object of interest; determining detection locations and detection times of the primary and secondary objects of interest within a predetermined distance from one or more of the SCO terminals using a human tracking module; generating a motion trigger based on detecting a change in location of the primary and secondary objects of interest within a predetermined distance of any one of the SCO terminals using a motion detection unit; Based on the generated motion trigger, receiving transaction data from the SCO unit, the transaction data including transactions received by scanning one or more of the secondary objects of interest; comparing the transaction data from the SCO unit with the detected one or more secondary objects of interest; generating a non-scan event alert based on a discrepancy in comparing the transaction data with the detected one or more secondary objects of interest; including, a set of instructions executed by the processing unit for: A method for providing the above.
2. classifying the detected subject of interest based on the age group of the subject of interest includes classifying the subject of interest as a SCO administrator, a child, or an adult; the secondary object of interest is associated with the detected primary object of interest, the secondary object of interest comprising a shopping cart or a stack of items; The method of claim 1.
3. The human tracking module: identifying a physical feature of the primary subject of interest in one of the plurality of video frames; forming person identification data based on the physical characteristics of the identified primary subject of interest; linking a biometric signature of said primary subject of interest to said person identifying data; storing said person identification data in an internal repository of said human tracking module; further configured as follows: The method of claim 1.
4. The human tracking module: forming query specific data for the primary subject of interest in consecutive frames of a plurality of the video frames; comparing the query-specific data of the primary subject of interest with the person-specific data stored in the internal repository; Based on the comparison, if the query identifying data does not match the person identifying data, assigning new person identifying data to the primary subject of interest in another video frame; further configured as follows: The method of claim 3.
5. the control unit further comprises an object recognition module; The object recognition module an object detection module, Detecting an object as the secondary object of interest; forming object detection data associated with the detected object; an object detection module configured as follows: a cropping module, processing a plurality of the video frames based on the object detection data; and cropping regions of the plurality of the video frames to form product crop regions; a cutting module configured as follows: Including, The method of claim 1.
6. the cutting module is further connected to an embedding unit; The embedding unit comprises: forming embedded data for said object in inventory of said SCO environment; receiving the product cut-out area; generating query embedding data in response to the received product cutout area; comparing the query embedded data and the embedded data; determining a match between the query embedding data and the embedding data if the similarity between the query embedding data and the embedding data exceeds a predetermined threshold; It is configured as follows: The method of claim 5.
7. The control unit SCO administrator locator module, Receive non-scan event alerts, determining a distance of a SCO administrator within a predetermined distance from the SCO terminal at which the non-scan event alert was generated; locking the SCO terminal if the distance of the SCO administrator is greater than a predetermined distance for a predetermined time interval; a SCO administrator locator module configured to: Further provided with The method of claim 1.
8. 1. A system for managing one or more self-checkout (SCO) terminals in a SCO environment, comprising: one or more video sensors positioned at predetermined locations from one or more of said SCO terminals configured to capture a plurality of video frames; an SCO unit communicatively coupled to each of one or more of said SCO terminals configured to obtain status data for each of said SCO terminals; a control unit coupled to the one or more video sensors and the SCO unit, the control unit comprising a processing unit connected to a memory, the memory including a set of instructions executed by the processing unit to determine one or more frames of interest from a plurality of the video frames using a machine learning model, wherein determining the one or more frames of interest comprises: Detecting a primary subject of interest using a human classification module; classifying the detected primary subjects of interest based on an age group of the primary subjects of interest; detecting one or more secondary objects of interest after the appearance of the primary object of interest; determining, using a human tracking module, the locations and times of appearance of the primary and secondary objects of interest within a predetermined distance from one or more of the SCO terminals; generating a motion trigger based on detecting a change in location data of the primary and secondary objects of interest within a predetermined distance of any one of the SCO terminals using a motion detection unit; Based on the generated motion trigger, receiving transaction data from the SCO unit, the transaction data including transactions received by scanning one or more of the secondary objects of interest; comparing the transaction data from the SCO unit with the detected secondary objects of interest using a non-scan event detector; generating a non-scan event alert based on a discrepancy in comparing the transaction data with the detected one or more secondary objects of interest; a control unit including: A system comprising:
9. classifying the detected subject of interest based on the age group of the subject of interest includes classifying the subject of interest as a SCO administrator, a child, or an adult; the secondary object of interest is associated with the detected primary object of interest, the secondary object of interest comprising one of a shopping cart or a stack of items; The system of claim 8.
10. The human tracking module: identifying a physical feature of the primary subject of interest in one of the plurality of video frames; forming person identification data based on the physical characteristics of the identified primary subject of interest; linking a biometric signature of said primary subject of interest to said person identifying data; storing said person identification data in an internal repository of said human tracking module; further configured as follows: The system of claim 8.
11. The human tracking module: forming query specific data for the primary subject of interest in consecutive frames of a plurality of the video frames; comparing the query-specific data of the primary subject of interest with the person-specific data stored in the internal repository; Based on the comparison, if the query identifying data does not match the person identifying data, assigning new person identifying data to the primary subject of interest in another video frame; further configured as follows: The system of claim 8.
12. the control unit further comprises an object recognition module; The object recognition module an object detection module, Detecting an object as the secondary object of interest; forming object detection data associated with the detected object; an object detection module configured as follows: a cropping module, processing a plurality of the video frames based on the object detection data; and cropping regions of the plurality of the video frames to form product crop regions; a cutting module configured as follows: Including, The system of claim 8.
13. the cutting module is further connected to an embedding unit; The embedding unit comprises: forming embedded data for said object in inventory of said SCO environment; receiving the product cut-out area; generating query embedding data in response to the received product cutout area; comparing the query embedded data and the embedded data; determining a match between the query embedding data and the embedding data if the similarity between the query embedding data and the embedding data exceeds a predetermined threshold; It is configured as follows: The system of claim 12.
14. The control unit SCO administrator locator module, Receive non-scan event alerts, determining a distance of a SCO administrator within a predetermined distance from the SCO terminal at which the non-scan event alert was generated; locking the SCO terminal if the distance of the SCO administrator is greater than a predetermined distance for a predetermined time interval; a SCO administrator locator module configured to: Further provided with The system of claim 8.
15. 1. A non-transitory computer-readable medium storing a set of instructions, the set of instructions, when executed by a computer, causing the computer to: capturing a plurality of video frames using one or more video sensors installed at predetermined locations from one or more SCO terminals; acquiring status data for each of one or more of said SCO terminals using an SCO unit communicatively coupled to each of said SCO terminals; coupling a control unit to one or more of the video sensors and the SCO unit, the control unit comprising a processing unit connected to a memory, the memory comprising: determining one or more frames of interest from the plurality of video frames using a machine learning model, wherein determining the one or more frames includes: Detecting a primary subject of interest using a human classification module; classifying the detected primary subjects of interest based on an age group of the primary subjects of interest; detecting one or more secondary objects of interest after detecting the primary object of interest; determining detection locations and detection times of the primary and secondary objects of interest within a predetermined distance from one or more of the SCO terminals using a human tracking module; generating a motion trigger based on detecting a change in location of the primary and secondary objects of interest within a predetermined distance of any one of the SCO terminals using a motion detection unit; Based on the generated motion trigger, receiving transaction data from the SCO unit, the transaction data including transactions received by scanning one or more of the secondary objects of interest; comparing the transaction data from the SCO unit with the detected one or more secondary objects of interest; generating a non-scan event alert based on a discrepancy in comparing the transaction data with the detected one or more secondary objects of interest; including, a set of instructions executed by the processing unit for: performing operations for managing one or more self-checkout (SCO) terminals in an SCO environment, including: Non-transitory computer-readable medium.
16. The human tracking module: identifying a physical feature of the primary subject of interest in one of the plurality of video frames; forming person identification data based on the physical characteristics of the identified primary subject of interest; linking a biometric signature of said primary subject of interest to said person identifying data; storing said person identification data in an internal repository of said human tracking module; further configured as follows:
16. The non-transitory computer-readable medium of claim 15.
17. The human tracking module: forming query specific data for the primary subject of interest in consecutive frames of a plurality of the video frames; comparing the query-specific data of the primary subject of interest with the person-specific data stored in the internal repository; Based on the comparison, if the query identifying data does not match the person identifying data, assigning new person identifying data to the primary subject of interest in another video frame; further configured as follows:
17. The non-transitory computer-readable medium of claim 16.
18. the control unit further comprises an object recognition module; The object recognition module an object detection module, Detecting an object as the secondary object of interest; forming object detection data associated with the detected object; an object detection module configured as follows: a cropping module, processing a plurality of the video frames based on the object detection data; and cropping regions of the plurality of the video frames to form product crop regions; a cutting module configured as follows: Including, 16. The non-transitory computer-readable medium of claim 15.
19. the cutting module is further connected to an embedding unit; The embedding unit comprises: forming embedded data for said object in inventory of said SCO environment; receiving the product cut-out area; generating query embedding data in response to the received product cutout area; comparing the query embedded data and the embedded data; determining a match between the query embedding data and the embedding data if the similarity between the query embedding data and the embedding data exceeds a predetermined threshold; It is configured as follows:
20. The non-transitory computer-readable medium of claim 18.
20. The control unit SCO administrator locator module, Receive non-scan event alerts, determining a distance of a SCO administrator within a predetermined distance from the SCO terminal at which the non-scan event alert was generated; locking the SCO terminal if the distance of the SCO administrator is greater than a predetermined distance for a predetermined time interval; a SCO administrator locator module configured to: Further provided with 16. The non-transitory computer-readable medium of claim 15.
Citation Information
Patent Citations
Scan-miss identification method and device, self-service cash register terminal and system
JP2021536619A
Systems and methods for detecting scan irregularities at self-checkout terminals
JP2022519191A
System and method for managing a retailer's SCO workspace
JP2022520239A
Skip-scanning identification method, apparatus, and self-service checkout terminal and system
US20210183212A1