Retail checkout using multi-signal batch product identification

The hybrid checkout system addresses the inefficiencies of traditional self-checkout and manned checkouts by using computer vision and RFID to provide a faster and more accurate batch product identification process.

JP2026052052APending Publication Date: 2026-03-23エヌ·シー·アール·ヴォイクス·コーポレイション
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025157305
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-10
Filing Date
2025-09-22
Publication Date
2026-03-23

AI Technical Summary

Technical Problem

Traditional self-checkout systems require a sequential product identification process that is cumbersome for consumers with many items, and computer vision-based systems can be inaccurate and burdensome for large baskets, while manned checkouts are inefficient for those avoiding contact.

Method used

A hybrid checkout system combining computer vision with multiple sensors, such as RFID, to generate multi-signal input for batch product identification, capturing images from multiple angles and integrating sensor data for improved accuracy.

Benefits of technology

The hybrid system transforms the sequential identification process into a faster and more accurate batch process, enhancing the checkout experience for both self-checkout and manned checkout users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026052052000001_ABST
    Figure 2026052052000001_ABST
Patent Text Reader

Abstract

We provide a hybrid checkout system that enables on-the-go, multi-signal, and batch product identification. [Solution] The system includes a computer vision device with multiple cameras fixed at different positions on a conveyor belt, each camera capturing multiple images of products placed on the conveyor belt, including markings to assist the user in placing products, from different viewpoints as the products move. Furthermore, an RFID sensor provided in the hybrid checkout device acquires RFID data from RFID tags attached to products as the products move. Image data and other sensor data are provided as multi-signal inputs to a machine learning model trained to recognize products and output product identifiers for the products. Price information corresponding to the product identifiers received from the model is determined, and the products and their respective prices are added to the transaction record.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a product identification technology in a retail store, and more particularly, to a method and apparatus for collectively identifying products in motion by computer vision and multi-signal input from a plurality of sensors.

Background Art

[0002] In recent years, due to various factors, the introduction of self-checkout by consumers has been accelerating. These factors include the COVID-19 pandemic and the associated risk of disease transmission during transactions at manned cash registers. Furthermore, due to the labor shortage that has worsened due to the pandemic and continues thereafter, retailers have had to reduce the number of available manned lanes, thereby forcing consumers who would otherwise have used manned lanes to complete transactions at self-checkout terminals instead. Apart from the pandemic-related reasons mentioned above, many consumers prefer self-checkout because they view it as a more efficient way to complete transactions, particularly when the number of items to be purchased is limited or when the queues at manned lanes are long.

[0003] On the other hand, some consumers are still negative about using self-checkout. Conventional self-checkout involves a sequential product identification process, similar to conventional manned cash registers, in which the operator is required to individually add each product in the basket of products to be purchased to the transaction record. In such a sequential product identification process, each product is identified sequentially, and the identification is based on, for example, barcode scanning of the product, product lookup (PLU) codes entered by keystrokes, or PLU codes determined based on the operator's selection from a set of candidate options. When a consumer has a large number of products to purchase, this sequential product identification process can become cumbersome and requires the consumer to individually input (e.g., barcode scan) each product.

Summary of the Invention

[0004] Traditional self-checkout systems involved a sequential product identification process, which was cumbersome when consumers had a large number of items. [Means for solving the problem]

[0005] This invention provides a hybrid checkout system that generates a multi-signal input by combining computer vision with other sensor signals. [Effects of the Invention]

[0006] This invention significantly improves transaction speed and provides more accurate product recognition capabilities. [Brief explanation of the drawing]

[0007] [Figure 1] This is a block diagram of a computer vision-based multi-signal hybrid checkout system architecture in an exemplary embodiment. [Figure 2] This figure shows a reinforced checkout conveyor belt in an exemplary embodiment. [Figure 3] This figure shows a computer vision device and reinforced conveyor belt for a hybrid checkout system in an exemplary embodiment. [Figure 4A] This diagram illustrates how, in an exemplary embodiment, data about a set of products is acquired by different sets of sensors based on the relative position of the products to the sensors as the set of products is transported along a moving surface. [Figure 4B] This diagram illustrates how, in an exemplary embodiment, data about a set of products is acquired by different sets of sensors based on the relative position of the products to the sensors as the set of products is transported along a moving surface. [Figure 4C]This diagram illustrates how, in an exemplary embodiment, data about a set of products is acquired by different sets of sensors based on the relative position of the products to the sensors as the set of products is transported along a moving surface. [Figure 5] This is a flowchart of the product identification process using the hybrid checkout system shown in Figure 1, in an exemplary embodiment. [Modes for carrying out the invention]

[0008] Consumers use self-checkout for a variety of reasons, including experiencing a more efficient payment process, avoiding contact with cashiers to reduce the risk of disease transmission, and avoiding long lines in manned lanes (especially when the consumer's basket contains only a few items).

[0009] However, at the same time, some consumers remain reluctant to use self-checkout. Some consumers may not be accustomed to the self-checkout process, particularly scanning items and / or manually entering product lookup (PLU) codes. Furthermore, traditional self-checkout, like traditional manned checkouts, involves a sequential item identification process, which requires the operator to individually add each item in the basket of intended purchases to the transaction record. In such a sequential item identification process, each item is identified sequentially, based on, for example, scanning the item's barcode, manually entering a PLU code, or selecting a PLU code based on the operator's choice from a set of candidate options. When consumers have many items to purchase, this sequential item identification process can become cumbersome, requiring consumers to individually enter information for each item (e.g., by scanning a barcode). While traditional manned checkouts are also sequential in nature, consumers who normally use self-checkout and have no other reason to avoid contact with cashiers (e.g., to avoid disease transmission) may prefer to use manned checkouts to avoid having to scan and / or manually enter PLU codes for numerous items themselves.

[0010] Computer vision-based self-checkout technologies are being developed that allow multiple items to be placed simultaneously in a scanning area and identify and add items to the transaction record using computer vision. While these computer vision-based self-checkout technologies reduce the sequential nature of the checkout process by eliminating the need to scan items individually, they can still be burdensome for consumers if the scanning area has limited space and the consumer has a large basket of items, as they must (sometimes repeatedly) remove items from the scanning area and place a new set of items on top of it. In some cases, computer vision-based self-checkout may not even be able to handle batch item recognition in a single transaction. Furthermore, some checkout technologies, whether computer vision-based or not, allow for identification without removing items from the basket or cart, but these technologies are expensive and suffer from various technical problems that affect the accuracy of item recognition, such as difficulties in item identification due to occlusion.

[0011] Embodiments of the technology disclosed herein address the aforementioned technical problems associated with both self-checkout and manned checkouts by providing a hybrid checkout system that combines computer vision with other sensor signals to generate a multi-signal input that is analyzed to identify and recognize goods in motion, thereby transforming the sequential item identification process into a batch process with a higher level of accuracy. The hybrid checkout system provides an attractive checkout experience for both consumers using self-checkout and those using manned checkouts. In some embodiments, the hybrid checkout system may include a cashier who can assist with the checkout process, including bagging items that are identified and added to the transaction record as they come off the conveyor belt, assisting in placing items back on the conveyor belt if items are not recognized based on the multi-signal input, and selectively providing instructions that there are no more items to add to the transaction. In other embodiments, the hybrid checkout system may be operated by the consumer as a self-checkout system without the need for a cashier.

[0012] In an exemplary embodiment, the hybrid checkout system includes a conveyor belt on which goods can be placed, a computer vision device including a plurality of upward cameras positioned at different locations on the computer vision device and having at least partially non-overlapping fields of view (FOV), and one or more other types of sensors, such as one or more radio frequency identification (RFID) sensors. The conveyor belt may be a reinforced checkout conveyor belt that includes markings on it that indicate / indicate the placement of goods on the belt. For example, the reinforced conveyor belt may be marked with a grid pattern that indicates to the operator of the hybrid checkout system (e.g., a consumer, cashier, or employee) that a single item should be placed in each cell of the grid. Other forms of markings may be used in addition to or as an alternative to the grid pattern. For example, a marking (e.g., "X") may be applied to the surface of the belt to indicate a specific location where goods should be placed.

[0013] In some embodiments, the conveyor belt may automatically begin moving when it detects one or more items placed on it. In other embodiments, an operator may need to provide some form of input (e.g., touch input to a display, physical button press, etc.) to initiate the belt's movement. While the items are moving on the belt, multiple cameras (e.g., overhead cameras) may capture images of the items in motion. Due to their diverse arrangement, the cameras may have different fields of view and capture images of the items from multiple different angles. Since images are captured while the items are moving, each camera captures images of the items from different viewpoints / angles as the items pass through its field of view. Each camera may capture images at a configurable frame rate. In this way, each camera captures multiple images of each item from different viewpoints.

[0014] In addition to image data generated by the camera, the hybrid checkout system may also include one or more other types of sensors that acquire additional sensor data, which can be used in combination with the image data to perform batch product identification. For example, one or more RFID sensors / readers may be embedded between the top and bottom surfaces of a conveyor belt. Each RFID reader may include a scanning antenna and a transceiver. The embedded RFID reader may query an RFID tag attached to a product. This may involve activating the tag and sending a radio wave signal to the tag that causes the RFID reader to send back a signal that can be translated to identify the tag's identifier. The product can then be identified based on the stored correspondence between the tag identifier and the product identifier.

[0015] In an exemplary embodiment, a multi-signal input is provided as input to a product identification machine learning model (MLM). In an exemplary embodiment, the multi-signal input includes images of products captured from multiple camera angles by multiple different cameras, RFID data indicating signals detected by an embedded RFID sensor, and optionally other types of sensor data, such as infrared (IR) data acquired by an IR sensor of a hybrid checkout system. The product identification MLM may be pre-trained to a desired level of accuracy based on at least labeled image data. The product identification MLM may receive the multi-signal data as input and output product identification / recognition data including a set of product identifiers for the detected products, which optionally includes confidence values ​​indicating the likelihood that the product identifiers accurately identify the products on the belt. The product identifiers can then be used to retrieve product pricing information and add the products and their corresponding prices to the transaction record of the current transaction.

[0016] Embodiments of hybrid checkout technology disclosed herein provide technical solutions to the aforementioned technical problems associated with self-checkout and manned checkouts. For example, the hybrid checkout system disclosed herein enables batch item identification, which significantly improves transaction speed compared to the typical sequential item identification process associated with conventional self-checkout and conventional manned checkouts. Furthermore, batch item identification is performed while the items are moving, by capturing images of the items from multiple different angles / viewpoints using multiple fixed-position cameras with at least partially non-overlapping fields of view as the items move along the belt surface. By capturing images from multiple viewpoints for each item, the item identification MLM is able to recognize items with a higher level of confidence / accuracy than possible with computer vision self-checkout, for example, where each camera captures an image of the item from a single viewpoint. Furthermore, according to embodiments of the disclosed technology, this image data from multiple viewpoints is complemented by additional sensor data (e.g., RFID data) to further improve the confidence / accuracy of item recognition.

[0017] Thus, the in-moving, multi-signal, bulk item identification process realized by the hybrid checkout system according to the embodiments of the disclosed technology eliminates the burden on operators associated with sequential checkout processes, while significantly improving transaction speed and providing more accurate item recognition capabilities. The higher accuracy is at least in part due to the fact that the image data includes images from multiple viewpoints of each item, and the input data to the item identification / recognition model is multi-signal data from multiple types of sensors.

[0018] Figure 1 is a block diagram of a computer vision-based multi-signal hybrid checkout system architecture in an exemplary embodiment of the disclosed technology. Note that each component is shown schematicly in a greatly simplified form, and only components relevant to understanding the embodiment are shown. Furthermore, the various components and their arrangements illustrated in Figure 1 are presented for illustrative purposes only. Note that other arrangements with more or fewer components are possible without departing from this specification and the teachings presented below.

[0019] A portion of the architecture 100 is provided in a retail environment 110, which may include physical stores such as supermarkets, discount stores, wholesale retailers, department stores / specialty stores, and gas stations. A hybrid checkout device 120 is installed within the store 110. Although Figure 1 depicts a single hybrid checkout device 120, it should be understood that multiple hybrid checkout devices 120 may be installed.

[0020] The hybrid checkout device 120 includes a point-of-sale (POS) system 180. The POS system 180 includes one or more processors 181 and a memory 182. The POS system 180 further includes one or more network interfaces 183 that enable communication with one or more store servers 140, for example, via an internal network 130. The POS system 180 further includes a camera 185, an RFID sensor 186, and one or more other peripheral devices 184. The other peripheral devices 184 may include a payment card reader capable of accepting contactless payments via, for example, near field communication (NFC), and this payment card reader may optionally include a PIN pad. The POS system 180 is configured to receive financial transaction information from a payment means such as a mobile device or a financial card such as a credit card, debit card, or gift card via the payment card reader. Thus, the POS system 180 can obtain consumer financial account-related information via one or more of a number of input means.

[0021] The other peripheral devices 184 may further include a barcode reader / scanner, a weighing instrument, a receipt printer, a display, and the like. The POS system 180 may be configured to receive product identification information from a barcode scanner as a result of an operator scanning a barcode affixed to a product using the scanner. The POS system 180 may be configured to receive weight data from a weighing instrument for products whose price is set based on weight.

[0022] In an exemplary embodiment, the camera 185 may include a plurality of overhead cameras attached, fixed, or otherwise integrated to one or more support structures. The overhead cameras may have at least partially non-overlapping fields of view. The RFID sensor 186 may be configured to query an RFID tag that can be affixed to a product and receive a response therefrom. The hybrid checkout device 120 further includes an enhanced checkout conveyor belt 195. The belt 195 will be described in more detail with reference to FIG. 2.

[0023] While the product is moving on the belt 195, the camera 185 may capture an image of the product disposed on the belt 195. By being disposed at different upper positions, the camera 185 has different fields of view and thus captures images from different angles. Further, since the camera 185 is configured to capture images while the belt 195 is moving, any camera can capture images of the product within its field of view from a plurality of different viewpoints / angles. The captured images may be stored as image data 150A in one or more data stores 150. Although embodiments of the disclosed technology may be mainly described with reference to upper cameras, it should be understood that one or more cameras may be arranged such that their fields of view substantially include side views of the product, upward perspective views of the product, and the like.

[0024] In addition to the captured product images, the RFID sensor 186 may query and obtain RFID data from an RFID tag / label attached to the product. The RFID data may include a stored association between the tag identifier and the product identifier of the product, and may be stored in the data store 150 as RFID data 150B. Similar to the image data 150A, the RFID data 150B may be obtained by the RFID sensor 186 from an RFID tag attached to the product while the product is moving on the belt 195. It should be understood that the POS system 180 may include additional types of sensors, such as IR sensors and weight sensors, to obtain additional sensor data regarding the product while the product is moving.

[0025] The POS system 180 may communicate with the store server 140 and, selectively, with other POS systems within the store 110 via the internal network 130. The POS system 180 may also communicate with other internal and external entities directly, via the internal network 130, or via the external network 160. The POS system 180 may communicate, for example, via one or more micro base stations (BS), pico base stations, or nano base stations. Multiple POS systems 180 may communicate with each other and with external devices using one of a number of different technologies, such as WiFi, Bluetooth, or Zigbee. In some cases, the POS system 180 may match product information and price information with another entity, for example, the internal store server 140 via the internal network 130, and / or one or more remote servers 170 via the external network 160. The POS system 180 may further attempt to obtain financial information related to a transaction and verify such information by transmitting the obtained financial information to one or more servers via at least one of the internal network 130 and one or more external networks 160.

[0026] In various embodiments, networks 130, 160 may include one or more wired and / or wireless networks. The external network 140 may be, for example, the Internet or a private network. The internal network 130 may be, for example, a wired or wireless local area network (LAN). In some embodiments, the internal network 130 may not be provided, and the POS system 180 may communicate directly with one or more external networks 160. In other embodiments, the POS system 180 may be able to communicate with the external network 160, but only indirectly through the store server 140. Please understand that other equipment used for communication through networks 130, 140, such as base stations, routers, access points, and gateways, is not shown for convenience.

[0027] The product identification machine learning model (MLM) 190 is described exemplarily as being hosted and executed on the remote server 170. The product identification MLM 190 may be implemented as executable instructions programmed and resident in memory and / or non-temporary computer-readable (processor-readable) storage media, and may be executed by one or more processors of one or more devices (e.g., the processor of the remote server 170). The same entity or different entities may operate the remote server 170, the store server 140, and / or the POS system 180. In some embodiments, the POS system 180 may transmit at least image data 150A acquired by the camera 185 and RFID data 150B acquired by the RFID sensor 186 as multi-signal input data to the remote server 170 via the internal network 130 and the external network 160. The multi-signal data may be provided as input data to the product identification MLM 190, which has been pre-trained with at least labeled product image data. The product identification MLM 190 may receive multi-signal data as input and output product recognition data 150C, which may be received by the POS system 180 and / or store server 140 and stored in the data store 150. The product recognition data 150C may include the product identifier of the recognized product (e.g., the product's SKU), and may also include a confidence value associated with the SKU output.

[0028] In some embodiments, the image data 150A and / or RFID data 150B do not have to be stored locally in the data store 150, but rather on the remote server 170, or in an unillustrated remote data store accessible by the remote server 170. In other embodiments, the image data 150A and / or RFID data 150B may be stored both locally and remotely. In some embodiments, the product identification MLM 190 (and / or one or more other machine learning models) may additionally or alternatively reside and run locally in the store 110, such as on the store server 140, or on the storage medium of the POS system 180.

[0029] The POS system 180 may compare product recognition data 150C (e.g., a scanned barcode, a product identifier received from a product identification MLM 190) with corresponding price data 150D stored in the data store 150, which may be selectively stored in memory 182 and retrieved from there. Alternatively, the POS system 180 may communicate with another entity (e.g., a store server 140) via, for example, the internal network 130 to obtain price data, and then use this price data to add products and their corresponding prices to the transaction record. The price information may be displayed on the display of the POS system 180.

[0030] The POS system 180 may be configured to access the data store 150 directly via a wired connection, via an internal network 130, and / or indirectly via a store server 140. Furthermore, in some embodiments, the image data 150A may include annotated / labeled image data that associates known product identifiers (e.g., SKUs) with corresponding images of products, which may be provided as training data for the model 190.

[0031] The datastore 150 may include any storage configured to retrieve and store data. Examples of such storage include, but are not limited to, flash drives, hard drives, optical drives, cloud storage, and / or magnetic tapes. The datastore 150 may include, but are not limited to, databases (e.g., relational, object-oriented, etc.), file systems, flat files, distributed datastores where data is stored on multiple nodes of a computer network, peer-to-peer network datastores, etc. The datastore 150 may store one or more database management systems (DBMS). The DBMS may be loaded into memory 182 and may support functions for accessing, retrieving, storing, and / or manipulating data. The DBMS may use any of the various database models (e.g., relational model, object model, etc.) and may support any of the various query languages. The DBMS may access data that is represented in one or more data schemas and stored in any appropriate data repository.

[0032] In some embodiments, the product recognition data 150C may represent a mapping between a product identifier (e.g., SKU) and a corresponding image in the image data 150A. The product recognition data 150C may include data that associates a product identifier, such as the SKU of a product detected by the product identification MLM 190, with the name, visual / graphical representation (e.g., thumbnail image), description, etc., of the corresponding product. In some embodiments, the product recognition data 150C may link the product identifier (e.g., SKU) to a corresponding image in the image data 150A for products where the output of the MLM 190 did not meet an acceptable confidence threshold. In some embodiments, these images may be fed back to the product identification MLM 190 along with an indication of whether the correct product identifier (e.g., SKU) was recognized for the product, thereby allowing the MLM 190 to learn from the received feedback data and improve its product recognition accuracy.

[0033] In some embodiments, the quantity of goods detected using computer vision can be cross-referenced with the quantity of goods detected using RFID to determine if they match. More specifically, a count of the number of goods detected by the MLM 190 can be generated using computer vision and compared with a count of the number of RFID tags detected by the RFID sensor 186. If there is a discrepancy, this may indicate a potential inventory depletion event and may be investigated further by store employees. More specifically, if a discrepancy is detected, an alert may be generated (e.g., a message displayed on the POS system 180, a message sent to the store employee's mobile device) to initiate employee intervention.

[0034] In some embodiments, the MLM 190 may fail to detect / recognize one or more items with an appropriate level of confidence during the initial passage of items on the belt 195, in which case the operator may be notified (e.g., via a message on the display of the POS system 180, an audible instruction, etc.), and the operator (e.g., a consumer or a retailer employee) may place the items back on the belt 195 so that an image of the items is captured again by the camera 185 as the items move across the field of view of the camera 185. In such exemplary scenarios, the RFID sensor 186 may receive a signal from an RFID tag attached to an item that was not recognized by computer vision, in which case the item count by the RFID sensor 186 may be higher than the item count by computer vision-based item detection, even though these scenarios are unlikely to indicate inventory depletion. Therefore, in such exemplary scenarios, the comparison between the item count by RFID detection and the item count by computer vision may not be performed until all items are properly detected / recognized by computer vision, or until any unrecognized items are identified separately by other means, such as barcode scanning.

[0035] Figure 2 is a diagram showing in more detail the reinforced checkout conveyor belt 195 in an exemplary embodiment. The belt 195 includes markings 202 that guide the operator of the hybrid checkout device 120 where to place items on the belt 195. The markings 202 may take the form of a grid pattern, such that a single item is intended to be placed in each cell of the grid. It should be understood that any appropriate markings (e.g., any appropriate graphical representation / marking) can be applied to the belt 195 to assist the operator (e.g., the consumer) in placing items. For example, graphical markings may be applied to the belt 195 in addition to or as an alternative to the grid pattern. In some embodiments, graphical markings may be provided at or near the center of each cell of the grid to further instruct the user to place items on the graphical markings. In other embodiments, a grid pattern may not be provided, but specific markings corresponding to specific locations where items should be placed on the belt 195 may be provided. In yet another embodiment, digital markings may be projected onto the belt surface to indicate where items should be placed.

[0036] An exemplary RFID sensor 204 is also shown in Figure 2, namely one of the RFID sensors 186. The RFID sensor 204 may be embedded in or otherwise integrated into the hybrid checkout device 120, such as in the area between the top and bottom surfaces of the belt 195. The belt 195 is formed of or designed of a material that allows the passage of RFID frequencies, while at the same time being reasonably durable and providing protection against debris or liquids falling onto the belt 195. Although a single exemplary RFID sensor 204 is shown, it should be understood that multiple RFID sensors 204 may be provided. In some embodiments, the RFID sensor 204 may be positioned facing upward toward the top surface of the belt 195 and toward the ceiling of the store 110 to help minimize the reading of RFID tags associated with items not placed on the belt 195. In some embodiments, the RFID sensor 204 has reading power strong enough to read tags several feet away, but weak enough not to read beyond the distance of imaging hardware (e.g., an overhead camera).

[0037] Figure 3 shows a computer vision device 300 of a hybrid checkout device 120 in an exemplary embodiment. The computer vision device 300 has an open frame support structure 302, which in the depicted embodiment has a substantially trapezoidal cross-section that is slightly inclined upward. Various upward cameras 306 may be mounted on the support structure 302 in various positions. For example, the cameras 306 may be mounted on the sides of the support structure 302, or further on support bars extending from opposing sides of the support structure 302. The computer vision device 300 further includes a side arm 304, to which one or more cameras 306 may be mounted, such as at opposing ends of the side arm 304. The various cameras 306 may capture images of each item 308 placed on the belt surface from different angles as the items 308 move along the belt 195. Simultaneously, an RFID sensor located below the upper surface of the belt may also acquire RFID data from RFID tags attached to the items as the items move. In some embodiments, the RFID sensors may be located elsewhere. For example, one or more RFID sensors may be fixed to the computer vision device 300 (e.g., a support structure 302 and / or a side arm 304).

[0038] Figures 4A, 4B, and 4C illustrate, in an exemplary embodiment, how data about a set of goods is acquired by different sets of sensors based on the relative position of the goods to the sensors as the set of goods is transported along a moving surface. Figures 4A, 4B, and 4C show various cameras 402, 404, 406, and 408 positioned at different locations on a computer vision device, which may be the device 300 of Figure 3. In particular, cameras 402 and 404 are depicted exemplary as being mounted on different parts of the support structure of the computer vision device. Cameras 406 and 408 are depicted as being mounted on the side arms of the computer vision device.

[0039] Referring first to Figure 4A, a set of goods is shown as being placed on a conveyor belt. The belt is assumed to be moving, and therefore, in the snapshot shown in Figure 4A, the set of goods is momentarily at a first position 400A. This point in time corresponds to the first point in time in this disclosure. At this point in time, the set of goods may be within the field of view of cameras 402, 404, 406, and 408, respectively. Thus, each of these cameras may take an image of a set of goods, each image taken from a different viewpoint, capturing a different view of the goods. As the goods continue to move with the belt, the goods may remain within the field of view of one or more cameras, and therefore additional images of the goods may be taken from different viewpoints. In some embodiments, a camera may take an image as long as a threshold number of goods are at least partially within the camera's field of view (or as long as a threshold amount of one or more goods is within the camera's field of view). Alternatively, each camera may periodically take an image of its field of view, regardless of whether any part of the goods is within its field of view.

[0040] Referring now to Figure 4B, in this snapshot, the product is at a second position 400B relative to the computer vision device. This point in time corresponds to the second point in time in this disclosure. At this point in time, cameras 402, 406, and 408 may capture images of the product from an angle / viewpoint different from the angle / viewpoint from which the cameras captured images of the product when the product was at the first position 400A relative to the computer vision checkout device. At this point in time, a query signal 412 transmitted by the RFID sensor 410 may be received by one or more RFID tags on one or more products, and these RFID tags may send back RFID data that identifies the tag and the corresponding product to which they are attached. In some embodiments, camera 404 may be directly above the product at this point in time, and such an image may not be suitable for product recognition, so it does not need to capture an image of the product. The viewpoints from which images of the product are captured at the first and second points in time include the "first viewpoint," "second viewpoint," "third viewpoint," and "fourth viewpoint" in this disclosure.

[0041] Next, referring to Figure 4C, in this snapshot, the product is in a third position 400C relative to the computer vision device. At this point, each of the cameras 402, 404, 406, and 408 may again capture images of the product from a different angle / viewpoint than the images of the product captured by the camera when the product was in the first and second positions 400A and 400B relative to the computer vision device.

[0042] Please understand that Figures 4A-4C are illustrative only. In some embodiments, all cameras may capture images of the goods throughout their movement on the belt. In addition, all RFID sensors may continuously transmit query signals while the goods are being transported on the belt. In some embodiments, the cameras and / or RFID sensors may acquire data according to a predetermined activation schedule optimized to generate images that lead to more accurate product recognition results.

[0043] Figure 5 is a flowchart of a product identification / recognition method 500 using the hybrid checkout system of Figure 1 in an exemplary embodiment. The method 500 may be performed at least partially by a processor 181 of a POS system 180 that executes computer executable instructions loaded into memory 182.

[0044] In step S502, as the goods are transported on a moving surface such as a moving conveyor belt, multiple images of the set of goods may be captured by multiple cameras. Each camera may capture images of the goods from different viewpoints / angles as the goods move through the camera's field of view. The processor 181 may execute computer executable instructions to trigger the cameras to capture images. In some embodiments, the cameras may be triggered according to a predefined schedule designed to optimize the quality of the captured images.

[0045] In step S504, additional sensor data about the product is acquired as the product is transported on the moving surface. The additional sensor data may include, for example, RFID data acquired by one or more RFID sensors, which may be located on the hybrid checkout device 120, for example, below the upper surface of the conveyor belt. In some embodiments, the processor 181 may execute a computer executable instruction to trigger the RFID sensors to acquire RFID data. In some embodiments, the additional sensor data may further include IR data, weight data, and so on.

[0046] In step S506, the processor 181 may execute instructions to combine, integrate, or otherwise provide the image data and additional sensor data to the product identification MLM as multi-signal input data. The product identification MLM is pre-trained with at least labeled image data and selectively other sensor data to recognize products and output product identification data.

[0047] In step S508, the processor 181 receives product recognition data as output from the product identification MLM, which includes the product identifier (e.g., SKU) of the recognized product and the confidence value associated with the prediction.

[0048] In step S510, the processor 181 determines the price information of the recognized product, at least partially based on the product recognition data. In particular, the processor 181 may perform a search for the corresponding price of the recognized product based on the product identifier received from the MLM.

[0049] In step S512, the processor 181 adds the goods and their respective prices to the transaction record for the current transaction. In some embodiments, a text or graphical representation of the recognized goods and their corresponding prices may be displayed on the display of the POS system 180.

[0050] The above description is illustrative and not limiting. By examining the above description, many other embodiments will become apparent to those skilled in the art. Therefore, the scope of each embodiment should be determined by reference to the appended claims and the entire scope of equivalents to be given thereto.

[0051] When software is described in a specific form (such as components or modules), this is merely to aid understanding and is not intended to limit how the software implementing those functions may be architectured or structured. For example, modules are illustrated as separate modules, but they may be implemented as uniform code, as individual components, as some (but not all) of these modules combined, or as software structured in any other convenient form. Furthermore, while software modules are illustrated as running on a single piece of hardware, the software may be distributed across multiple processors, or in any other convenient form.

[0052] In the above description of the embodiments, various features are grouped into a single embodiment for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as reflecting that the above embodiments have more features than are explicitly described in each claim. Rather, as reflected in the following claims, the subject matter of the invention lies in fewer features than all the features of the single disclosed embodiment combined. Accordingly, the following claims are incorporated herein into the description of the embodiments, and each claim stands as an independent exemplary embodiment, and any combination of the claimed subject matter also constitutes an embodiment.

Claims

1. When multiple products are transported along a moving surface, image data including multiple images of the multiple products is acquired by one or more cameras. When the aforementioned multiple products are transported along the moving surface, one or more sensors are used to acquire additional sensor data relating to the aforementioned multiple products. The aforementioned image data and the additional sensor data are provided to a product recognition machine learning model (MLM) as multi-signal inputs. The product identifier corresponding to the product recognized by the product identification MLM is received from the product identification MLM, To determine the price corresponding to each of the received product identifiers, Add the received product identifier and the respective product representations corresponding to their respective prices to the transaction record of the current transaction, Methods that include...

2. The one or more cameras mentioned above include a first camera, Acquiring the aforementioned image data involves, at a first time point, capturing a first image of the first product among the plurality of products from a first viewpoint using the first camera, and at a second time point, capturing a second image of the first product from a second viewpoint using the first camera. The method according to claim 1, including the method described in claim 1.

3. The aforementioned one or more cameras further include a second camera, Acquiring the aforementioned image data involves capturing a third image of the first product from a third viewpoint using the second camera at the first time point, and capturing a fourth image of the first product from a fourth viewpoint using the second camera at the second time point. The method according to claim 2, further comprising:

4. The moving surface is a conveyor belt, The first camera and the second camera are fixed cameras positioned at different locations relative to the conveyor belt. The method according to claim 3.

5. The conveyor belt includes one or more markings indicating the locations where the plurality of products should be placed. The method according to claim 4.

6. The first viewpoint, the second viewpoint, the third viewpoint, and the fourth viewpoint are different from each other. The method according to claim 3.

7. The one or more sensors are RFID sensors, the additional sensor data is RFID data received from each RFID tag attached to each of the one or more of the multiple products, and the RFID data includes the identifier of each RFID tag and the identifier of each of the one or more of the multiple products. The method according to claim 1.

8. A first count of the plurality of products is determined based at least partially on the product identifier corresponding to the recognized product, A second count of the plurality of products is determined based at least partially on the RFID data, To determine whether a potential inventory depletion event has occurred, the first count and the second count are compared, The method according to claim 7, further comprising:

9. If, in the comparison described above, it is determined that the first count and the second count do not match, an alert is generated to initiate intervention by an officer. The method according to claim 8, further comprising:

10. A hybrid checkout device, Point-of-sale (POS) systems and Equipped with a mechanism for transporting goods. The aforementioned POS system is Processor and Memory for storing executable instructions, One or more cameras, Equipped with one or more sensors, The processor accesses the memory and executes the executable instruction. When multiple of the aforementioned products are transported by the mechanism, image data including multiple images of the multiple products taken by one or more cameras is received from one or more cameras, The system receives additional sensor data relating to the multiple products, which is acquired when the multiple products are transported by the mechanism, from one or more sensors. The aforementioned image data and the additional sensor data are provided to a product recognition machine learning model (MLM) as multi-signal inputs. The product identifier corresponding to the product recognized by the product identification MLM is received from the product identification MLM, To determine the price corresponding to each of the received product identifiers, Add the received product identifier and the respective product representations corresponding to their respective prices to the transaction record of the current transaction, Configured to perform, Hybrid checkout device.

11. The one or more cameras mentioned above include a first camera, The first camera captures a first image of the first product among the plurality of products from a first viewpoint at a first time point, and captures a second image of the first product from a second viewpoint at a second time point. The hybrid checkout device according to claim 10.

12. The aforementioned one or more cameras further include a second camera, The second camera captures a third image of the first product from a third viewpoint at the first time point, and captures a fourth image of the first product from a fourth viewpoint at the second time point. The hybrid checkout device according to claim 11.

13. The aforementioned mechanism is a conveyor belt, The first camera and the second camera are fixed cameras positioned at different locations relative to the conveyor belt. The hybrid checkout device according to claim 12.

14. The conveyor belt includes one or more markings indicating the locations where the plurality of products should be placed. The hybrid checkout device according to claim 13.

15. The first viewpoint, the second viewpoint, the third viewpoint, and the fourth viewpoint are different from each other. The hybrid checkout device according to claim 12.

16. The one or more sensors are RFID sensors, the additional sensor data is RFID data received from each RFID tag attached to each of the one or more of the multiple products, and the RFID data includes the identifier of each RFID tag and the identifier of each of the one or more of the multiple products. The hybrid checkout device according to claim 10.

17. The processor executes the executable instruction, A first count of the plurality of products is determined based at least partially on the product identifier corresponding to the recognized product, A second count of the plurality of products is determined based at least partially on the RFID data, To determine whether a potential inventory depletion event has occurred, the first count and the second count are compared, Configured to do the following: The hybrid checkout device according to claim 16.

18. The processor executes the executable instruction, If, in the comparison described above, it is determined that the first count and the second count do not match, an alert is generated to initiate intervention by an officer. Configured to do the following: The hybrid checkout device according to claim 17.

19. A non-temporary computer-readable medium that stores executable instructions causing the at least one processor to perform a method, when executed by the at least one processor, The aforementioned method, When multiple products are transported along a moving surface, image data including multiple images of the multiple products is acquired by one or more cameras. When the aforementioned multiple products are transported along the moving surface, one or more sensors are used to acquire additional sensor data relating to the aforementioned multiple products. The aforementioned image data and the additional sensor data are provided to a product recognition machine learning model (MLM) as multi-signal inputs. The product identifier corresponding to the product recognized by the product identification MLM is received from the product identification MLM, To determine the price corresponding to each of the received product identifiers, Add the received product identifier and the respective product representations corresponding to their respective prices to the transaction record of the current transaction, Non-temporary computer-readable media, including [specific examples of such media].

20. The one or more cameras include a first camera and a second camera fixedly positioned at different locations. Acquiring the aforementioned image data At a first point in time, a first image of the first product among the multiple products is taken from a first viewpoint using the first camera, At a second point in time, a second image of the first product is taken from a second viewpoint using the first camera. At the first point in time, a third image of the first product is taken from a third viewpoint by the second camera, At the second point in time, a fourth image of the first product is taken from a fourth viewpoint using the second camera, A non-temporary computer-readable medium according to claim 19, including the following: