Self-checkout device
The self-checkout device uses cameras and machine learning to streamline product registration and payment, addressing inefficiencies in existing systems by enabling quick and easy processing of multiple items, including those without scannable identifiers.
Patent Information
- Application Number
- JP2024577175
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-03
- Filing Date
- 2023-07-24
- Publication Date
- 2025-08-20
AI Technical Summary
Existing self-checkout systems require significant customer interaction, are inefficient in processing multiple items, and struggle with items lacking scannable identifiers, leading to delays in high-throughput sales environments.
A self-checkout device equipped with cameras and machine learning algorithms to detect and identify products based on video footage, allowing for simultaneous registration and payment processing without manual intervention.
Facilitates fast and convenient product registration, reducing delays and interaction requirements, especially for items without scannable identifiers, enhancing the shopping experience in high-throughput sales environments.
Smart Images

Figure 2025527114000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure generally relates to a self-checkout (SCO) device that allows a customer to simply register one or more products by placing the products, each of which is identified, its individual price determined, and the total amount to be paid for the products calculated. [Background technology]
[0002] In a traditional retail environment, customers select various items for purchase and then bring those items to a store associate to check out. Over the past few decades, retail point-of-sale systems have become heavily automated to expedite the checkout process. Computer-based point-of-sale systems are now the standard in retail environments. However, ultimately, such point-of-sale systems may still be operated by store associates. The labor hours it takes to service a checkout counter significantly impact a retail store's overall costs. Reducing or eliminating the time it takes for customers to process and scan their purchases could significantly reduce the labor required for a retail store, thereby alleviating this growing challenge.
[0003] To reduce operational costs, some businesses are implementing self-checkout systems that replace store associates at each checkout terminal. A self-checkout system is a terminal that customers operate themselves without the direct assistance of a store associate. Self-checkout systems typically include a barcode (RFID or other identifier) reader (also called a scanner), a scale for weighing items such as fruits and vegetables, and an interactive screen for selecting products from a predetermined list or entering product codes for products that do not have scannable identifiers (e.g., perishable products such as fruits, vegetables, meats, and bakery items). Self-checkout systems also typically include a payment system, usually accepting cash and card transactions (or other touchless payment mechanisms).
[0004] In a fixed self-checkout system, a customer brings the products they wish to purchase to a fixed location in the store. The customer then presents the products to the self-checkout system, which then registers a record of the products presented. Specifically, the customer individually presents each product to the self-checkout system by scanning each product with a self-checkout scanner (or, if available, the self-checkout system's scanner gun) to detect and interpret any identifiers (i.e., barcodes, RFID tags, etc.) on the products. The self-checkout system then consolidates the registered item details, calculates a total price, and facilitates the customer's payment process.
[0005] However, existing self-checkout systems often still require a high level of interaction from store associates. Furthermore, existing self-checkout systems suffer from various issues, including a poor user interface, an inability to process multiple items at once, and an inability to guide customers to place items at the register. For example, in high-throughput sales environments such as convenience stores, grocery store express lanes, or lunch and takeout sections, customers are often in a hurry and need to quickly register and pay for products. However, these same customers may present multiple products for registration with the self-checkout system. The need to register each of these products separately introduces delays into the sales transaction. These delays are even greater when customers need to register products that do not have scannable identifiers, such as items whose price varies by weight, such as a bunch of bananas or a lunch bowl. To register such presented products, customers must use the self-checkout system's touchscreen component to manually search one or more product lists to find and select matching products. Alternatively, customers may use the touchscreen component to manually enter the product codes of the presented products. In either case, the process of registering such products is very time-consuming and tedious. These delays can be a significant inconvenience and a deterrent to time-pressed customers who want to pay for their purchase quickly and move on.
[0006] It is with this in mind that the present disclosure has been made, and its object is to provide a self-checkout device that creates a fast, easy, and innovative experience for shoppers in express lanes of convenience and grocery stores, such as the lunch or takeout section, by reducing delays and inconvenience in high-throughput sales environments caused by the need to individually register multiple products in a sales transaction. Summary of the Invention
[0007] In an aspect of the present disclosure, a self-checkout device is disclosed. The self-checkout device includes a detection plate adapted to allow placement of a product. The self-checkout device further includes one or more cameras positioned to have a field of view surrounding at least the detection plate, the one or more cameras configured to provide video footage. The self-checkout device further includes a motion detection module configured to detect the presence of motion in the video footage; a sequence selection module configured to select a sequence of video frames over a time interval corresponding to the detection of the presence of motion in the video footage; an appearance interpretation module configured to register one or more products present in the sequence of video frames; a billing module configured to obtain prices of the registered one or more products, generate a total bill based on the obtained prices, and process payment of the total bill; and a controller module operatively connected to the one or more cameras and communicatively coupled to the motion detection module, the sequence selection module, the appearance interpretation module, and the billing module to control their operation and facilitate communication therebetween.
[0008] In one or more embodiments, the appearance interpretation module comprises an object detection module configured to analyze the sequence of video frames to detect one or more objects therein, a cropping module configured to isolate the detected one or more objects within the sequence of video frames and extract visual features of the detected one or more objects, an embedding module configured to convert the extracted visual features of the detected one or more objects into an embedded feature vector, and an expert system module configured to compare the embedded feature vector with pre-stored feature vectors in an embedded database and identify the detected one or more objects based on the comparison, wherein the identified one or more objects are registered as one or more products.
[0009] In one or more embodiments, the expert system module is further configured to determine whether any one of the one or more objects identified from the one or more products is a weight-dependent bulk product item.
[0010] In one or more embodiments, the self-checkout device further comprises a weighing module configured to activate a weighing scale unit to measure a weight of a weight-dependent bulk product item from the one or more products placed on the detection plate, wherein the billing module is configured to generate a total bill based on the measured weight of the weight-dependent bulk product item.
[0011] In one or more embodiments, the self-checkout device further comprises a barcode processing module configured to detect one or more barcodes in the sequence of selected video frames and decode the detected barcodes corresponding to the one or more registered products, wherein the billing module is configured to obtain prices of the one or more registered products based on the decoded barcodes.
[0012] In one or more embodiments, the self-checkout device further comprises a guidance module operatively connected to the design display unit, the guidance module configured to activate the design display unit to display the design on the detection plate and provide visual guidance to a user for optimal placement of the product on the detection plate.
[0013] In one or more embodiments, the sensor further comprises a recessed mounting member disposed upright relative to the detection plate, wherein the one or more cameras are mounted in the recessed mounting member.
[0014] In one or more embodiments, the recessed mounting member houses an illumination device for illuminating the detection plate.
[0015] In one or more embodiments, the one or more cameras comprise a first camera and a second camera oriented at different angles to capture video footage of the product from multiple perspectives.
[0016] In one or more embodiments, the billing module is further configured to generate an itemized list based on the registered product or products.
[0017] In one or more embodiments, the self-checkout device further comprises a display screen configured to display the itemized list and the total charge.
[0018] In one or more embodiments, the self-checkout device further comprises a management module configured to support updates to the configuration of the self-checkout device, including the product database.
[0019] In one or more embodiments, the appearance interpretation module employs machine learning models to facilitate the detection, cropping, embedding, and identification processes.
[0020] In one or more embodiments, the self-checkout device operates as a standalone device.
[0021] In another aspect, a method implemented by a self-checkout device is disclosed. The method comprises receiving video footage of a detection plate of the self-checkout device from one or more cameras. The method further comprises detecting the presence of motion in the video footage by processing the video footage. The method further comprises selecting a sequence of video frames over a time interval corresponding to the detection of the presence of motion in the video footage. The method further comprises detecting and decoding one or more barcodes displayed in the sequence of video frames. The method further comprises calculating a total bill amount corresponding to the decoded one or more barcodes. The method further comprises displaying the total bill amount on a display screen of the self-checkout device.
[0022] In one or more embodiments, the method also includes detecting items displayed in the sequence of video frames when one or more barcodes are not displayed. The method further includes distinguishing between sale items and non-sale items among the detected items. The method further includes issuing a first alert upon detection of the one or more non-sale items, the first alert including a message to remove any non-sale items placed on the detection plate of the self-checkout device.
[0023] In one or more embodiments, the method also includes determining a distribution of the detected vend items on the detection plate of the self-checkout device, and issuing a second alert if the determined distribution of the detected vend items is not proper.
[0024] In one or more embodiments, the method also includes cropping one or more regions from each of the plurality of video frames that substantially surround each detected sales item. The method further includes generating, from each of the one or more cropped regions, an embedded representation of the sales item displayed therein. The method further includes comparing the generated embedded representation to records of embedded representations of products to find a record that matches the embedded representation of the product. The method further includes determining a price corresponding to the record that matches the embedded representation of the product. The method further includes calculating a total bill as the sum of the determined prices corresponding to the records that match the embedded representation of the product for all of the detected sales items. The method further includes displaying the total bill on a display screen.
[0025] In one or more embodiments, the method also comprises receiving payment of the total bill amount.
[0026] In another aspect, a computer program product is disclosed having machine-readable instructions stored thereon that, when executed by one or more processing units, cause the one or more processing units to perform the aforementioned method.
[0027] The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the exemplary aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description. [Brief explanation of the drawings]
[0028] For a more complete understanding of the exemplary embodiments of the present disclosure, reference is now made to the following descriptions taken in conjunction with the accompanying drawings, in which:
[0029] [Figure 1] FIG. 1 illustrates a schematic perspective view of a self-checkout device according to one or more exemplary embodiments of the present disclosure. [Figure 2] FIG. 2 illustrates a schematic side view of a self-checkout device according to one or more exemplary embodiments of the present disclosure. [Figure 3] FIG. 3 illustrates a schematic top view of a self-checkout device according to one or more exemplary embodiments of the present disclosure. [Figure 4] FIG. 4 is an exploded view of a stand unit of a self-checkout device, illustrating its various components, according to one or more exemplary embodiments of the present disclosure. [Figure 5] FIG. 5 is an exploded view of an interaction unit of a self-checkout device, illustrating its various components, in accordance with one or more exemplary embodiments of the present disclosure. [Figure 6] FIG. 6 shows a schematic block diagram of a self-checkout device according to a first exemplary embodiment of the present disclosure. [Figure 7] FIG. 7 shows a schematic block diagram of a system including multiple self-checkout devices according to a second exemplary embodiment of the present disclosure. [Figure 8A] FIG. 8A illustrates a flowchart of a method implemented by a self-checkout device according to one or more exemplary embodiments of the present disclosure. [Figure 8B] FIG. 8B illustrates a flowchart of a method implemented by a self-checkout device, according to one or more exemplary embodiments of the present disclosure. [Figure 9] FIG. 9 shows a schematic perspective view of a self-checkout device according to an alternative embodiment of the present disclosure. [Figure 10] FIG. 10 illustrates an example depiction of a self-checkout device implemented with a product placed thereon, according to one or more example embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0030] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure is not limited to these specific details.
[0031] References herein to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the present disclosure. Appearances of the phrase "in one embodiment" in various places herein do not necessarily all refer to the same embodiment, nor are separate or alternative embodiments mutually exclusive of other embodiments. Furthermore, the terms "a" and "an" herein do not denote a limitation of quantity, but rather indicate the presence of at least one of the referenced item. Furthermore, various features are described that may be exhibited in some embodiments and not in other embodiments. Similarly, various features are described that may be requirements of some embodiments and not in other embodiments.
[0032] Unless otherwise specified in the following description, terms such as "perform," "calculate," "computer-assisted," "compute," "establish," "generate," "configure," "reconstruct," and the like preferably relate to operations and / or processes and / or processing steps that modify and / or generate data and / or transform data into other data, where data may be represented or exist in the form of physical variables, particularly electrical impulses. In particular, the term "computer" should be interpreted as broadly as possible to cover all electronic devices having data processing properties. Thus, a computer may be, for example, a personal computer, a server, a programmable logic controller (PLC), a handheld computer system, a pocket PC device, a mobile wireless device, and other communication devices, processors, and other electronic data processing devices capable of processing data in a computer-assisted manner.
[0033] Furthermore, in particular, a person skilled in the art with knowledge of the method claim / method claims will naturally recognize every routine possibility or possibility of implementation for realizing a product in the prior art, and therefore, there is no need for a separate disclosure in the specification. In particular, these routine implementation variations known to those skilled in the art can be realized solely by hardware (components) or solely by software (components). Alternatively and / or additionally, a person skilled in the art can, within the scope of his or her professional ability, select any combination of hardware (components) and software (components) according to the embodiments of the present invention to the maximum extent according to the embodiments of the present invention to implement the implementation variations according to the embodiments of the present invention.
[0034] The embodiments described herein may be described in the general context of computer-executable instructions residing on some form of computer-readable storage medium, such as program modules, executed by one or more computers or other devices. By way of example, and not limitation, computer-readable storage media may include non-transitory computer-readable storage media and communication media; non-transitory computer-readable media includes all computer-readable media except transitory propagating signals. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The functionality of the program modules may be combined or distributed as desired in various embodiments.
[0035] Some portions of the detailed descriptions which follow are presented and described in terms of processes or methods. Although steps and an order thereof are disclosed in the figures herein that illustrate operation of the methods, such steps and orders are exemplary. Embodiments are suitable for performing various other steps or variations of steps recited in the flowcharts of the figures herein in orders other than those shown and described herein. Some portions of the detailed descriptions which follow are presented in terms of procedures, logic blocks, processing, and other symbolic representations of operations on data bits within a computer memory. These descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. In the present application, a procedure, logic block, process, or the like, is conceived to be a self-consistent sequence of steps or instructions leading to a desired result. The steps are steps employing physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated in a computer system. It is sometimes convenient, principally for reasons of common usage, to refer to these signals as transactions, bits, values, elements, symbols, characters, samples, pixels, or the like.
[0036] In some embodiments, any suitable computer-usable or computer-readable medium (or media) may be utilized. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-usable or computer-readable storage medium (including a storage device associated with a computing device) may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable media may include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a digital versatile disc (DVD), a static random access memory (SRAM), a memory stick, a floppy disk, a mechanically encoded device having instructions recorded on punch cards or grooved structures, media such as media supporting the Internet or an intranet, and a magnetic storage device. It should be noted that a computer usable or computer readable medium may also be a suitable medium on which a program may be stored, scanned, compiled, interpreted, or otherwise processed in a suitable manner, as appropriate, and then stored in a computer memory. In the context of the present disclosure, a computer usable or computer readable storage medium may be any tangible medium that can store or preserve a program for use by or in connection with an instruction execution system, apparatus, or device.
[0037] In some embodiments, a computer-readable signal medium may include a propagated data signal having computer-readable program code embodied therein, for example, as part of baseband or a carrier wave. In some embodiments, such a propagated signal may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. In some embodiments, the computer-readable program code may be transmitted using any suitable medium, including, but not limited to, the Internet, wire, fiber optic cable, RF, etc. In some embodiments, a computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium but can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0038] In some embodiments, computer program code for performing operations of the present disclosure may be either source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Java, Smalltalk, C++, etc. Java and all Java-based trademarks and logos are trademarks or registered trademarks of Oracle and / or its affiliates. However, computer program code for performing operations of the present disclosure may also be written in conventional procedural programming languages such as the "C" programming language, PASCAL, or similar programming languages, or in scripting languages such as JavaScript, PERL, or Python. In this embodiment, the language used for training may be Python, Tensorflow, Bazel, C, or C++. Additionally, a decoder in the user device (described below) may use C, C++, or a processor-specific ISA. Additionally, certain operations may utilize assembly code in C / C++. Also, the ASR (Automatic Speech Recognition) and G2P decoder can run alongside the entire user system without restriction on embedded Linux (any distribution), Android, iOS, Windows, etc. The program code can run completely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server.In the latter scenario, the remote computer may be connected to the user's computer via a local area network (LAN) or wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field programmable gate arrays (FPGAs) or other hardware accelerators, microcontroller units (MCUs), or programmable logic arrays (PLAs) may execute computer-readable program instructions / code to perform aspects of the present disclosure by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry.
[0039] In some embodiments, the flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of apparatuses (systems), methods, and computer program products according to various embodiments of the present disclosure. Each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, may represent a module, segment, or portion of code comprising one or more executable computer program instructions for performing the specified logical functions / acts. These computer program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, and the computer program instructions, executable by the processor of the computer or other programmable data processing apparatus, produce the capability to perform one or more functions / acts specified in the block or combination of blocks in the flowcharts and / or block diagrams. It should be noted that in some embodiments, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may in fact be executed substantially concurrently or may be executed in the reverse order, depending on the functionality involved.
[0040] In some embodiments, these computer program instructions may be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture that includes instruction means that implement the functions / acts specified in the flowchart and / or block diagram blocks or combinations thereof.
[0041] In some embodiments, computer program instructions can be loaded into a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus (not necessarily in a particular order) to generate a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions / acts (not necessarily in a particular order) specified in a block or combination of blocks in the flowcharts and / or block diagrams.
[0042] Furthermore, in the following detailed description of the present disclosure, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be understood that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail as not to unnecessarily obscure aspects of the present disclosure.
[0043] 1-3, different schematic diagrams of a self-checkout device 10 are combined, in accordance with one or more exemplary embodiments of the present disclosure. The self-checkout device 10 is a device used in high-throughput sales environments and is configured to provide an interface that allows customers to pay for services and merchandise without the direct assistance of an employee (unless required). When using the self-checkout device 10, customers are responsible for registering the products they wish to purchase and paying for them. For purposes of this disclosure, the process of registering one or more products with the self-checkout device 10 refers to the process in which each product is presented to the self-checkout device 10 during a transaction, a record of the presented products is created, and a total bill is calculated. For further clarity, this process can be divided into a sequence of connected events, hereinafter referred to as episodes. For example, the process in which a customer registers one of their selected products with the self-checkout device 10 is hereinafter referred to as a "product registration episode." Additionally, the process in which a customer pays for all registered products is hereinafter referred to as a "payment episode."
[0044] Also, for purposes of this disclosure, products that do not have a scannable identifier will hereinafter be referred to as "bulk products." For further clarity, bulk product items whose price is determined by weight will hereinafter be referred to as "weight-dependent bulk product items." For example, weight-dependent bulk product items include a bunch of bananas and a lunch bowl. Those skilled in the art will appreciate that the examples of bulk product items and weight-dependent bulk product items described herein are provided for illustrative purposes only. In particular, those skilled in the art will appreciate that the present disclosure is in no way limited to the above examples. Conversely, the present disclosure is operable to register any product item that does not have a scannable identifier and whose price may or may not depend on the weight of the product item. For consistency, the process of registering a bulk product item will hereinafter be referred to as a "bulk product registration episode."
[0045] As shown, the self-checkout device 10 includes a stand unit 12 and an interaction unit 14. The interaction unit 14 includes a detection plate 32 and a display screen 34 (best shown in FIG. 5 ). As shown in FIGS. 1-3 , the stand unit 12 is coupled to the interaction unit 14 in the self-checkout device 10 of the present disclosure. As particularly shown in FIG. 2 , the stand unit 12 includes a recessed mounting member 16, the lower end 18 of which is attached to a base member 20. As shown, the recessed mounting member 16 is disposed upright relative to the interaction unit 14 (specifically, the detection plate 32 therein). The base member 20 may include a stand-mating structure (not shown) for enabling coupling of the stand unit 12 and the interaction unit 14, which will be described in more detail below. In this embodiment, either or both of the recessed mounting member 16 and the base member 20 are formed from a metal or plastic material. Both the recessed mounting member 16 and the base member 20 have non-reflective surfaces. Preferably, the surfaces of the recessed mounting member 16 and the base member 20 are light-absorbing. More preferably, the surfaces of the recessed mounting member 16 and the base member 20 are black.
[0046] Referring to FIG. 4 , an exploded view illustrating various components of the stand unit 12 is shown, in accordance with one or more exemplary embodiments of the present disclosure. As shown, the inner surface 22 of the concave mounting member 16 provides a recess configured to accommodate a lighting device (not shown). In one example, the lighting device may include multiple light-emitting diodes (LEDs). Preferably, the lighting device includes multiple RGB (red, green, blue) LEDs to enhance illumination of nearby components (the interaction unit 14, more specifically the detection plate 32 (shown in FIG. 3 ) and products placed thereon (described below)). In this configuration, the lighting device is further configured to reduce the effects of glare and ambient light in the environment surrounding the self-checkout device. The stand unit 12 further includes a light-diffusing case member 24. As shown, the light-diffusing case member 24 has a concave shape. Here, the curvature of the light-diffusing case member 24 substantially matches the curvature of the concave mounting member 16. In one example, the light-diffusing housing member 24 is formed from polycarbonate, acrylic or polymethyl methacrylate (PMMA), polystyrene, or any other suitable plastic material. The light-diffusing housing member 24 is preferably white and may have a milky or translucent surface finish. Those skilled in the art will recognize that the above-described materials and surface finishes for the light-diffusing housing member 24 are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the self-checkout device 10 of the present disclosure is in no way limited to the use of these materials or surface finishes for the light-diffusing housing member 24. Conversely, the present disclosure is operable with any material or surface finish capable of diffusing light from the lighting device. In use, the light-diffusing housing member 24 is attached to the inner surface 22 of the recessed mounting member 16, with the light-diffusing housing member 24 effectively forming a cap on top of the inner surface 22 of the recessed mounting member 16. In this manner, the lighting device is sandwiched between the inner surface 22 of the recessed mounting member 16 and the light-diffusing housing member 24. Thus, the lighting device effectively forms a backlight member for the light-diffusing housing member 24.
[0047] Additionally, as shown, the stand unit 12 includes one or more cameras. Specifically, the one or more cameras include a first camera 26 and a second camera 28 oriented at different angles to capture video footage of the product from multiple perspectives. Generally, the cameras 26, 28 are positioned to have a field of view that encompasses at least the detection plate 32 of the interaction unit 14. In this configuration, the first camera 26 is mounted on the upper end 27 of the inner surface 22 of the recessed mounting member 16. Specifically, the first camera 26 is mounted facing downward on the recessed mounting member 16, and its field of view encompasses a top-down view of the area below it and one or more objects contained within that area. The first camera 26 is preferably an RGB-D camera. For example, the first camera 26 may include a time-of-flight (TOF) sensor, a structured light sensor, or a stereoscopic sensor. Those skilled in the art will recognize that the above examples are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the self-checkout device of the present disclosure is not limited to the above examples. Conversely, the present disclosure is operable with any one or more sensors whose output signals provide a three-dimensional representation of the observed scene. Preferably, the first camera 26 has a 4K or equivalent resolution. Preferably, the first camera 26 has an effective autofocus function. The second camera 28 is mounted on the inner surface 22 of the concave mounting member 16 at a height approximately midway between the upper end 27 and the lower end 18 of the concave mounting member 16. In such a case, the light-diffusing housing member 24 is provided with an opening 30 positioned at a position that substantially coincides with the position of the second camera 28 when the light-diffusing housing member 24 is mounted on the concave mounting member 16. The opening 30 ensures that the field of view of the second camera 28 is not obstructed by the light-diffusing housing member 24. Preferably, the second camera 28 includes an RGB camera. More preferably, the second camera 28 has a 4K or equivalent resolution. More preferably, the second camera 28 has an effective autofocus function.
[0048] 1-4 , the base member 20 of the stand unit 12 is mechanically coupled to the interaction unit 14. Specifically, the base member 20 is attachable to either side of the interaction unit 14 by an interfitting structure (not shown) including a stand unit mating structure (not shown) and corresponding interaction unit mating structures (not shown) formed on one or more sides of the base member 20 and the interaction unit 14, respectively. The interfitting structure may include one or more tongue-and-groove arrangements between the corresponding sides of the base member 20 and the interaction unit 14, or a slot or other recess on either one or more sides of the interaction unit 14, where the slot or recess is configured to receive at least a portion of the corresponding side of the base member 20. Those skilled in the art will understand that the above-described coupling means are provided for illustrative purposes only. In particular, those skilled in the art will understand that the present disclosure is not limited to the above-described coupling mechanisms. Conversely, the present disclosure is operable with any mechanism for coupling the base member 20 to the interaction unit 14 that is robust enough to hold the base member 20 in a fixed position relative to the interaction unit 14 for extended periods of time, even under impacts and bumps from users of the self-checkout device 10 and objects placed on the interaction unit 14. It will also be appreciated that coupling of the base member 20 and the interaction unit 14 need not be achieved mechanically. Instead, the base member 20 may be magnetically coupled to the interaction unit 14, for example. The coupling mechanism is further configured such that when coupled with the interaction unit 14, the base member 20 is oriented such that the arc of the concave mounting member 16 curves inward toward the interaction unit 14.Here, the interaction unit mating structure includes interaction unit contacts (not shown) that are configured to contact stand unit contacts (not shown) attached to the outer surface of the base member 20 in the stand unit mating structure upon mating of the interaction unit 14 with the base member 20, supporting electrical and communicative coupling between the cameras 26, 28 of the stand unit 12 and the control circuitry (described below) of the base member 20. Furthermore, the cameras 26, 28 and the lighting device are communicatively and electrically coupled to the stand unit contacts, respectively.
[0049] In another embodiment, the recessed mounting member 16 may include two substantially coincident, spaced apart arcuate members (not shown) mounted parallel and in an upright position from the interaction unit 14. The arcuate members are periodically joined by a plurality of crossbars (not shown) to provide support, structural reinforcement, and stability to the two arcuate members. In this embodiment, the cameras 26, 28 are supported at the top and mid-height of the crossbars, respectively.
[0050] 9, a depiction of a self-checkout device 10 according to an alternative embodiment of the present disclosure is shown. As shown in FIG. 9, the self-checkout device 10 may have a stand unit 12 attached to a side (here, the left side) of the interaction unit 14, as opposed to a rear side (as shown and described with reference to FIGS. 1-3) of the self-checkout device 10. In yet another embodiment, two stand units 12 (not shown) may be provided that may be mechanically coupled to the interaction unit 14. Specifically, the base member 20 of each of the two stand units 12 is attachable to any two of one or more side surfaces of the interaction unit 14 by one or more interfitting structures including a stand unit mating structure on each stand unit and corresponding interaction unit mating structures formed on any two of the one or more side surfaces of the base member 20 and the interaction unit 14, respectively. As previously mentioned, the interaction unit mating structure includes interaction unit contacts configured to contact each of the stand unit contacts upon mating of the interaction unit 14 with the two base members 20, supporting electrical and communication coupling between the cameras of the two stand units 12 and the control circuits (described below) of the two base members 20. This embodiment is particularly useful in difficult ambient conditions (e.g., lighting, glare, external elements) because it doubles the number of cameras in the self-checkout device 10 and increases the amount and type of illumination. This improves the reliability of detecting products placed on the self-checkout device 10 by at least 10%.
[0051] FIG. 5 illustrates an exploded view showing various components of the interaction unit 14, in accordance with one or more exemplary embodiments of the present disclosure. Referring to FIG. 5 in combination with FIGS. 1-4 , the interaction unit 14 includes a detection plate 32 and a display screen 34. Here, the detection plate 32 is adapted to allow products to be placed thereon. In one example, the top surface of the detection plate 32 is configured as a square measuring 39 cm by 39 cm. However, those skilled in the art will appreciate that the above configuration of the top surface of the detection plate 32 is provided for illustrative purposes only. In particular, those skilled in the art will appreciate that the self-checkout device 10 of the present disclosure is not limited to the above configuration of the top surface of the detection plate 32. Conversely, the present disclosure is operable with a detection plate 32 of a size or shape that allows multiple products (not shown) to be placed thereon without stacking or crowding, and that allows the cameras 26, 28 to identify product features without being obstructed by other products. For example, the detection plate 32 may have a larger size, such as 50 cm x 50 cm, to accommodate more products without departing from the spirit and scope of the present disclosure.
[0052] In one embodiment, the detection plate 32 may be backlit. The backlighting of the detection plate 32 is configured to eliminate reflections and shadows that may cause or contribute to erroneous results from the self-checkout device 10. In one example, the backlighting of the detection plate 32 is implemented by multiple backlight elements (not shown). The backlighting elements are spatially distributed across the horizontal plane of the detection plate 32, thereby illuminating substantially the entire top surface of the detection plate 32. Furthermore, in some examples, individual backlighting elements may be individually activated to provide configurable and variable illumination patterns in different regions of the detection plate 32. In particular, areas of the top surface of the detection plate 32 where no backlighting elements are activated may appear darker in color than the remainder of the top surface of the detection plate 32. Thus, by controllably activating the individual backlighting elements, various designs may be effectively displayed on the top surface of the detection plate 32. These designs may be configured to provide guidance to a user regarding where to place products on the detection plate 32 to facilitate detection by the self-checkout device 10. In another example, backlighting of the detection plate 32 is not achieved by multiple individually actuatable backlight elements, but instead is implemented such that the operation of one or more backlight elements is synchronized, such that one or all backlight elements are activated or deactivated simultaneously. In yet another example, a design may be displayed on the detection plate 32 by a projection device (not shown) attached to the stand unit 12. For simplicity, the individually actuatable backlight elements and the projection unit will hereinafter be referred to as a "design display unit."
[0053] Also as shown, display screen 34 includes a top surface 36 operable to display information, including a list of products being processed in a transaction and messages to the user or store associate. Display screen 34 may be a touchscreen configured to detect a user's touch on top surface 36 of display screen 34 and the location of that touch.
[0054] Further, as shown, the interaction unit 14 includes a 1D barcode reader 38, a contactless card reader 40, and a multi-function button 42. The 1D barcode reader 38, the contactless card reader 40, and the multi-function button 42 are operable to provide additional mechanisms for a user to interact with the self-checkout device 10 of the present disclosure. Accordingly, for brevity, the 1D barcode reader 38, the contactless card reader 40, and the multi-function button 42 are hereinafter referred to as “additional user interface elements.” The 1D barcode reader 38 may be configured to detect, read, and decode a barcode presented thereon and, in response, output a detected barcode signal containing information about the decoded barcode. The contactless card reader 40 may be configured to transmit payment instructions to a presented payment card and to receive and output payment details from the payment card. The multi-function button 42 may be configured to output a signal corresponding to a detected press. For brevity, the signals output from each of the additional user interface elements are hereinafter referred to as “additional user interface signals.”
[0055] In one example, the multi-function button 42 is omitted from the "additional user interface elements," and user interaction with the self-checkout device 10 is replaced by detection of the user's finger movements relative to the display screen 34 by the cameras 26, 28, or an IR device (not shown) of the stand unit 12, thereby converting the display screen 34 into a basic touchscreen device. Specifically, no button presses are required to advance the operation of the self-checkout device 10 (e.g., to unload and load product batches from the detection plate 32, to proceed through the payment process, or to access administrative functions). Instead, well-defined gestures serve to move forward or backward through the process or change transaction stages or administrative functions. In such cases, the display screen 34 may be configured to display permissible gestures to the user. The gestures should be easily distinguishable, easily detectable, and undone by the user in the event of an error. Eliminating mechanical components, such as the multi-function button 42, increases the durability of the self-checkout device 10 and reduces its maintenance requirements. Those skilled in the art will recognize that the above-described members of the additional user interface elements are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the self-checkout device 10 of the present disclosure is not limited to the above-described members of additional user interface elements. Rather, the present disclosure is operable with any other mechanism that allows a user to interact with the self-checkout device 10. For example, a preferred embodiment may include a speaker and microphone system configured to play predefined messages to the user and to detect and receive speech from the user.
[0056] The interaction unit 14 further includes an open-ended casing member 44. The casing member 44 is configured to receive and house the display screen 34 and the detection plate 32 substantially side-by-side, with the top surfaces 36 of the display screen 34 and the detection plate 32 exposed through the open end of the casing member 44. The casing member 44 is further configured to arrange and house additional user interface elements such that a user can access each of the additional user interface elements through the open end of the casing member 44. For example, the additional user interface elements are arranged substantially side-by-side with the detection plate 32 and the display screen 34, such that, proceeding from left to right of the casing member 44 (as best shown in FIG. 3 ), the display screen 34 is effectively sandwiched between the detection plate 32 and the additional user interface elements. Alternatively, the additional user interface elements may be arranged around the periphery of the display screen 34 and / or the detection plate 32. Furthermore, those skilled in the art will recognize that if the relative spatial order of the display screen 34 and the detection plate 32 were reversed, proceeding from the left side to the right side of the casing member 44, the detection plate 32 would effectively be sandwiched between the display screen 34 and the additional user interface element.
[0057] In alternative embodiments, to reduce the footprint of the interaction unit 14 and increase its usability in tight spaces, additional user interface elements may be formed in the recessed mounting member 16 of the stand unit 12 rather than in the interaction unit 14. Furthermore, the portions of the interaction unit 14 that are not part of the detection plate 32 may be reduced by moving the contactless card reader 40 closer to the display screen 34, omitting the 1D barcode reader 38 from the additional user interface elements, and relying on the second camera 28 of the stand unit 12 to read the barcodes of presented products. Doing so may reduce the footprint of the self-checkout device 10 and further increase its usability in space-constrained environments. Similarly, providing a simplified user interface with fewer interaction options may improve the user experience operating the self-checkout device 10.
[0058] The interaction unit 14 further includes a transparent protective plate member 46. The protective plate member 46 is configured with dimensions that substantially match the open edge of the casing member 44. The protective plate member 46 is further configured to mount over the open edge of the casing member 44 and form a substantially waterproof seal, so that the display screen 34, the detection plate 32, and additional user interface elements are effectively sandwiched between the protective plate member 46 and the casing member 44. In one embodiment, the protective plate member 46 is formed from scratch-resistant or tempered glass. In another embodiment, the protective plate member 46 is formed from a transparent, impact-resistant plastic material, such as clear polycarbonate, clear acrylic, clear polyethylene terephthalate glycol (PETG), or clear polyvinyl chloride (PVC). In one example, the protective plate member 46 is painted with an opaque pigment or covered with an opaque adhesive foil. Preferably, either or both of the opaque pigment and the opaque adhesive foil are black. Here, the opaque pigment or the opaque adhesive foil is not present in the display area 48 of the protective plate member 46. As shown, the display area 48 is aligned with the display screen 34 and the detector plate 32 when the protective plate member 46 is installed over the open end of the casing member 44. Additionally, the protective plate member 46 includes a plurality of cutout areas 50 that are aligned with additional user interface elements when the protective plate member 46 is installed over the open end of the casing member 44. This allows a user to clearly and unobstructedly view the display screen 34 and the detector plate 32, while being protected from impacts, scratches, and accidental spills by the protective plate member 46. Additionally, the user has unobstructed access to each of the additional user interface elements.
[0059] In this embodiment, the interaction unit 14 is configured to have a height-to-width ratio of 16:9. Preferably, the interaction unit 14 is configured to have a diagonal of 32 inches. However, those skilled in the art will recognize that the above-mentioned dimensions and relationships therebetween are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the self-checkout device 10 of the present disclosure is in no way limited to these dimensions and relationships therebetween. To the contrary, the present disclosure is operable with any physical dimensions and / or relationships therebetween sufficient to simultaneously accommodate multiple products of different sizes and shapes without being crowded or stacked, and can be conveniently incorporated into and adapted to the existing environment of a high-throughput sales environment with minimal disruption to customers, staff, and existing processes and systems operating in that environment.
[0060] In particular, during use, the interfitting structure of the base member 20 and the interaction unit 14 is configured to position the recessed mounting member 16 in an overhanging, aligned position with the detection plate 32. Specifically, the interfitting structure is configured to position the first camera 26 of the recessed mounting member 16 substantially directly above the detection plate 32, such that the first camera 26 can provide a top-down view of the detection plate 32 and any products placed thereon. In one example, the recessed mounting member 16 is configured to position the first camera 26 60 cm above the detection plate 32. However, those skilled in the art will recognize that the above height is provided for illustrative purposes only. In particular, those skilled in the art will recognize that the self-checkout device 10 of the present disclosure is not limited to the above height. Conversely, the recessed mounting member 16 can be configured at any height suitable to provide sufficient height for the first camera 26 to provide a clear, unobstructed top-down view of even the tallest products that may be placed on the detection plate 32. Additionally, the curvature of the recessed mounting member 16 is configured to maximize its stability and minimize obscuration of its front area by the recessed mounting member 16, thereby allowing a user to access the detection plate 32 and for the user to conveniently place a product on the detection plate 32. The curvature of the recessed mounting member 16 and the placement of the second camera 28 are further configured to provide the second camera 28 with a wide field of view that forms a lateral view of the detection plate 32 and a product placed thereon.
[0061] The interaction unit 14 further includes a control circuit (not shown), wherein substantially each of the detection plate 32, the display screen 34, and the additional user interface elements is communicatively and electrically coupled to the control circuit. The control circuit is also communicatively and electrically coupled to the interaction unit contacts. The casing member 44 is configured to house the control circuit (not shown) alongside the display screen 34, the detection plate 32, and at least one of the additional user interface elements. Alternatively, the casing member 44 may be configured with a plurality of vertically separated internal slots adapted to house the control circuit (not shown) in a sandwiched arrangement between the bottom of the casing member 44 and the underside of either or both of the display screen 34 and the detection plate 32. In some examples, the casing member 44 is further provided with a plurality of vents to allow heat to escape from the display screen 34 to prevent overheating of the display screen and / or the control circuit. In this embodiment, the casing member 44 may be formed from a durable, lightweight, waterproof plastic or rubber material suitable for withstanding everyday wear and preventing spilled liquids from reaching the display screen 34. For hygiene reasons, the casing member 44 should also be easily cleanable.
[0062] In one or more embodiments, the interaction unit 14 of the self-checkout device 10 of the present invention further includes a weighing scale unit (not shown) communicatively coupled to the control circuitry. The weighing scale unit (not shown) may be configured to be housed within the casing member 44 in a sandwiched arrangement between the bottom of the casing member 44 and the underside (not shown) of the detection plate 32 to enable weighing of products placed on the detection plate 32.
[0063] The control circuitry includes an LED receiver (not shown) communicatively coupled to the lighting device via the interaction unit contacts and the stand unit contacts. The LED receiver can be configured to receive control signals from a controller (not shown), which is located remotely from the LED receiver and wirelessly connected to the LEDs of the lighting device via the LED receiver to control the color and intensity of the LEDs and turn one or more LEDs on or off as needed. The control circuitry can further include a microprocessor (not shown) configured to receive video footage from the cameras 26, 28 and additional user interface signals from additional user interface elements. The microprocessor can be further configured to receive signals from a weigh scale unit, if present, indicating the weight of a product placed on the detection plate 32. The microprocessor can be further configured to process the received video footage and additional user interface signals, as well as signals from the weigh scale unit, and, based on such processing, issue control signals to at least one of the detection plate 32, the display screen 34, and the contactless card reader 40. Specific functions of the microprocessor supporting the operation of the self-checkout device 10 are described in more detail in the previous paragraph.
[0064] Referring now to FIG. 6, a schematic block diagram of a self-checkout device 10 according to a first exemplary embodiment of the present disclosure is shown. Specifically, FIG. 6 illustrates the self-checkout device 10 as a standalone device. As shown in FIG. 6, in combination with FIGS. 1-5 described in the previous paragraph, the self-checkout device 10 includes a controller module 102 operably connected to first and second cameras 26 and 28 mounted on the recessed mounting member 16 of the self-checkout device 10 and communicatively coupled to various modules / components (described below) therein. The self-checkout device 10 further includes a motion detection module 104 configured to receive video footage via the controller module 102 and detect the presence of motion in the video footage. The self-checkout device 10 further includes a sequence selection module 106 configured to receive video footage via the controller module 102 and select a sequence of video frames over a time interval corresponding to the detection of the presence of motion in the video footage. In other words, the sequence selection module 106 selects a sequence of video frames over a time interval beginning after the detection of motion and after the detection of the end of the motion by the motion detection module 104 (i.e., from the start of the detection of the presence of motion in the video footage to the end of the detection of the presence of motion in the video footage). The self-checkout device 10 further includes a barcode processing module 108 configured to detect the presence of a barcode in the sequence of video frames and decode the barcode. The self-checkout device 10 further includes an appearance interpretation module 114 configured to detect, recognize, and identify objects displayed in the sequence of video frames based on the appearance of the objects, thereby registering one or more products present in the sequence of video frames. The self-checkout device 10 further includes a weighing module 126 configured to operate a scale unit thereof to measure the weight of a weight-dependent bulk product item.The self-checkout device 10 further includes a billing module 128 configured to obtain prices of one or more registered products, generate an itemized list based on the registered one or more products, generate a total bill based on the obtained prices, and process payment of the total bill. The self-checkout device 10 further includes a guidance module 130 operatively connected to the design display unit and configured to display a design on the detection plate 32 to provide a user with visual guidance for optimally placing products on the detection plate 32. The controller module 102 is communicatively coupled to the motion detection module 104, the sequence selection module 106, the barcode processing module 108, the appearance interpretation module 114, the weighing module 126, the billing module 128, and the guidance module 130 to control their operation and facilitate communication therebetween. The self-checkout device 10 further includes a management module 132 communicatively coupled to the controller module 102 to support updating the configuration of the software of the self-checkout device 10, including the product database 110, and resetting the software of the self-checkout device 10 as needed. Each of these modules and the relationships between them are described in more detail in the previous paragraphs.
[0065] Motion Detection Module 104 In a high-throughput sales environment, significant hand and product movement may occur in the area adjacent to the self-checkout device 10 and the detection plate 32 as products are placed on or removed from the detection plate 32. These movements may make detection of the product more difficult. Therefore, to improve operational performance, monitoring of the detection plate 32 should occur only when a customer has placed a product on the detection plate 32.
[0066] In this configuration, the motion detection module 104 is coupled to a first camera 26 positioned above the detection plate 32 to obtain a bird's-eye view thereof. The motion detection module 104 is further coupled to a second camera 28 positioned to the side of the detection plate 32 at the same height as the detection plate 32 to obtain a side view thereof. The motion detection module 104 is adapted to receive two video image streams from the first camera 26 and the second camera 28, respectively. As can be understood, the video image from the video camera includes a plurality of video frames captured consecutively, where p is the number of video frames in the captured video image. For a given video frame
[0067]
number
[0068] is recorded by the video camera.
[0069]
number
[0070]
number
[0071] is the time interval between the capture of the first video frame and the capture of the next video frame (also called the sampling interval). Using this notation, the video footage captured by the first camera 26 is
[0072]
number
[0073] where:
[0074]
number
[0075] is the sampling time
[0076]
number
[0077] For simplicity, the stream of video footage received from the first camera 26 will be referred to hereafter as "first video stream VID1." Similarly, the video footage captured by the second camera 28 will be referred to hereafter as "first video stream VID1."
[0078]
number
[0079] where:
[0080]
number
[0081] is the sampling time
[0082]
number
[0083] 1 and 2. The video frames VID1 and VID2 are captured from the second camera 28 at the same time. For simplicity, the stream of video footage received from the second camera 28 is hereinafter referred to as the "second video stream VID2." In one embodiment, the video frames in the first video stream VID1 and the second video stream VID2 are encoded using the H.264 video compression standard. The H.264 video format uses motion vectors as a key element of compressed video footage. The motion detection module 104 detects motion in the first video stream VID1 and the second video stream VID2 using the motion vectors obtained from decoding the H.264-encoded video frames. In another embodiment, the first video stream VID1 and the second video stream VID2 are each sampled at predefined intervals. The sampling intervals for the first video stream VID1 and the second video stream VID2 are configured to be long enough to avoid falsely detecting small, fast movements, such as finger movements, rather than large movements corresponding to the placement or removal of a product from the detection plate 32.
[0084] Consecutive samples of a video frame from the first video stream VID1
[0085]
number
[0086] Similarly, consecutive samples of a video frame from the second video stream VID2 are compared to detect the difference between them.
[0087]
number
[0088] and compare the samples to detect differences between them. A difference exceeding a predefined threshold is taken to indicate that motion occurred during the intermediate period between successive samples. The threshold is configured to avoid temporary changes, such as light flicker, being mistaken for motion. When the start and completion of motion is detected by the motion detection module 104, a "motion trigger" signal is sent by the motion detection module 104 to the controller module 102.
[0089] Sequence Selection Module 106 The sequence selection module 106 is communicatively coupled to the first and second cameras 26 and 28 to receive a first video stream VID1 and a second video stream VID2 therefrom. The sequence selection module 106 is also communicatively coupled to the controller module 102 to receive a motion trigger signal therefrom. Upon receiving the motion trigger signal, the sequence selection module 106 is configured to extract first and second sequences of video frames from the first and second video streams VID1 and VID2, respectively. For brevity, the sequence of video frames selected from the first video stream VID1 will hereinafter be referred to as the first selected sequence VS1. Similarly, the sequence of video frames selected from the second video stream VID2 will hereinafter be referred to as the second selected sequence VS2. The first selected sequence VS1 and the second selected sequence VS2 include video frames that begin at the time of issuance of the motion trigger signal and continue for a predetermined interval thereafter. For ease of understanding and consistency with the notation above, the time of issuance of the motion trigger signal is represented as
[0090]
number
[0091] where:
[0092]
number
[0093] is the corresponding sampling interval at which the motion trigger signal is emitted. Similarly, the selected video sequence is
[0094]
number
[0095] where α is a value sufficient to identify the product placed on the detection plate.
[0096] Using this nomenclature, the first selected sequence VS1 is:
[0097]
number
[0098] Similarly, the second selected sequence VS2 can be written as:
[0099]
number
[0100] In other words, the first selected sequence VS1 includes α consecutively sampled video frames captured from the first camera 26, and the second selected sequence VS2 includes α consecutively sampled video frames captured from the second camera 28. Furthermore, the start times of the first selected sequence VS1 and the second selected sequence VS2 are
[0101]
number
[0102] and the end time
[0103]
number
[0104] The sequence selection module 106 is communicatively coupled to the controller module 102 and transmits the first selected sequence VS1 and the second selected sequence VS2 to the controller module 102.
[0105] Barcode Processing Module 108 The barcode processing module 108 is coupled to the controller module 102 and receives therefrom the first selected sequence VS1 and the second selected sequence VS2. The barcode processing module 108 includes a barcode analysis algorithm configured to receive video frames from either or both of the first selected sequence VS1 and the second selected sequence VS2.
[0106] The barcode analysis algorithm is configured to detect the presence of barcodes in the received video frames and decode any such detected barcodes into a corresponding text representation, which for brevity is hereinafter referred to as a "barcode ciphertext." In one embodiment, the barcode analysis algorithm includes the known "Quick Browser" model (as described in T. Do and D. Kim, "Quick Browser: A Unified Model to Detect and Read Simple Objects in Real-time," 2021 International Joint Conference on Neural Networks (IJCNN), 2021, pp. 1-8.). In another embodiment, the barcode analysis algorithms include (a) barcode detection, localization, and rotation algorithms (such as those described in "Real-time barcode detection and Classification Using Deep Learning," DK Hansen, Nasrollahi K., Rasmussen CB, Moeslund TB, Proc. 9th International Joint Conference on Computational Intelligence, 2017), and (b) barcode decoding algorithms (such as those described in the ZXing (Zebra Crossing) barcode scanning library) or deformable barcode digit models (such as those described in O. Gallo and R Manduchi "Reading Challenging Barcodes with Cameras," Proceedings of the IEEE Workshop on Applications of Computer Vision, 2009 (7-8): 1-6). Those skilled in the art will understand that these algorithms are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the present disclosure is not limited to the use of these barcode detection and decoding algorithms.Conversely, the present disclosure is operable with any algorithm capable of detecting, locating, and decoding barcodes visible in video frames captured by first camera 26 or second camera 28. For example, the barcode analysis algorithm may use a single-shot detector (SSD) algorithm (such as described in Y. Ren and Z. Liu, “Barcode detection and decoding method based on deep learning,” 2nd International Conference on Information Systems and Computer Aided Education (ICISCAE), 2019, 393-396) to detect the presence of a barcode.
[0107] The barcode processing module 108 is communicatively coupled to a product database 110 that stores a plurality of tuples containing barcode cryptogram elements for each product in the store's inventory, a corresponding identifier for each product, and a price for each product. At least 2,000 of the tuples also contain embedding vectors, which will be described later in connection with the embedding module for each product / bulk product. Those skilled in the art will appreciate that this number of tuples containing embedding vectors is provided for illustrative purposes only. In particular, those skilled in the art will appreciate that the self-checkout device 10 of the present disclosure is not limited to the number of tuples in the product database 110 containing embedding vectors. Conversely, the self-checkout device 10 of the present disclosure can operate with any number of tuples in the product database 110 that contain sufficient embedding vectors to identify at least a portion of the products / bulk products in the store's inventory based on their appearance.
[0108] If the barcode processing module 108 detects the presence of a barcode in a video frame of the first selected sequence VS1 or the second selected sequence VS2 and decodes the detected barcode, the barcode processing module 108 is adapted to use the resulting barcode cryptogram to query the product database 110 and obtain therefrom a product identifier and a product price corresponding to the barcode cryptogram. The barcode processing module 108 is further configured to communicate the identifier and corresponding price to the controller module 102. In contrast, if the barcode processing module 108 fails to detect the presence of a barcode in a video frame of the first selected sequence VS1 or the second selected sequence VS2, the barcode processing module 108 is communicatively coupled with the controller module 102 and emits an “appearance activation” signal.
[0109] Appearance Interpretation Module 114 The appearance interpretation module 114 includes an object detection module 116 configured to analyze the sequence of video frames from the sequence selection module 106 to detect one or more objects therein. The appearance interpretation module 114 further includes a cropping module 118 configured to isolate the detected one or more objects within the sequence of video frames and extract visual features of the detected one or more objects. The appearance interpretation module 114 further includes an embedding module 120 configured to convert the extracted visual features of the detected one or more objects into an embedded feature vector. The appearance interpretation module 114 further includes an expert system module 122 configured to compare the embedded feature vector with feature vectors pre-stored in an embedding database 124 and identify the detected one or more objects based on the comparison. The identified one or more objects are then registered as one or more products. In an embodiment of the present disclosure, the appearance interpretation module 114 uses machine learning models to facilitate the detection, cropping, embedding, and identification processes.
[0110] Specifically, the appearance interpretation module 114 is communicatively coupled to the controller module 102 and receives the first selected sequence VS1 and the second selected sequence VS2, as well as an appearance activation signal, from the controller module 102. The appearance interpretation module 114 is also communicatively coupled to the guidance module 130, as described below. Upon receiving the appearance activation signal, the appearance interpretation module 114 is adapted to communicate the first selected sequence VS1 and the second selected sequence VS2 to the object detection module 116.
[0111] Object Detection Module 116 A customer may approach the self-checkout device 10 with store products and their personal belongings (e.g., handbag, carry-on bag, wallet, cell phone, etc.). When placing the products on the detection plate, the customer may inadvertently place their personal belongings within the field of view of one or both of the first camera 26 and the second camera 28. The purpose of the object detection module 116 is to detect the presence of objects other than the detection plate in the received video frames, determine whether the detected objects are products, and, if all of the detected objects are products, determine the location of the products.
[0112] For purposes of the present invention, the object detection module 116 implements an object detection algorithm configured to receive video frames from the received first selected sequence VS1 and detect the presence of objects in the video frames. The object detection algorithm is further configured to classify the detected objects as either "products for sale" or "other," where objects classified as "other" may include customer belongings. Similarly, the object detection algorithm is configured to receive video frames from the second selected sequence VS2 and detect the presence of objects in the video frames and classify the detected objects as either "products for sale" or "other." Thus, for a given video frame from the first selected sequence VS1,
[0113]
number
[0114] (where,
[0115]
number
[0116] The object detection module 116 generates a first label vector
[0117]
number
[0118] where:
[0119]
number
[0120] is the video frame from the first selected sequence VS1
[0121]
number
[0122] is the number of objects detected in
[0123]
number
[0124] is the label corresponding to the classification of the j-th detected object.
[0125] Similarly, for a given video frame from a second selected sequence VS2,
[0126]
number
[0127] (where,
[0128]
number
[0129] ),
[0130] The object detection module 116 generates a second label vector
[0131]
number
[0132] where:
[0133]
number
[0134] is a video frame from the second selected sequence VS2
[0135]
number
[0136] is the number of objects detected in
[0137]
number
[0138] is the label corresponding to the classification of the j-th detected object.
[0139] The object detection algorithm is also configured to determine the coordinates of a bounding box positioned to enclose the detected object in the video frame. The bounding box coordinates are established relative to the coordinate system of the received video frame of the first selected sequence VS1 or the second selected sequence VS2, as appropriate. Specifically, for a given video frame from the first selected sequence VS1,
[0140]
number
[0141] (where,
[0142]
number
[0143] ), the object detection algorithm uses bounding boxes
[0144]
number
[0145] configured to output one or more details of the set of
[0146]
number
[0147] is the video frame
[0148]
number
[0149] is the number of objects detected in
[0150]
number
[0151] is the bounding box surrounding the j-th detected product.
[0152] Similarly, for a given video frame from a second selected sequence VS2,
[0153]
number
[0154] (where,
[0155]
number
[0156] ), the object detection algorithm uses bounding boxes
[0157]
number
[0158] configured to output one or more details of the set of
[0159]
number
[0160] is the video frame
[0161]
number
[0162] is the number of objects detected in
[0163]
number
[0164] is the bounding box surrounding the j-th detected object.
[0165] Each bounding box
[0166]
number
[0167] The details of include four variables, namely [x,y], h and w, where [x,y] are the time domains of the video frame.
[0168]
number
[0169] are the coordinates of the upper left corner of the bounding box relative to the upper left corner of the object (coordinates are [0,0]), and h and w are the height and width of the bounding box, respectively.
[0170] If the number of detected objects in a video frame is more than 6 (i.e.,
[0171]
number
[0172] ), the object detection module 116 is adapted to issue an "excess object alert" signal to the controller module 102. Those skilled in the art will appreciate that the above object counts, which trigger the issuance of an excess object alert signal in a video frame, are provided for illustrative purposes only. In particular, those skilled in the art will appreciate that the self-checkout device 10 of the present disclosure is not limited to the number of objects detected in a video frame that triggers the issuance of an excess object alert signal. Conversely, the self-checkout device 10 of the present disclosure is operable to trigger the issuance of an excess object alert signal for any number of objects detected in a video frame, thereby meeting the need to maximize the probability of successful product identification while maximizing the number of products that can be simultaneously identified in this manner.
[0173] In one or more embodiments, the object detection algorithm implements a deep neural network whose architecture is substantially based on EfficientDet (described in M. Tan, R. Pang and Q.V. Le, EfficientDet: Scalable and Efficient Object Detection, 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2020, pp. 10778-10787). The deep neural network architecture may also be based on YOLOv4 (described in A. Bochkovskiy, C.Y. Wang and H.Y. M. Liao, 2020 arXiv: 2004.10934). However, those skilled in the art will understand that these object detection deep neural network architectures are provided for illustrative purposes only. In particular, those skilled in the art will understand that the self-checkout device 10 of the present invention is not limited to these deep neural network architectures. Conversely, the self-checkout device 10 is operable with any object detection architecture and / or training algorithm suitable for detecting, classifying, and locating products within an image or video frame.
[0174] It will be appreciated that the goal of training the deep neural network is for the deep neural network to establish an internal representation of products / loose items that will enable the deep neural network to recognize the presence of products / loose items in subsequently received video footage. To this end, the dataset used to train the deep neural network is comprised of video footage of various scenarios in which one or more of each of the products / loose items from inventory in a retail environment and / or other objects are placed on the detection plate of the self-checkout device 10. Specifically, the video footage may include scenarios in which one or more personal items are placed alone on the detection plate 32, scenarios in which one or more products / loose items are placed alone on the detection plate 32, and scenarios in which one or more products / loose items and one or more personal items are placed on the detection plate 32.
[0175] The video footage, hereafter referred to as the training dataset, is assembled with the goal of providing robust, class-balanced information about the target product / loose product from different views of the product / loose product acquired at different positions and orientations of the product relative to the first and second cameras 26, 28, the different positions and orientations of the product / loose product representing the intended use environment of the self-checkout device 10. The members of the training dataset are selected to create sufficient diversity to overcome challenges to subsequent product / loose product recognition posed by changing lighting conditions, changing viewpoints, cluttered backgrounds, and most importantly, intra-class variation.
[0176] In one or more examples, before use in the training dataset, the video footage is processed to remove highly similar video frames / images. Further data augmentation techniques (such as rotation, flipping, or brightness changes) may be applied to the members of the training dataset to increase their diversity and thereby the robustness of the final trained object detection model. In a further preprocessing step, each image / video frame in the video footage of the training dataset is further provided with a bounding box, with each bounding box positioned to enclose an object visible in the image / video frame. Each image / video frame is also optionally provided with a "product" or "other" label corresponding to each bounding box in the respective image / video frame.
[0177] The object detection module 116 is further configured to concatenate the bounding box details of each object detected in the video frame and the corresponding label classification of the detected object to form a detected object vector. Specifically, the output from the object detection module 116 is one or more first detected object vectors:
[0178]
number
[0179] and one or more second detected object vectors.
[0180]
number
[0181] where the object detection module 116 is further configured to communicate this output to the controller module 102.
[0182] Cutting Module 118 The cropping module 118 is communicatively coupled to the controller module 102 and receives the first selected sequence VS1 and the second selected sequence VS2 therefrom. The cropping module 118 is also configured to receive the first product vector PV1(y), the second product vector PV2(y), and the selection timestamp vector STS(y) from the controller module 102. The cropping module 118 is adapted to crop one or more first cropping regions from each video frame of the first selected sequence VS1, the perimeters of which are established by the bounding box coordinates of the first product vector PV1(y) whose timestamps determined from the selection timestamp vector STS(y) coincide with the timestamps of the video frames. The cropping module 118 is further adapted to crop one or more second cropping regions from each video frame of the second selected sequence VS2, the perimeters of which are established by the bounding box coordinates of the second product vector PV2(y) whose timestamps determined from the selection timestamp vector STS(y) coincide with the timestamps of the video frames.
[0183] The cropping module 118 is further configured to resize each first crop region and each second crop region to the same predetermined size, which has been empirically established as the size that provides optimal product / loose product recognition by the embedding module 120, as described in the preceding paragraph. For clarity, this size is hereinafter referred to as the "processed image size." Additionally, data augmentation techniques (such as rotation, flipping, brightness change, etc.) may optionally be applied. The cropping module 118 is further configured to send the resulting first crop region and second crop region to the embedding module 120.
[0184] Embedded Module 120 The embedding module 120 is coupled to the cropping module 118 and receives the first crop region and the second crop region therefrom during runtime operation. The embedding module 120 uses a deep metric learning module, reviewed in K. Musgrave, S. Belongie, and S.-N. Li, A Metric Learning Reality Check (retrieved August 19, 2020, from https: / / arxiv.org / abs / 2003.08505), to learn a unique representation in the form of an embedding vector from each product and loose product video frame in the store's inventory. This allows for identification of either or both of the product and loose product that subsequently appear in a video frame of the first video stream VID1 or the second video stream VID2. Specifically, the deep metric learning module is configured to generate embedding vectors in response to product / loose product images, where the embedding vectors are close to each other (in the embedding space) if the images contain the same product, and are far apart as measured by a similarity or distance function (e.g., dot-product similarity or Euclidean distance) if the images contain different products. A query image can then be validated based on a threshold similarity or distance in the embedding space.
[0185] During use, the embedded module 120 has two distinct phases of operation: an initial configuration phase and a run-time phase.
[0186] Initial Configuration of Embedded Module 120 During the initial configuration stage, the embedding module 120 identifies the products / loose products in the store's inventory. i One or more embedding vectors E that form a unique representation of iThe network is trained to learn the following: (a) the initial construction stage includes several distinct phases: a training data preparation phase and a network training phase. These phases are implemented sequentially in a cyclical, iterative manner to train the embedding module 120. Each of these phases is described in detail in the previous paragraphs.
[0187] Training data preparation phase The dataset used to train the embedding module 120 includes video footage of a scenario in which one or more of each of the products / loose items from the inventory of a retail environment are placed on the detection plate of the self-checkout device 10. This video footage, hereafter referred to as the training dataset, is assembled with the goal of providing robust, class-balanced information about the target products / loose items obtained from various views of the products / loose items acquired at different positions and orientations of the products / loose items relative to the first and second cameras 26 and 28, where the different positions and orientations of the products / loose items are representative of the intended use environment of the self-checkout device 10. The members of the training dataset are selected to create sufficient diversity to overcome subsequent product / loose item recognition challenges caused by changing lighting conditions, changing viewpoints, cluttered backgrounds, and, most importantly, intra-class variation.
[0188] In one or more examples, video footage is processed to remove highly similar video frames before use in the training dataset. Members of the training dataset may also be subjected to data augmentation techniques (such as rotation, flipping, or brightness modification) to increase diversity, thereby enhancing the robustness of the final trained deep neural network in the embedding module 120. Polygonal regions encompassing individual products / loose products visible in the video frames are cropped therefrom. The cropped regions are resized to fit the processed image size, generating cropped product / loose product images. Each cropped product / loose product image is also provided with a class label identifying the corresponding product / loose product.
[0189] Model training phase For simplicity, the deep neural network (not shown) in embedding module 120 is hereafter referred to as an "embedded neural network (ENN)." An embedded neural network includes a deep neural network (e.g., ResNet, Inception, EfficientNet) in which the last layer or layers (which typically output a classification vector) are replaced with a linear normalization layer that outputs a unit-norm (embedding) vector of the required dimensionality (the dimensionality is a parameter set when creating the embedded neural network).
[0190] During the model training phase, positive and negative pairs of cropped / loose product images are constructed from the training dataset. A positive pair contains two cropped / loose product images with the same class label, and a negative pair contains two cropped / loose product images with different class labels. For brevity, the resulting cropped / loose product images are hereafter referred to as "paired cropped images." Paired cropped images are sampled according to a pair mining strategy (e.g., MultiSimilarity or ArcFace as outlined in R. Manmatha, C.-Y. Wu, A.J. Smoia, and P. Krahenbuhl, "Sampling Matters in Deep Embedded Learning," 2017 IEEE International Conference on Computer Vision (CCV2017) Venice, 2017, pp. 2859-2867, doi: 10.1109 / ICCV.2017.309). Next, a pairwise metric learning loss is calculated from the sampled paired video frames (as described in K. Musgrave, S. Belongie and S.-N. Lim, A Metric Learning Reality Check, 2020, https: / / arxiv.org / abs / 2003.08505). The weights of the embedding neural network are then optimized using a backpropagation approach that minimizes the pairwise metric learning loss value.
[0191] Every pair of cropped images is processed by the embedding neural network to generate a corresponding embedding vector. The resulting embedding vectors are organized in pairs, just like the pair of cropped images. The resulting embedding vectors are stored in the embedding database 124. Thus, given an image of each product / loose item in the store's inventory, the trained embedding neural network generates an embedding vector E computed for each product / loose item. i are input into the embedding database 124. Thus, the embedding database 124 contains multiple sets of embedding vectors (E i , Id i ) and all products / loose products in store inventory i The embedding vectors and corresponding identifiers Idi are included.
[0192] Run-time phase of the embedding module 120 For clarity, runtime is defined as the normal business hours of the associated store. During runtime, the embedding neural network (not shown) generates an embedding vector for each product / loose item visible in the video frames captured by the first and second cameras 26 and 28 of the product / loose item placed at the self-checkout device 10. Accordingly, the embedding vector generated from the first cropped region received from the video footage captured by the first camera 26 is hereinafter referred to as the first query embedding QE1. Similarly, the embedding vector generated from the second cropped region received from the video footage captured by the second camera 28 is hereinafter referred to as the second query embedding QE2. The embedding module 120 is communicatively coupled to the expert system module 122 and transmits either or both of the first query embedding QE1 and the second query embedding QE2 to the expert system module 122.
[0193] Expert System Module 122 The expert system module 122 is coupled to the embedding module 120 and receives either or both of the first query embedding QE1 and the second query embedding QE2 generated by the embedded neural network during the runtime operation of the embedding module 120. Upon receiving the first query embedding QE1 or the second query embedding QE2, the expert system module 122 queries the embedding database 124 and derives therefrom an embedding vector E i The expert system module 122 uses a similarity or distance function (e.g., dot product similarity or Euclidean distance) to convert the first query embedding QE1 or the second query embedding QE2 into an embedding vector E i The expert system module 122 compares each first query embedding QE1 or second query embedding QE2 with the obtained embedding vector E using a similarity or distance function (e.g., dot product similarity or Euclidean distance). i A preset percentage (Per) of the plurality of first query embeddings QE1 of the corresponding plurality of first product vectors PV1(y) from the first selected sequence VS1 is compared with the obtained embedding vector E i If the similarity between the first query embedding QE1 and the second query embedding QE2 exceeds a preset threshold (Th), the first query embedding QE1 is obtained as the embedding vector E i Similarly, a predetermined percentage (Per) of the plurality of second query embeddings QE2 of the corresponding plurality of second product vectors PV2(y) from the second selected sequence VS2 and the obtained embedding vector E i If the similarity between the two exceeds a preset threshold (Th), the second query embedding QE2 is used to obtain the embedding vector E i It can be concluded that it is consistent with
[0194] The values of the percentage (Per) and threshold (Th) parameters used by the expert system module 122 are established using a grid search method, with targets empirically defined according to the operator's risk tolerance, balanced with the operator's desire to speed up the checkout process at the self-checkout device 10 of the present disclosure. Additionally, the embedding database 124 is queried to obtain the embedding vector E i The process of comparing QE1 with the received first query embedding QE2 and QE2 continues until a match is found or all embedding vectors E from the embedding database 124 are compared. i The process is repeated until the first query embedding QE1 and the embedding vector E from the embedding database 124 are obtained. i If a match is found between i is hereafter referred to as the first matching embedding ME1. Similarly, the multiple second query embeddings QE2 and the embedding vectors E from the embedding database 124 are i If a match is found between i is hereafter referred to as the second matching embedding ME2. In contrast, the first query embedding QE1 and the embedding vector E from the embedding database 124 i If no match is found between the second query embedding QE2 and the embedding vector E from the embedding database 124, i If no match is found between the expert system module 122 and the controller module 102, the expert system module 122 is configured to issue an "unconfirmed product alert" signal to the controller module 102.
[0195] The expert system module 122 is further adapted to use the first matching embedding to retrieve from the product database 110 a product identifier corresponding to the first matching embedding ME1, where the product identifier is the identifier of the product / loose product represented by the first matching embedding ME1. For simplicity, this product identifier is hereinafter referred to as the first matching class label. Similarly, the expert system module 122 is also adapted to retrieve from the product database 110 a product identifier corresponding to the second matching embedding ME2, where the product identifier is the identifier of the product / loose product represented by the second matching embedding ME2. For simplicity, this product identifier is hereinafter referred to as the second matching class label.
[0196] The expert system module 122 is also configured to compare each first matching class label with all second matching class labels. If either the first matching class label or the second matching class label matches, the expert system module 122 is configured to retrieve a price corresponding to the first matching class label or the second matching class label from the product database 110. The expert system module 122 is further configured to send the first matching class label and the corresponding price to the controller module 102. If either the first matching class label or the second matching class label matches, the expert system module 122 is further configured to determine whether the first matching class label corresponds to a weight-dependent bulk product item. If the first matching class label corresponds to a weight-dependent bulk product item, the expert system module 122 is further configured to determine whether the first matching class label and the second matching class label corresponding to the first and second clipping regions of the first product vector PV1(y) and the second product vector PV2(y) include items that are not weight-dependent bulk product items. In other words, the expert system module 122 is configured to determine whether both a weight-dependent bulk product item and a product whose price is not weight-dependent are placed on the detection plate 32 of the self-checkout device 10 at the same time. If this occurs, the expert system module 122 is configured to issue a "mixed product alert" signal to the controller module 102. Otherwise, the expert system module 122 is configured to issue a "weight-dependent product" signal to the controller module 102. In contrast, if either the first matching class label and the second matching class label do not match, the expert system module 122 is configured to issue a "product mismatch alert" signal to the controller module 102.
[0197] Guidance Module 130 The guidance module 130 is communicatively coupled to the interaction unit 14 and the controller module 102. In use, the guidance module 130 is operable to activate the design display unit to display a design on the detection plate 32 before motion is detected near the self-checkout device 10 or the detection plate 32. The design includes at least two circles or ovals positioned in opposite quadrants of the top surface of the detection plate 32. The circles may be configured to be sized sufficiently to enclose the bottom of an average bottle or can (as shown in FIG. 10 ). Those skilled in the art will appreciate that the above design is provided for illustrative purposes only. In particular, those skilled in the art will appreciate that the self-checkout device 10 of the present disclosure is not limited to the details of the above pattern. Conversely, the self-checkout device 10 of the present disclosure can be operated with any design before motion is detected near the self-checkout device 10 or the detection plate 32, with the design serving to provide initial guidance regarding product placement on the detection plate 32. For example, the design may include a polygon rather than a circle or oval. Similarly, the design may include four or more circles, ovals, or polygons spaced approximately evenly across all four quadrants of the detection plate.
[0198] Upon receiving a motion trigger signal from the motion detection module 104, the guidance module 130 operates to receive the first product vector PV1(y) and the second product vector PV2(y) from the controller module 102. The guidance module 130 is configured to check the bounding boxes of the first product vector PV1(y) and the second product vector PV2(y) and, if there is more than one detected product, determine whether the corresponding bounding boxes are located in one or more quadrants of the detection plate 32. If the guidance module 130 detects otherwise, it suggests that the user has placed all of the products in only one quadrant of the detection plate 32. This may obstruct the view of one or more products and hinder product identification. Therefore, if the guidance module 130 detects that the bounding boxes of two or more detected products are located in one quadrant of the detection plate 32, the guidance module 130 is adapted to issue a “quadrant misuse warning” signal to the controller module 102.
[0199] The guidance module 130 is also operable to examine the bounding boxes of the first product vector PV1(y) and the second product vector PV2(y) and compare the distance between adjacent bounding boxes. If the distance is less than a predetermined threshold, hereinafter referred to as the bounding box separation threshold, the guidance module 130 is configured to issue a “product distance warning” signal to the controller module 102. The “product distance warning” includes coordinates of bounding boxes separated by a distance less than the predetermined threshold. For simplicity, these bounding boxes are hereinafter referred to as “too close bounding boxes.” Thus, the “product distance warning” includes coordinates of bounding boxes that are too close. The bounding box separation threshold is empirically set depending on the size of products typically sold by the operator, lighting conditions in the retail environment, and other conditions that may hinder the performance of the self-checkout device 10. The predetermined distance is configured to balance the requirement to maximize product identification accuracy with the speed improvement achievable by allowing a customer to simultaneously place multiple products on the detection plate 32 and have them registered by the self-checkout device 10.
[0200] Weighing Module 126 The scale module 126 is communicatively coupled to the controller module 102 and receives an activation signal therefrom. Upon receiving the activation signal, the scale module 126 is configured to activate a weigh scale unit (not shown) of the self-checkout device 10 to measure the weight of the weight-dependent bulk product item. The scale module 126 is further configured to transmit the value of the weight measurement to the controller module 102.
[0201] Billing Module 128 The billing module 128 is configured to receive from the controller module 102 the price of each product whose barcode is detected in the first selected sequence VS1 or the second selected sequence VS2. Alternatively, the billing module 128 is configured to receive from the controller module 102 the price of each product whose barcode is detected by the 1D barcode reader 38 of the self-checkout device 10. The billing module 128 is configured to sum these prices to calculate a total bill for the products. Otherwise, if the barcode processing module 108 fails to detect the presence of a barcode in the video frames of the first selected sequence VS1 or the second selected sequence VS2, the billing module 128 is configured to receive from the controller module 102 the price of each product / bulk product recognized in the first selected sequence VS1 and the second selected sequence VS2 by the appearance interpretation module 114, where the price is the price corresponding to the first matching class label determined by the expert system module 122. The billing module 128 is further configured to receive from the controller module 102 weight measurements of the detected weight-dependent bulk product items. The billing module 128 is further configured to calculate a total billing amount for the weight-dependent bulk product items from the weight measurements of the weight-dependent bulk product items and the price per unit weight of the weight-dependent bulk product items received from the controller module 102. The billing module 128 is further configured to calculate a total billing amount for the products by summing the prices of all products recognized by the appearance interpretation module 114 in the first selected sequence VS1 and the second selected sequence VS2.
[0202] Controller Module 102 In this embodiment, the controller module 102 is configured to receive a motion trigger signal from the motion detection module 104 and communicate the motion trigger signal to the sequence selection module 106 to activate the sequence selection module 106. The controller module 102 is also configured to receive a first selected sequence VS1 and a second selected sequence VS2 from the sequence selection module 106 and communicate the first selected sequence VS1 and the second selected sequence VS2 to the barcode processing module 108. The controller module 102 is configured to receive an appearance activation signal from the barcode processing module 108 and communicate the appearance activation signal to the appearance interpretation module 114 to activate the appearance interpretation module 114.
[0203] The controller module 102 receives a first detected object vector from the object detection module 116.
[0204]
number
[0205] and a second detected object vector
[0206]
number
[0207] a first detected object vector
[0208]
number
[0209] The first label vector of
[0210]
number
[0211] or the second detected object vector
[0212]
number
[0213] The second label vector of
[0214]
number
[0215] includes "other," the controller module 102 is configured to display a message on the display screen 34 of the interaction unit 14, the message alerting the user that a non-sales product has been placed on the detection plate 32 and needs to be removed therefrom. The controller module 102 is further operable to activate the motion detection module 104 to detect motion in an area near the self-checkout device 10 and the detection plate 32. If a motion trigger signal is not received from the motion detection module 104 within a predetermined time interval, the controller module 102 is operable to alert an operator to indicate that the customer needs assistance. For simplicity, this predetermined time interval is hereinafter referred to as the "non-sales object reset period." However, upon receiving a motion trigger signal from the motion detection module 104 within the non-sales object reset period, the controller module 102 is operable to activate the object detection module 116 to check a further first selected sequence VS1 and a further second selected sequence VS2. The resulting first detected object vector
[0216]
number
[0217] The first label vector of
[0218]
number
[0219] or the resulting second detected object vector
[0220]
number
[0221] The second label vector of
[0222]
number
[0223] includes "other," the controller module 102 is further operable to issue an alert to an operator to indicate that the customer needs assistance. Similarly, upon receiving an object excess alert signal from the object detection module 116, the controller module 102 is configured to display a message on the display screen 34 of the interaction unit 14, the message alerting the user that some of the products placed on the detection plate 32 need to be removed.
[0224] For ease of understanding, for a given first selected sequence VS1 or a given further first selected sequence VS1, the video frames
[0225]
number
[0226] Let np1 contain objects that are labeled only as sales products, i.e.,
[0227]
number
[0228] Similarly, for a given second selected sequence VS2 or a given further second selected sequence VS2,
[0229]
number
[0230] Let np2 contain objects that are labeled only as sales products, i.e.,
[0231]
number
[0232] Furthermore, the first intermediate product vector IPV1(w1) (1≦w1≦np1) is a first label vector
[0233]
number
[0234] The first detected object vector contains only the "Sales Product" element
[0235]
number
[0236] and the first intermediate timestamp TS1(w1) (1≦w1≦np1) is defined as the first detected object vector
[0237]
number
[0238] timestamp
[0239]
number
[0240] is defined as:
[0241] Thus, for a given first intermediate product vector IPV1(w1) (1≦w1≦np1), there is a matching first intermediate timestamp TS1(w1) (1≦w1≦np1), and a second intermediate product vector IPV2(w2) (1≦w2≦np2) has a matching second label vector TS1(w2) (1≦w2≦np2).
[0242]
number
[0243] A second detected object vector containing only the "Sales Products" element
[0244]
number
[0245] and the second intermediate timestamp TS2(w2) (1≦w2≦np2) is defined as the first detected object vector
[0246]
number
[0247] timestamp
[0248]
number
[0249] Thus, for a given second intermediate product vector IPV2(w2), where 1 ≤ w2 ≤ np2, there is a matching second intermediate timestamp TS2(w2), where 1 ≤ w2 ≤ np2. Ideally, np1 should match np2, but depending on the complexity of the displayed scene, either the first or second camera 26, 28 may detect the presence of an object on the detection plate that is not detected by the other camera.
[0250] Here, the controller module 102 is configured to compare the first intermediate timestamp TS1(w1) (1≦w1≦np1) with the second intermediate timestamp TS2(w2) (1≦w2≦np2). The controller module 102 is further configured to select a first intermediate product vector IPV1(w1) (1≦w1≦np1) and a second intermediate product vector IPV2(w2) (1≦w2≦np2) whose corresponding first intermediate timestamp TS1(w1) (1≦w1≦np1) matches the corresponding second intermediate timestamp TS2(w2) (1≦w2≦np2). For brevity, the selected first intermediate product vector IPV1(w1) and second intermediate product vector IPV2(w2) are referred to as first product vector PV1(y) and second product vector PV2(y), respectively, where 1≦y≦sel, sel is the number of selected first product vectors PV1(y) or second product vectors PV2(y), and sel≦min(np1, np2). The timestamps of each first product vector PV1(y) and corresponding second product vector PV2(y) are compiled into a selection timestamp vector STS(y). This approach is employed to exclude video frames with transient changes, such as light flicker or fast movement, that were not detected by the motion detection module 104 from subsequent consideration by the self-checkout device 10.
[0251] The controller module 102 is configured to send the first product vector PV1(y), the second product vector PV2(y), and the selection timestamp vector STS(y) to the guidance module 130. If the controller module 102 does not receive a quadrant misuse warning or a product distance warning from the guidance module 130, the controller module 102 is configured to send the first product vector PV1(y), the second product vector PV2(y), and the selection timestamp vector STS(y) to the cutting module 118. However, if the controller module 102 receives a quadrant misuse warning signal from the guidance module 130, the controller module 102 is configured to activate the design display unit of the interaction unit 14 to change the design displayed on the detection plate 32 so that a circle, an oval, or a polygon is highlighted in each of the four quadrants of the detection plate 32. The controller module 102 is further configured to display a message on the display screen 34 of the interaction unit 14, alerting the user that a portion of the product placed on the detection plate 32 needs to be moved to another quadrant. The controller module 102 is further operable to activate the motion detection module 104 to detect product movement. If a motion trigger signal is not received from the motion detection module 104 within a predetermined time interval, the controller module 102 is operable to issue an alert to an operator, indicating that the customer requires operator assistance. For simplicity, this predetermined time interval is hereinafter referred to as the "first product movement reset period." However, upon receiving a motion trigger signal from the motion detection module 104 within the first product movement reset period, the controller module 102 is operable to activate the guidance module 130 to confirm placement of the product's bounding box on the detection plate 32. If the guidance module 130 reissues a quadrant misuse alert, the controller module 102 is operable to issue an alert to an operator, indicating that the customer requires operator assistance.Here, the first product movement reset period and the second product movement reset period are each empirically determined according to the operator's requirements, and minimize delays in the registration process while ensuring sufficient time for the user to move the product to a new location.
[0252] Upon receiving the product distance alert, the controller module 102 is configured to use the coordinates of the too-close bounding box included in the product distance alert to identify a circle, oval, or polygon displayed on the detection plate 32 that is closest to the product corresponding to the “too-close bounding box.” For simplicity, each of these circles, ovals, or polygons is hereinafter referred to as the “nearest guidance handle.” The controller module 102 is further configured to activate the design display unit of the interaction unit 14 to move the nearest guidance handle to another location within a predetermined distance from the coordinates of the too-close bounding box, such that the nearest guidance handle is separated by a distance that exceeds the bounding box separation threshold. For clarity, the predetermined distance of movement of the nearest guidance handle is hereinafter referred to as the guidance handle movement distance. The guidance handle movement distance is empirically determined based on the need to balance the requirement to maximize product identification accuracy with the speed improvement achievable by allowing a customer to simultaneously place multiple products on the detection plate and register them simultaneously with the self-checkout device 10.
[0253] The controller module 102 is further configured to display a message on the display screen 34 of the interaction unit 14, alerting the user that the product closest to the nearest guidance handle needs to be moved to that position. The controller module 102 is operable to activate the motion detection module 104 to detect product movement. If a motion trigger signal is not received from the motion detection module 104 within a predetermined time interval, the controller module 102 is operable to issue an alert to an operator indicating that the customer needs assistance. For simplicity, the predetermined time interval is hereinafter referred to as the "second product movement reset period." However, upon receiving a motion trigger signal from the motion detection module 104 within the second product movement reset period, the controller module 102 is operable to activate the guidance module 130 to confirm the placement of the product's bounding box on the detection plate. If the guidance module 130 reissues the product distance alert, the controller module 102 is operable to issue an alert to an operator indicating that the customer needs assistance.
[0254] The controller module 102 is further configured to receive an unidentified product alert signal, a mixed product alert signal, or a product mismatch alert signal from the expert system module 122. Upon receiving any of these, the controller module 102 is further configured to activate the display screen 34 of the self-checkout device 10 to prompt the user to present the product to the 1D barcode reader 38 of the self-checkout device 10 to read the product's barcode. If the 1D barcode reader 38 is unable to read the product's barcode, the controller module 102 is configured to issue an alert to the operator indicating that the customer requires operator assistance.
[0255] The controller module 102 is further configured to receive a product label and its corresponding price from the barcode processing module 108 or a first matching class label and its corresponding price from the expert system module 122. The controller module 102 is further configured to receive a “weight-dependent product” signal from the expert system module 122, which, upon receiving the signal, causes the self-checkout device 10 of the present disclosure to either activate the scale module 126 if a scale unit is present or to issue an alert to an operator indicating that the customer requires operator assistance. The controller module 102 is further configured to receive a weight measurement from the scale module 126. The controller module 102 is further configured to transmit the price received from the barcode processing module 108, or the price received from the 1D barcode reader 38 of the self-checkout device 10, or the price received from the expert system module 122, and the weight measurement from the scale module 126, if available, to the billing module 128. In response, the controller module 102 is further configured to receive from the billing module 128 a total bill for the products registered by the self-checkout device 10 of the present disclosure.
[0256] Additionally, the controller module 102 is configured to operate the display screen 34 of the self-checkout device 10 to display either (a) the product label and corresponding price received from the barcode processing module 108 or 1D barcode reader (not shown) of the self-checkout device 10, or (b) the first matching class label and corresponding price from the expert system module 122. The controller module 102 is configured to operate the display screen 34 of the self-checkout device 10 to display an itemized list of products along with the total bill amount (shown in FIG. 10 ).
[0257] Alternatively, if weight-dependent bulk product items are recognized in the first selected sequence VS1 and the second selected sequence VS2 by the appearance interpretation module 114, the billing module 128 is configured to activate the display screen 34 of the self-checkout device 10 to display the first matching class label and the corresponding total bill amount of the weight-dependent bulk product items (as shown in FIG. 10 ), where the billing module 128 is further configured to activate the contactless card reader 40 to receive payment of the total bill amount of the recognized products or the total bill amount of the weight-dependent bulk product items.
[0258] Management Module 132 The management module 132 is adapted to allow an operator to access the software of the self-checkout device 10 to update either or both of the software and its configuration, for example, to update tuples in the product database 110 or to update the training of the embedded neural network of the embedding module 120 or the deep neural network of the object detection module 116. The management module 132 may include a PIN function or other access control mechanism to restrict access to the software of the self-checkout device 10 to certain operators.
[0259] The above description of the first embodiment of the software for the self-checkout device 10 (as shown in FIG. 6 ) focused on a standalone implementation in which a single specific self-checkout device 10 is provided with its own product database 110 and embedded database 124. The single self-checkout device 10 could operate without interacting with other infrastructure or other self-checkout devices 10 within the store. The process of updating the software and configuration of an individual self-checkout device 10, including training the embedded neural network and object detection neural network, is feasible when there are only a few self-checkout devices in the store. However, in a large store with multiple such self-checkout devices 10, the process of updating the software and configuration of each individual self-checkout device 10 becomes problematic. In such cases, the multiple self-checkout devices 10 are individually and communicatively coupled to a central controller that includes a centralized software update scheduler and a centralized record of the store's total product / loose product inventory. The central controller issues software configuration updates that include updates to the internal representations / embeddings formed by the embedded neural network and object detection neural network.
[0260] Referring now to FIG. 7, a schematic block diagram of a system 100 including multiple self-checkout devices 10 according to a second exemplary embodiment of the present disclosure is shown. As shown in FIG. 7, in combination with FIGS. 1-6 described in the previous paragraph, the system 700 includes individual self-checkout devices 10 communicatively coupled in a distributed network. The embedded neural networks and object detection neural networks of the self-checkout devices 10 can be trained using different subsets of products / loose products in the store's inventory, and the product databases 110 and embedded databases 124 of the individual self-checkout devices 10 can be populated accordingly. The individual self-checkout devices 10 are configured to share the embedded representations formed by the embedded neural networks and object detection neural networks, respectively, with each other. The individual self-checkout devices 10 are further configured to share members of the product databases 110 and embedded databases 124 with each other. The system 700 eliminates the need to individually train the embedded neural network and object detection neural network of each self-checkout device 10 using members of the store's total product / loose product inventory, while also eliminating the need to maintain a centralized software update scheduler and a centralized record of the store's total product / loose product inventory.
[0261] Referring now to FIG. 8, a flowchart of a method 200 implemented by the self-checkout device 10 of the present disclosure is shown. Method 200 will now be described with reference to the components defined in FIGS. 1-7. In step 202, method 200 includes receiving video footage including a plurality of video frames from each of first and second cameras 26 and 28. In step 204, method 200 includes detecting the presence of motion in the received video footage. In step 206, method 200 includes selecting a predetermined number of video frames from the received video footage following detection of the end of motion in the received video footage. In step 208, method 200 includes detecting and decoding a barcode displayed in the selected video frames. In step 210, method 200 includes calculating a total bill amount corresponding to the decoded barcode. In step 212, method 200 includes detecting the presence of an object displayed in the selected video frames if a barcode is not displayed in the selected video frames. At step 214, method 200 includes distinguishing between sale items and non-sale items in the detected objects. At step 216, method 200 includes issuing an alert upon detection of a non-sale item, the alert including a message to remove the non-sale item from self-checkout device 10. At step 218, method 200 includes determining locations of the detected sale items. At step 220, method 200 includes determining a distribution of the detected sale items from the determined locations. At step 222, method 200 includes issuing an alert upon detecting improper distribution of the detected sale items. At step 224, method 200 includes cropping one or more regions from each received video frame that substantially enclose each detected sale item. At step 226, method 200 includes generating an embedded representation of the sale item visible therein from each cropped region.At step 228, method 200 includes comparing the generated embedded representation to records of embedded representations of products contained in the retail environment's product inventory to find a match with any of the members of the records. At step 230, method 200 includes determining a price corresponding to the member matching the record of embedded representations of products contained in the retail environment's product inventory. At step 232, method 200 includes calculating a total bill for all products / loose products appearing in the received video footage that match members of the record of embedded representations of products contained in the retail environment's product inventory. At step 234, method 200 includes displaying the total bill to the user. At step 236, method 200 also includes receiving payment of the total bill from the user. It will be understood that steps 202 through 236 described herein are merely exemplary, and other alternatives, such as adding one or more steps, deleting one or more steps, or providing one or more steps in a different order, may be provided without departing from the scope of the claims herein.
[0262] In one or more examples, receiving 202 video footage including a plurality of video frames from each of first and second cameras 26 and 28 precedes displaying a design on detection plate 32 of self-checkout device 10. Preferably, displaying a design includes displaying one or more of a circle, an oval, or a polygon in each quadrant of detection plate 32.
[0263] In one or more examples, detecting 204 the presence of motion in the received video footage includes using motion vectors obtained from decoding the H.264 video frames, or alternatively, detecting 204 the presence of motion in the received video footage includes the substeps of comparing consecutive samples of the video frames received from the first camera 26 to detect differences therebetween, comparing consecutive samples of the video frames received from the second camera 28 to detect differences therebetween, and determining that motion occurred in the period between the consecutive samples if the detected difference exceeds a predetermined threshold.
[0264] In one example, the threshold is set so that a temporary change, such as a flickering light, is mistaken for movement. In one example, the interval between successive samples of video frames received from the first camera 26 and the second camera 28 is set to be of sufficient duration to avoid falsely detecting small, fast movements, such as finger movements, rather than larger movements corresponding to the placement or removal of a product from the detection plate.
[0265] In one or more examples, if a barcode does not appear in the selected video frame, step 212 of detecting the presence of an object displayed in the selected video frame and step 214 of distinguishing between for-sale and non-for-sale items in the detected object may include the following substeps prior to step 204 of detecting the presence of motion in the received video footage: training an object detection neural network model (e.g., training an EfficientDet neural network) using a labeled training dataset to form an internal representation of for-sale and non-for-sale items placed on the detection plate 32 of the self-checkout device 10; presenting the selected video frame to the trained object detection neural network model to form a representation of the for-sale and non-for-sale items visible therein; and obtaining from the object detection neural network model labels corresponding to the for-sale items displayed in the selected video frame and coordinates of a bounding box substantially surrounding each of the for-sale items.
[0266] In another example, if no barcode is present in the selected video frame, step 212 of detecting the presence of an object appearing in the selected video frame and step 214 of distinguishing between for-sale and non-for-sale items of the detected object further include the following substeps: - a substep of counting the number of objects detected in any one of the selected video frames; - issuing an alert when it is detected that the number of objects exceeds a predetermined threshold, the alert being a message requesting the removal of the part of the product placed at the time of the detection.
[0267] In one or more examples, issuing an alert upon detecting an improper distribution of the detected sales items 222 includes the following substeps: - issuing a warning when two or more items for sale are placed in only one quadrant of the detection plate 32 of the self-checkout device 10; and - issuing an alarm if the distance between adjacent sales items on the detection plate 32 is less than a predetermined threshold.
[0268] Furthermore, step 222 of detecting and issuing a warning that two or more sales items are placed in only one quadrant of the detection plate 32 of the self-checkout device 10 includes the following substeps: - modifying the design displayed on the detection plate 32 so that a circle, an oval, or a polygon is highlighted in each of the four quadrants of the detection plate 32; - displaying a message on the display screen 34 of the interaction unit 14, the message warning the user that some of the sales items placed on the detection plate 32 need to be moved to another quadrant; - detecting whether a sales item has been moved into one or more quadrants of the detection plate 32; and - issuing an alert to an operator indicating that the customer needs assistance if the sales item has not been moved within a predetermined time interval.
[0269] Furthermore, the step 222 of issuing an alert when detecting that the distance between adjacent sales items on the detection plate 32 is less than a predetermined threshold includes the following substeps: - identifying the circle, oval or polygon of the design displayed on the detection plate 32 that is closest to the adjacent sales item, the distance being less than a predetermined threshold; - moving the identified circle, oval, or polygon to another location within the guidance handle movement distance of an adjacent sales item, the distance being less than a predetermined threshold; - displaying a message on the display screen 34 of the interaction unit 14, the message alerting the user that the nearest sales items to the moved circle, oval or polygon should be moved to those locations; - detecting whether a sales item has been moved to the location of the moved circle, oval, or polygon; and - issuing an alert to an operator if the sale item is not moved within a predetermined time interval, the alert indicating that the customer needs assistance.
[0270] In one or more examples, step 224 of cropping one or more regions substantially surrounding each detected for-sale item from each received video frame includes, if no barcode appears in the selected video frame, step 214 of cropping one or more regions from each received video frame whose perimeter is established by coordinates of a bounding box generated by step 212 of detecting the presence of an object appearing in the selected video frame, and step 214 of distinguishing between for-sale and non-for-sale items of the detected objects.
[0271] In one or more examples, generating 226 an embedded representation from each cropped region of the sales items visible therein includes the following substeps: - before the step 204 of detecting the presence of motion in the received video footage, a sub-step of training an embedding neural network using a training dataset to form an embedding representation of each of the plurality of products / loose products in the store's inventory; - presenting the cropped region to the trained embedding neural network to form an embedding representation of the product or loose product visible in the cropped region; - comparing the embeddings formed from the cropped regions with the embeddings of each product / loose product in the store's inventory; - determining whether the embedding formed from the cropped region matches the embedding of any of the products / loose products in the store's inventory; and - A substep of obtaining the label corresponding to the matching embedding representation and its price or price per unit weight.
[0272] Now, if no match is found between the embedded representation formed from the cropped region and any of the embedded representations of products in the store's inventory, the method 200 includes the following steps: - displaying a message on the display screen 34 of the interaction unit 14, the message requesting the user to present the barcode (if any) of the product displayed in the cropped area to the 1D barcode reader 38 of the self-checkout device 10; - activating the 1D barcode reader 38 to read the presented barcode; and - issuing an alert to an operator if the 1D barcode reader 38 is unable to read the presented barcode, the alert indicating that the customer needs assistance.
[0273] Additionally, if an embedded representation determined to match a crop region generated from video footage from the first camera 26 does not match a corresponding embedded representation determined to match a crop region generated from the second camera 28, the method 200 includes issuing an alert to an operator, the alert indicating that the customer needs assistance.
[0274] Further, if the embedded representation determined to match the cropped region generated from the video footage from the first camera or the second camera 26, 28 is of a loose product, the method 200 includes the following steps: - if the weighing scale unit is integral with the detection plate 32, displaying a message on the display screen 34 of the interaction unit 14, the message requesting the user to remove from the detection plate 32 of the self-checkout device 10 any remaining products placed on the detection plate 32; - if the weighing scale unit is a separate component from the detection plate 32 of the self-checkout device 10, displaying a message on the display screen 34 of the interaction unit 14, the message requesting the user to place the item on the weighing scale unit of the self-checkout device 10; - activating a weighing scale unit to measure the weight of the bulk product; and - calculating the price of the bulk product by multiplying the measured weight of the bulk product by the price per unit weight of the obtained bulk product.
[0275] Based on the method 200 described in the preceding paragraph, the following use cases are addressed by the self-checkout device 10 of the present disclosure: a) The customer places the product / loose item on the detection plate 32 where it is identified by barcode or appearance within 30 seconds. The self-checkout device 10 displays guidance to the customer to assist in positioning the product / loose item on the detection plate 32, increasing the likelihood that the product / loose item will be correctly identified. b) The customer uses the 1D barcode scanner of the self-checkout device 10 to scan a product that is not automatically identified by the software of the self-checkout device 10. c) The customer places the bulk product on the weighing scale unit of the self-checkout device 10 to measure the weight of the bulk product. The weighing scale unit may be integrated with the detection plate 32, in which case the customer removes all other products from the detection plate 32 before the weighing scale unit measures the weight of the bulk product. The software calculates the price of the bulk product by multiplying the weight of the bulk product by the price per unit weight. The price of the bulk product is added to the bill amount, which is calculated as the sum of the prices of the remaining identified products. d) If the customer has more than six products / loose items, the customer places the first six products / loose items on the detection plate 32. Once the products / loose items have been identified by the software of the self-checkout device 10, the customer presses the multi-function button on the self-checkout device 10 to allow the further products / loose items to be included in the total bill and places the further products / loose items on the detection plate 32, whereupon the further products / loose items are identified and their prices are added to the total bill. e) If the user is in trouble, they can press the multi-function button to put the self-checkout device 10 into standby mode and call a store attendant for help, who can be alerted by the illumination of colored lights mounted on either or both of the recessed mounting member and the interaction unit of the self-checkout device 10. f) If payment for the product / loose product detected by the detection plate 32 is not received at the self-checkout device 10, the self-checkout device 10 goes into standby mode and an attendant is alerted to come to the self-checkout device 10 in question and investigate the problem. g) Upon assessing why the self-checkout device 10 has gone into standby mode, the store attendant enters a PIN into the self-checkout device 10, thereby resetting the self-checkout device 10 for further use. h) A store associate can add more products / loose items for subsequent detection and identification by the software in the self-checkout device 10 by capturing a small number (up to three) of video frames of the SKU to train the object detection neural network and the embedded neural network. In an embodiment of a distributed network of self-checkout devices 10, the added products / loose items are automatically communicated to the remaining self-checkout devices 10 in the store and, if necessary, to self-checkout devices 10 in other stores.
[0276] 9, there is shown a diagram of a first alternative embodiment of the self-checkout device 10 of the present disclosure. As shown, according to the first alternative embodiment, the self-checkout device 10 may have a stand unit 12 attached to a side side (here, the left side) of the interaction unit 14, as opposed to a rear side of the self-checkout device 10 (as shown and described with reference to FIGS. 1-3).
[0277] The present disclosure provides a space-saving self-checkout device 10 capable of quickly registering one or more products, some of which may be bulk products. Specifically, the self-checkout device 10 includes a detection plate 32 on which multiple products can be placed together. The self-checkout device 10 also includes two cameras 26, 28 configured to capture video footage of the detection plate 32 and the products placed thereon. The video footage from the cameras 26, 28 is processed using computer vision to enable identification and registration of each product. The self-checkout device 10 implements a robust product recognition algorithm configured to identify and recognize products placed on the plate regardless of the product's orientation. Thus, the product recognition algorithm allows for product identification without the need to scan the product's barcode. The self-checkout device 10 also provides a weighing scale unit for weighing weight-dependent bulk product items. This allows the self-checkout device 10 to register up to six products substantially simultaneously in a single transaction without the customer having to scan or manually identify the items. Thus, the present self-checkout device 10 significantly increases the speed at which multiple products can be registered in a single transaction, thereby reducing delays in high throughput sales environments.
[0278] The self-checkout device 10 can operate with a significantly fewer number of cameras than other solutions. Specifically, the self-checkout device 10 can operate with two cameras 26, 28 mounted on an upright, recessed mounting member 16. In contrast, other solutions require four to six cameras. In one example, the two cameras 26, 28 are positioned such that the first camera 26 is positioned approximately above the detection plate 32, looking down on the detection plate 32, and the second camera 28 views the side of the detection plate 32. The inward curvature of the recessed mounting member 16 provides a wider field of view for the second camera 28. The inward curvature of the recessed mounting member 16 also makes the detection plate 32 more accessible and easier to locate than would be possible with a vertical, upright mounting member. The recessed mounting member 16 also includes a light-diffusing casing member 24 (or reflective element) that further illuminates the products on the detection plate 32, thereby assisting the computer vision algorithms of the present disclosure. Here, the detection plate 32 includes an upward-facing interactive display unit configured to display markings that guide the placement of products on the detection plate 32 for subsequent identification by a computer vision algorithm. The display unit is configured to display a white background to eliminate reflections of products placed on the detection plate 32 and further eliminate reflections of the surrounding environment on the display unit. Here, the placement of the displayed markings is determined by a placement algorithm operable using video footage from the first camera 26. The placement algorithm determines an optimal placement of one or more products on the detection plate 32, and the optimal placement is established to maximize the view of the products by the cameras 26, 28 while minimizing occlusion of the products by other products placed on the detection plate 32.
[0279] The foregoing description of specific embodiments of the present disclosure has been presented for purposes of illustration and description. The terms "including," "comprising," "incorporating," "having," "being," and the like, used to describe and claim the present disclosure, are intended to be interpreted in a non-exclusive manner, i.e., there may be items, components, or elements not expressly described. References to the singular should also be construed as relating to the plural. The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." An embodiment described as "exemplary" should not necessarily be construed as preferred or advantageous over other embodiments and / or exclude the incorporation of features of other embodiments. They are not intended to be exhaustive or to limit the disclosure to the precise forms disclosed, and obviously, many modifications and variations are possible in light of the above teachings. The exemplary embodiments have been chosen and described in order to best explain the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the present disclosure and various embodiments with various modifications suited to the particular uses envisioned.
Claims
1. a detection plate adapted to allow placement of the product; one or more cameras positioned to have a field of view surrounding at least the detection plate, the one or more cameras configured to provide video footage; a motion detection module configured to detect the presence of motion in the video footage; a sequence selection module configured to select a sequence of video frames over a time interval corresponding to detecting the presence of motion in the video footage; an appearance interpretation module configured to register one or more products present within the sequence of video frames; a billing module configured to retrieve prices of the one or more registered products, generate a total bill based on the retrieved prices, and process payment of the total bill; a controller module operatively connected to the one or more cameras and communicatively coupled to the motion detection module, the sequence selection module, the appearance interpretation module, and the billing module to control operation thereof and facilitate communication therebetween; A self-checkout device comprising:
2. The appearance interpretation module: an object detection module configured to analyze the sequence of video frames to detect one or more objects therein; a cropping module configured to isolate the one or more objects detected in the sequence of video frames and extract visual features of the one or more detected objects; an embedding module configured to convert the extracted visual features of the one or more detected objects into an embedded feature vector; an expert system module configured to compare the embedded feature vector with feature vectors pre-stored in an embedded database and to identify the one or more detected objects based on the comparison; Equipped with the identified one or more objects are registered as the one or more products; The self-checkout device of claim 1 .
3. the appearance interpretation module uses machine learning models to facilitate the detection, cropping, embedding, and identification processes; The self-checkout device of claim 2 .
4. the expert system module is further configured to determine whether any one of the one or more objects identified from the one or more products is a weight-dependent bulk product item. The self-checkout device of claim 2 .
5. a weighing module configured to activate a weighing scale unit to measure a weight of a weight-dependent bulk product item from the one or more products placed on the detection plate, and the billing module configured to generate a total bill based on the measured weight of the weight-dependent bulk product item. The self-checkout device of claim 4.
6. a barcode processing module configured to detect one or more barcodes in the selected sequence of video frames and decode the detected barcodes corresponding to the one or more registered products, wherein the billing module is configured to obtain prices of the one or more registered products based on the decoded barcodes. The self-checkout device of claim 1 .
7. and a guidance module operatively connected to the design display unit, the guidance module configured to activate the design display unit to display a design on the detection plate and provide a user with visual guidance for optimal placement of a product on the detection plate. The self-checkout device of claim 1 .
8. a recessed mounting member disposed upright relative to the detection plate, the one or more cameras being mounted to the recessed mounting member; The self-checkout device of claim 1 .
9. the recessed mounting member houses an illumination device for illuminating the detection plate; The self-checkout device of claim 8.
10. the one or more cameras comprising a first camera and a second camera oriented at different angles to capture the video footage of the product from multiple perspectives; The self-checkout device of claim 1 .
11. the billing module is further configured to generate an itemized list based on the registered product or products. The self-checkout device of claim 1 .
12. a display screen configured to display the itemized list and the total bill amount; The self-checkout device of claim 11.
13. further comprising a management module configured to support updates to the configuration of the self-checkout device, including a product database. The self-checkout device of claim 1 .
14. the self-checkout device operates as a standalone device; The self-checkout device of claim 1 .
15. 1. A method implemented by a self-checkout device, the method comprising: receiving video footage of a detection plate of the self-checkout device from one or more cameras; detecting the presence of motion in the video footage by processing the video footage; selecting a sequence of video frames over a time interval corresponding to detecting the presence of motion in the video footage; Detecting and decoding one or more barcodes displayed within the sequence of video frames; calculating a total bill amount corresponding to the one or more decoded barcodes; displaying the total charge on a display screen of the self-checkout device; A method for providing
16. detecting an item displayed in said sequence of video frames when one or more barcodes are not displayed; Distinguishing the detected items between sale items and non-sale items; issuing a first alert upon detection of one or more non-sale items, the first alert including a message to remove the non-sale items placed on the detection plate of the self-checkout device; The method of claim 15 further comprising:
17. determining a distribution of detected sale items on the detection plate of the self-checkout device; issuing a second alert if the determined distribution of the detected sale items is not appropriate; and 17. The method of claim 16, further comprising:
18. cropping one or more regions from each of the sequence of video frames substantially surrounding each detected item for sale; generating an embedded representation from each of the one or more cropped regions of the sales item displayed therein; comparing the generated embeddings with records of product embeddings to find records that match the product embeddings; determining a price corresponding to a record that matches the embedded representation of the product; calculating a total bill amount as the sum of determined prices corresponding to records that match the embedded representation of the product for all of the detected sale items; displaying the total bill amount on the display screen; 20. The method of claim 17, further comprising:
19. 20. The method of claim 18, further comprising receiving payment of the total bill amount.
20. 10. A computer program product having machine-readable instructions stored thereon, the machine-readable instructions, when executed by one or more processing units, causing the one or more processing units to perform the method of claim 1. Computer program products.
Citation Information
Patent Citations
Self-service computer system and method for using the same
JP2010182306A
Store system and program
JP2013050924A
Product sales data processing device
JP2016014937A
Adjustment and settlement device and automated store system
JP2020166639A
Information processing device, program and maintenance terminal
JP2021101300A