Customer support in self-checkout using computer vision

A computer vision and machine learning system identifies checkout problems and provides real-time instructions to customers, enhancing transaction accuracy and reducing computational resources.

JP2025175949APending Publication Date: 2025-12-03TOSHIBA TEC KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025041708
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-17
Filing Date
2025-03-14
Publication Date
2025-12-03

AI Technical Summary

Technical Problem

Customers face difficulties in completing transactions at self-checkout lanes due to issues like overlapping items, occlusions, and other factors that affect computer vision accuracy, leading to inaccurate item recognition and potential errors.

Method used

Implementing a system that uses computer vision and machine learning to identify potential checkout problems and provide real-time instructions to customers, such as moving items, to improve visibility and accuracy.

Benefits of technology

Enhances transaction success by reducing computational resources and avoiding errors, improving customer experience and merchant efficiency by addressing visibility issues proactively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025175949000001_ABST
    Figure 2025175949000001_ABST
Patent Text Reader

Abstract

To provide a technology to use machine learning (ML) having a point-of-sales (POS) system.SOLUTION: A technology includes identifying one or more images of one or more items to be purchased, which are captured in a point-of-sales (POS) system. The technology further includes predicting a checkout problem relating to the one or more items to be purchased, which includes using a trained ML model to determine the checkout problem on the basis of the one or more images. The technology further includes generating one or more instructions for a purchaser of the one or more items on the basis of the predicted checkout problem, and presenting the one or more instructions in a user interface of the POS system.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001]

[0001] The present disclosure relates to machine learning (ML), including computer vision. Automation with computer vision continues to increase in checkout lanes, and sometimes customers have difficulty following instructions to successfully complete a transaction (e.g., to successfully complete a transaction that uses computer vision to identify items for purchase in a checkout lane). For example, a customer's interaction with a lane or device determines the accuracy of the checkout experience (e.g., computer vision accuracy, output, predictions, or alerts). [Brief explanation of the drawings]

[0002] [Figure 1] FIG. 1 illustrates an exemplary checkout area with customer assistance using computer vision item detection, according to one embodiment. [Figure 2]

[0003] FIG. 2 is a block diagram illustrating a controller for customer assistance using computer vision item detection, according to one embodiment. [Figure 3]

[0004] FIG. 3 is a flowchart illustrating customer assistance using computer vision item detection, according to one embodiment. [Figure 4]

[0005] FIG. 4 is a flowchart illustrating predicting checkout problems using computer vision, according to one embodiment. [Figure 5]

[0006] FIG. 5 is a flowchart illustrating training an ML model for predicting checkout problems using computer vision, according to one embodiment. [Figure 6]

[0007] FIG. 6 is a flowchart illustrating inference using ML models to predict checkout problems using computer vision, according to one embodiment. [Figure 7]

[0008] FIG. 7 is a flowchart illustrating generating instructions for checkout assistance, according to one embodiment. Detailed Description

[0003]

[0009] As can be seen, customer behavior during checkout (e.g., how items are positioned in a point-of-sale (POS) environment) can affect the accuracy of computer vision analysis and the success of the checkout transaction. In embodiments, automated aspects of the checkout experience can be enhanced by providing relevant instructions (e.g., audio, video, or text messages) to the customer while they are using the automated checkout system. For example, the computer vision-based item detection process can be extended to include generating instructions for the customer. The computer vision component, or another suitable ML model, can consider attributes of the items being purchased (e.g., their placement, shape, size, boundaries, and any other suitable attributes) and use these attributes to identify potential problems with the checkout system. The system can then provide instructions to the customer (e.g., instructions regarding moving the items) to enhance the transaction. To provide accurate predictions, this can include having the computer vision system indicate how a sensor (e.g., a camera or other image capture device) should display the items to be purchased.

[0004]

[0010] In embodiments, when a POS system identifies a potential checkout problem, the customer may be provided with an immediate (or near-immediate) on-screen alert, allowing the customer to address the issue. For example, assume a customer places three items for checkout at a POS system (e.g., on the POS system's scale). Two of the items overlap without a clear boundary (e.g., to a computer vision system), blocking the third item from view. In existing systems, the computer vision system may fail to identify the third item (e.g., because it blocks the view to an image capture device or other sensor).

[0005]

[0011] In embodiments, the POS system can identify this potential problem (e.g., predict the problem using an appropriate ML model) and provide instructions to the customer. For example, the POS system can provide text, audio, or video instructions to the customer to instruct the customer to move the item so that sufficient clearance is presented to view the item. This can greatly improve the checkout experience for the customer (e.g., to avoid errors) and for the merchant (e.g., to avoid employee intervention or loss). As described below, in embodiments, the POS system uses computer vision to predict checkout problems and provide instructions to the customer. Alternatively, or additionally, the POS system uses additional sensor data (e.g., weight data captured by a scale) in an appropriate ML model to predict checkout problems.

[0006] The benefits of using computer vision for customer assistance

[0012] As described above, in embodiments, ML (e.g., a computer vision ML model and one or more additional suitable ML models) can be used to predict checkout problems and provide instructions to the customer. This has many technical advantages. For example, customer use of a POS system can cause computer vision techniques to be inaccurate or inadequate at accurately predicting items for purchase. As described further below with respect to FIG. 1 , an item may be blocked because it is not sufficiently visible to the image capture device (e.g., the item is overlapping, occluded, blocked, too far or too close to the image capture device, or some other suitable issue), the item is crumpled, rotten, broken, or damaged, or for some other suitable reason. Providing instructions to the customer to address these issues can result in significantly more accurate predictions. This solves a technical problem inherent to computer vision techniques—computer vision predictions are limited to visible data—by using computer vision (or other ML techniques) to predict problems and provide instructions to resolve the problems.

[0007]

[0013] Additionally, one or more of the techniques described below can reduce the computational resources used for prediction by reducing the resources used for computer vision item recognition. For example, fewer training resources can be used to train a computer vision ML model because the computer vision ML model uses limited data and the POS system is not burdened with identifying incorrectly placed items. Instead, the POS system can provide instructions to the customer to address the checkout problem, allowing the computer vision system to access more complete data to more easily identify the item. This also avoids wasted computation (including wasted power) by avoiding repeated, unsuccessful attempts to identify the item. Instead, the POS system can use an appropriate ML model to predict the problem and provide instructions, enabling a successful transaction while reducing wasted attempts at item recognition.

[0008]

[0014] 1 illustrates an exemplary checkout area 100 with customer assistance using computer vision item detection, according to one embodiment. In an embodiment, checkout area 100 relates to a retail environment (e.g., a grocery store or any other suitable retail environment). This is merely an example, and checkout area 100 may relate to any suitable environment.

[0009]

[0015] One or more purchasers 102 use checkout area 110 (e.g., to pay for purchases). In an embodiment, checkout area 110 includes multiple point-of-sale (POS) systems 120A-N. For example, one of the purchasers 102 can use one of POS systems 120A-N for self-checkout to purchase an item. Checkout area 110 also includes an employee station 126. For example, an employee (e.g., a retail store employee) can use employee station 126 to monitor the purchasers 102 and POS systems 120A-N. Self-checkout is merely one example, and POS systems 120A-N can be any suitable systems. For example, POS system 120A can be an assistance checkout kiosk where an employee assists purchasers with checkout, or checkout area 110 can be entirely purchaser-centric and not include employee assistance.

[0010]

[0016] In embodiments, each of POS systems 120A-N includes components used by a purchaser for self-checkout. For example, POS system 120A includes a scanner 122 and one or more sensors 124A-N (e.g., an image capture device, a weight sensor, a pressure sensor, or any other suitable sensor). In embodiments, purchaser 102 can use scanner 122 to scan a Universal Product Code (UPC) on an item. Furthermore, in embodiments, scanner 122 can be integrated with or replaced by one or more of sensors 124A-N. For example, sensors 124A-N can include one or more image capture devices (e.g., a visible spectrum camera, an infrared camera, or any other suitable image capture device), which can be used to identify an item (e.g., based on the UPC for the item or some other suitable aspect of the item).

[0011]

[0017] In embodiments, POS system 120A can communicate with management system 140 using network 130. Network 130 can be any suitable communication network, including a local area network (LAN), a wide area network (WAN), a cellular communication network, the Internet, or any other suitable communication network. POS system 122A can communicate with network 130 using any suitable network connection, including a wired connection (e.g., an Ethernet connection), a Wi-Fi connection (e.g., an 802.11 connection), or a cellular connection.

[0012]

[0018] In embodiments, POS system 120A can communicate with management system 140 to identify items scanned by a purchaser and perform other functions related to self-checkout. Management system 140 is described in further detail below in connection with FIG. 2. In embodiments, POS system 120A can use management system 140 to identify items (using scanner 122, sensors 124A-N, or any combination thereof).

[0013]

[0019] For example, POS system 120A can capture information using sensors 124A-N and can use computer vision to identify and predict problems with items and to provide instructions (e.g., to purchaser 102, an employee at employee station 126, or some other suitable human entity or automated system), as will be further described below with respect to FIG.

[0014]

[0020] These issues may include a variety of issues. For example, in embodiments, POS system 120A can predict issues affecting computer vision and item recognition. This may include overlapping items (e.g., multiple items are shown for purchase and one or more items overlap one or more other items, obscuring the sensor view for item prediction), occlusions (e.g., the sensor view is blocked or obstructed), lighting conditions (e.g., insufficient light, excessive glare, or other suitable lighting conditions), distance (e.g., an item that is too far or too close to the sensor), sensor issues (e.g., camera malfunction, incorrect camera orientation, or other suitable sensor issue), or any other suitable issue affecting computer vision and item recognition.

[0015]

[0021] As another example, in embodiments, the POS system 120 can predict problems that will affect a customer's purchasing or shopping experience. This may include crumpling or tearing of a product (e.g., a product bag), a spoiled or malfunctioning item (e.g., a bruised or spoiled product), a broken product, a product requiring employee intervention (e.g., a product requiring age verification), or any other suitable problem that will affect a customer's purchasing or shopping experience. As another example, in embodiments, the POS system 120 can predict problems that will affect a purchase transaction. This may include identifying an item that is lodged in a customer's arm, hand, or shopping container (e.g., a shopping bag, basket, cart, or other suitable shopping container). These are merely examples, and the POS system 120A can predict a wide variety of problems.

[0016]

[0022] 1 illustrates a management system 140 connected to the checkout area 110 using a communications network 130. The management system 140 can respond to the POS system 120A with predicted instructions (generated using an appropriate ML model) to address the checkout issue. The POS system 120A can then present the instructions to the purchaser (using text, audio instructions, video instructions, or some other appropriate interface). This is merely an example, and the management system 140 can be maintained, in whole or in part, on a local computer accessible to the POS 120A, without using a network connection (e.g., maintained on the POS system 120A itself or in a local storage repository).

[0017]

[0023] Additionally, in embodiments, sensors 124A-N are also components of POS system 120A and can be used to predict problems (e.g., using computer vision) and predict instructions to address the problems. For example, sensors 124A-N can include image capture devices used to capture one or more images of an item that purchaser 102 intends to purchase. POS system 120A can transmit the images to management system 140 to identify the item depicted in the image. Management system 140 can then use one or more suitable trained ML models (e.g., computer vision ML models and any other suitable ML models) to identify the item depicted in the image, predict the problem, predict instructions for the problem, and respond with the instructions to POS system 120A.

[0018]

[0024] For example, management system 140 can send instructions (e.g., audio, text, or video instructions) directly to POS system 120A, a code identifying a predefined instruction (from a corpus of available instructions), or any other suitable information. POS system 120A can use the received information to present the instructions to the purchaser. In an embodiment, the instructions are maintained at POS system 120A. Alternatively, this information can be maintained in another suitable location. For example, POS system 120A can communicate with any suitable storage location (e.g., a local storage location or a cloud storage location) to retrieve the information (e.g., using the identification code for the instruction). Alternatively or additionally, as described above, management system 140 can provide the information (e.g., instructions) to the user.

[0019]

[0025] 2 is a block diagram illustrating a controller 200 for customer assistance using computer vision item detection, according to one embodiment. In an embodiment, the controller 200 is used for the management system 140 illustrated in FIG. 1. The controller 200 includes a processor 202, a memory 210, and a network component 220. The processor 202 retrieves and executes programming instructions stored in the memory 210. The processor 202 may represent a single central processing unit (CPU), multiple CPUs, a single CPU with multiple processing cores, a graphics processing unit (GPU) with multiple execution paths, etc.

[0020]

[0026] Network component 220 includes components for controller 200 that interface with an appropriate communication network (e.g., communication network 130 shown in FIG. 1). For example, network component 220 may include wired, Wifi, or cellular network interface components and associated software. Although memory 210 is shown as a single entity, memory 210 may include one or more memory devices having blocks of memory associated with physical addresses, such as random access memory (RAM), read-only memory (ROM), flash memory, or other types of volatile and / or non-volatile memory.

[0021]

[0027] Memory 210 generally includes program code for performing various functions related to the use of controller 200. Although alternative implementations may have different functions and / or combinations of functions, the program code is generally described as various functional "applications" or "modules" within memory 210. Within memory 210, checkout issues service 212 uses computer vision item detection to facilitate customer assistance, as will be further described below with respect to Figures 3-7.

[0022]

[0028] 2 depicts checkout problem service 212 as located in memory 210, that representation is provided merely as an illustration for clarity. More generally, controller 200 may include one or more computing platforms, such as, for example, computer servers, which may be co-located, separate, or may form an interactively linked but distributed system, such as a cloud-based system (e.g., a public cloud, a private cloud, a hybrid cloud, or any other suitable cloud-based system). Consequently, processor 202 and memory 210 may correspond to memory resources in a distributed processing and computing environment. Furthermore, in embodiments, checkout problem service 212 may be divided across any suitable number of computer systems or computing nodes (e.g., in a cloud computing system), including being fully or partially integrated within a POS device.

[0023]

[0029] 3 is a flowchart 300 illustrating customer assistance using computer vision item detection, according to one embodiment. In block 302, a checkout issue service (e.g., the checkout issue service 212 illustrated in FIG. 2) captures the checkout environment. For example, the checkout issue service (or some other suitable software service) can capture the checkout environment (e.g., the POS system 120A illustrated in FIG. 1) using one or more sensors (e.g., the sensors 124A-N illustrated in FIG. 1).

[0024]

[0030] In block 304, the checkout problem service predicts a checkout problem using computer vision. In an embodiment, the checkout problem service may use one or more ML models to attempt to identify items using computer vision and predict a problem based on the output from the attempt. This is described further below with respect to FIG. 4. For example, the checkout problem service may use a trained computer vision ML model (e.g., a neural network) to attempt to identify items in the checkout environment based on the data captured in block 302. The checkout problem service may use another ML model (e.g., another neural network) to predict a problem based on the output from the first ML model (e.g., a computer vision ML model).

[0025]

[0031] As described above, in embodiments, the checkout issue service predicts checkout issues using computer vision (e.g., based on captured images or video of the checkout area). Alternatively, or in addition, the checkout issue service uses additional sensor data to predict checkout issues. For example, a POS system (e.g., POS system 120A illustrated in FIG. 1) may include a scale or other non-visual sensors. The checkout issue service may use data from one or more of these sensors to predict checkout issues (e.g., in addition to or instead of visual data). For example, a scale may be used to identify potentially hidden items not visible in the captured image or video, and the checkout issue service may use this weight discrepancy data to identify potential checkout issues. This is merely an example, and any suitable sensor data may be used.

[0026]

[0032] In block 306, the checkout problem service generates instructions for checkout assistance. For example, the checkout problem service can use the predicted problems generated in block 304 to generate or select appropriate instructions. These can be text instructions, audio instructions, video instructions, or any other suitable instructions. This will be described further below with respect to FIG. 7.

[0027]

[0033] In block 308, the checkout problem service presents instructions. For example, the checkout problem service may provide audio, text, or video instructions through a suitable user interface (e.g., a screen, a speaker, or any other suitable user interface). This is by way of example only, and the checkout problem service (or any suitable software service) may present any suitable type of instructions.

[0028]

[0034] Figure 4 is a flowchart illustrating predicting checkout problems using computer vision, according to one embodiment. In an embodiment, Figure 4 corresponds to block 304 illustrated in Figure 3. In block 402, a checkout problem service (e.g., checkout problem service 212 illustrated in Figure 2) attempts to identify an item using computer vision.

[0029]

[0035] For example, the checkout problem service may use a suitable computer vision ML model (e.g., a deep neural network (DNN), a support vector machine (SVM), or some other suitable ML model) to attempt to infer items for purchase from one or more images captured by the sensor. In embodiments where the POS kiosk includes multiple image capture devices, multiple images (e.g., captured from different angles) may be used together to predict items for purchase.

[0030]

[0036] In block 404, the checkout problem service predicts problems based on the computer vision output. For example, in block 402, the checkout problem service (e.g., using a computer vision ML model or another suitable ML model) can output a predicted item (e.g., a best guess at the predicted item) along with a confidence score or other indication of the likelihood of the accuracy of the prediction. As another example, the checkout problem service can output the identified problems. For example, in addition to (or instead of) item recognition, a computer vision ML model can be trained to recognize potential checkout problems in a POS environment. The checkout problem service can use the computer vision ML model to output these potential problems in addition to or instead of the predicted items. The checkout problem service can then use an appropriate ML model (e.g., the trained ML model) to predict problems based on this output. This will be described further below with respect to Figures 5-6.

[0031]

[0037] For example, as described above with respect to FIG. 1 , the checkout issue service can predict a wide range of issues affecting computer vision and item recognition. This can include overlapping items, occlusions, lighting conditions, distance from the image capture device, sensor issues, or any other suitable issues affecting computer vision and item recognition. As another example, the checkout issue service can predict issues affecting a customer's purchasing or shopping experience. This can include crumpled or torn products, rotten or malfunctioning products, broken products, products requiring employee intervention, items lodged in a customer's arm, hand, or shopping container, or any other suitable issues affecting a customer's purchasing or shopping experience. These are merely examples. Furthermore, as described above with respect to block 304 illustrated in FIG. 3 , computer vision is merely an example, and the checkout issue service can supplement or replace the captured image data using additional sensor data (e.g., weight data from a scale).

[0032]

[0038] FIG. 5 is a flowchart 500 illustrating training an ML model for predicting checkout problems using computer vision, according to one embodiment. This is merely an example, and in embodiments, any suitable unsupervised technique (without requiring training) can be used. In block 502, a training service (e.g., a human administrator or a software or hardware service) collects historical checkout data. For example, a checkout problem service (e.g., checkout problem service 212, as illustrated in FIG. 2) can be configured to function as a training service and collect historical data reflecting customer checkouts. In embodiments, the historical data includes predicted items and confidence scores (output by a computer vision ML model, as described above with respect to FIG. 4) along with identified problems corresponding to the predicted items and confidence scores. Alternatively or additionally, the historical data includes labeled image data reflecting checkout problems. These are merely examples, and any suitable historical checkout data or other training data can be used.

[0033]

[0039] In block 506, a training service (or other suitable service) preprocesses the collected historical checkout data. For example, the training service can create feature vectors that reflect the values ​​of various features for the historical checkout data. In block 508, the training service receives the feature vectors and uses them to train a trained checkout issue prediction ML model 510.

[0034]

[0040] In an embodiment, at block 504, the training service also collects additional data. For example, the training service may use UPC data, weight data, or any other suitable data to further identify products. At block 506, the training service may also preprocess this additional data. For example, feature vectors corresponding to the historical checkout data may be further annotated using the additional data. Alternatively or additionally, additional feature vectors corresponding to the additional product data may be created. At block 508, the training service uses the additional data preprocessed during training to generate a trained checkout issue prediction ML model 510.

[0035]

[0041] In embodiments, preprocessing and training can be performed as batch training. In this embodiment, data is preprocessed immediately (historical checkout data and additional data) and provided to the training service in block 508. Alternatively, preprocessing and training can be performed in a streaming manner. In this embodiment, data is streaming and continuously preprocessed and provided to the training service. For example, it may be desirable to take a streaming approach for scalability. The training data set may be very large, and therefore it may be desirable to preprocess the data and provide it to the training service in a streaming manner (e.g., to avoid computational and storage limitations). Additionally, in embodiments, a federated learning approach can be used in which multiple entities contribute to a training shared model.

[0036]

[0042] 6 is a flowchart 600 illustrating estimation using an ML model to predict checkout problems using computer vision, according to one embodiment. In an embodiment, a processing service 620 (the checkout problem service 212 shown in FIG. 2 or any other suitable software service) is associated with a checkout problem prediction ML model 510. In an embodiment, the checkout problem prediction ML model 510 is trained to infer predicted problems 630 for one or more identified items or problems 602. For example, the checkout problem prediction ML model 510 can predict problems based on identified items, confidence scores for the items, etc. Alternatively or additionally, the checkout problem prediction ML model 510 predicts problems based on identified problems (e.g., problems identified directly from images captured by a computer vision ML model).

[0037]

[0043] In an embodiment, the predicted problem 630 reflects a predicted problem for the identified product or problem 602. Alternatively or additionally, the predicted problem 630 identifies multiple suggested matches (e.g., a range of predicted problems). In an embodiment, the checkout problem prediction ML model 510 provides a confidence score for the likely predicted problem 630, which may take into account the instructions generated in block 306 illustrated in FIG. 3.

[0038]

[0044] 7 is a flowchart illustrating generating instructions for checkout assistance, according to one embodiment. In an embodiment, FIG. 7 corresponds to block 306 illustrated in FIG. 3. At block 702, a checkout problem service (e.g., checkout problem service 212 illustrated in FIG. 2) receives a predicted problem. For example, the checkout problem service can receive the predicted problem generated by one or more ML models, as described above with respect to FIGS. 4-6.

[0039]

[0045] In an embodiment, the checkout problem service receives one predicted problem (e.g., the most likely predicted problem). Alternatively, the checkout problem service receives multiple predicted problems. For example, the checkout problem service can receive a list of predicted problems, which can be ranked (e.g., based on likelihood, expected impact, or some other suitable criteria). Additionally, the checkout problem service can receive a score associated with the predicted problem (e.g., a confidence score generated by an ML model). The score can indicate the likelihood of the problem occurring, the severity of the problem, or some other suitable information.

[0040]

[0046] At block 704, the checkout problem service generates instructions based on the predicted problem. In an embodiment, the checkout problem service uses the predicted problem (or problems) to generate instructions. For example, the checkout problem service can use a suitable natural language processing (NLP) ML model to generate text or audio instructions. As another example, the checkout problem service can use a suitable ML model to generate visual instructions (e.g., an image or video containing the instructions). In this example, an ML model can be trained to generate instructions from the predicted problem, and the checkout problem service can use the ML model to infer the instructions.

[0041]

[0047] Alternatively, or additionally, the checkout service can use the predicted problem to select from previously generated instructions. For example, one or more instructions can be generated for various possible problems. The checkout service can use the predicted problem to select from these instructions.

[0042]

[0048] In embodiments, the checkout issue service can provide various instructions using various media (e.g., text instructions, audio instructions, video instructions, or any other suitable instruction format). For example, the checkout service can instruct the customer to move an item (e.g., "move the 24-pack of soda toward the side of the checkout area") to improve the data available to the computer vision ML model. These instructions can identify the item (e.g., an item recognized using the computer vision ML model) and provide the customer with relevant information (e.g., where or how to take action) in real time (or near real time) to improve the transaction.

[0043]

[0049] In embodiments, multiple ML models are used to predict checkout problems and present instructions to the user. For example, as described above with respect to FIG. 3, three models can be used. A computer vision ML model can be trained for item recognition and used to identify items in a checkout environment. A checkout problem prediction ML model (e.g., checkout problem prediction ML model 510 illustrated in FIGS. 5-6) can predict checkout problems (based on output from the computer vision ML model). Furthermore, an NLP ML model or other suitable ML model can be used to generate instructions from the predicted checkout problems. However, this is merely an example. In embodiments, fewer or more ML models can be used. For example, one ML model can be used to infer predicted checkout problems from captured environmental sensor data (e.g., images or video captured by a camera at the checkout). That is, one ML model receives the captured data as input and infers the predicted problem, rather than two separate ML models. Furthermore, this ML model can also be used to generate instructions. Alternatively, one or more instructions can be selected using rule-based techniques (e.g., selecting from pre-generated instructions) without using ML to generate the instructions. These are just examples.

[0044]

[0050] While various embodiments have been described for purposes of illustration, they are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles, practical applications, or technical improvements of the embodiments in the marketplace, or to enable those skilled in the art to understand the various embodiments disclosed herein.

[0045]

[0051] In the foregoing, reference has been made to the embodiments presented in this disclosure. However, the scope of the present disclosure is not limited to the described embodiments. Instead, any combination of the above-described features and elements, whether associated with different embodiments, is contemplated to implement and perform the contemplated embodiments. Moreover, the embodiments disclosed herein achieve advantages over other possible solutions or prior art, and whether or not an advantage is achieved by a given embodiment does not limit the scope of the present disclosure. Accordingly, the above-described aspects, features, embodiments, and advantages are merely illustrative and are not considered elements or limitations of the appended claims unless expressly recited in the claims. Similarly, references to "the present disclosure" should not be construed as a generalization of any inventive subject matter disclosed herein, nor should they be considered elements or limitations of the appended claims unless expressly recited in the claims.

[0046]

[0052] Aspects of the described embodiments may take the form of entirely hardware embodiments, entirely software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and hardware aspects, which may all be generally referred to herein as "circuits," "modules," or "systems."

[0047]

[0053] One or more of the described embodiments may be a system, a method, and / or a computer program product, which may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the embodiment.

[0048]

[0054] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanical coding devices such as punch cards or ridge structures in grooves having instructions recorded thereon, and any suitable combination of the foregoing. Computer-readable storage medium, as used herein, should not be construed as being a transitory signal itself, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through a fiber optic cable), or an electrical signal transmitted over a wire.

[0049]

[0055] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may comprise copper transmission cables, optical transmission fiber, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions to storage in a computer-readable storage medium within the respective computing / processing device.

[0050]

[0056] The computer-readable program instructions for carrying out the operations of the described embodiments may be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, or the like, and traditional procedural programming languages ​​such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or a connection to an external computer may be made (e.g., through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry to perform aspects of the described embodiments.

[0051]

[0057] Aspects of the described embodiments are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to the embodiments. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0052]

[0058] These computer-readable program instructions may be provided to a general-purpose computer, special-purpose computer processor, or other programmable data processing apparatus to produce a machine, whereby the instructions, executing via the computer processor or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram blocks. These computer-readable program instructions may also be stored in a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, and / or other device to function in the manner described, whereby the computer-readable storage medium having instructions stored therein comprises an article of manufacture containing instructions that implement aspects of the functions / acts specified in the flowchart and / or block diagram blocks.

[0053]

[0059] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to create a computer-implemented process, whereby the instructions executing on the computer, other programmable apparatus, or other device implement the functions / operations specified in the flowchart and / or block diagram blocks.

[0054]

[0060] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. In this regard, each block in the flowcharts or block diagrams may represent a portion, segment, or module of instructions, comprising one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may not occur in the order noted in the figures. For example, depending on the functionality involved, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may be executed in the reverse order. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.

[0055]

[0061] Embodiments may be provided to end users through a cloud computing infrastructure. Cloud computing generally refers to the provision of scalable computing resources as a service over a network. More formally, cloud computing may be defined as a computing capability that provides abstraction between computing resources and their underlying technical architecture (e.g., servers, storage, network), enabling convenient, on-demand network access to a shared pool of configurable computing resources that can be rapidly provisioned and released with minimal administrative effort or service provider interaction. Thus, cloud computing allows users to access virtual computing resources (e.g., storage, data, applications, and complete virtualized computing systems) in the "cloud" without regard to the underlying physical systems (or the location of these systems) used to provide the computing resources.

[0056]

[0062] Typically, cloud computing resources are provided to users on a pay-per-use basis, where users are charged for the computing resources they actually use (e.g., the amount of storage space consumed by the user or the number of virtual systems instantiated by the user). Users can access any of the resources present in the cloud anytime, anywhere from the Internet. In the context of the described embodiment, users can access applications (e.g., the checkout problem service 212 illustrated in FIG. 2) or related data available in the cloud. For example, the checkout problem service, or any aspect of the checkout problem service, can run on a computing system in the cloud to predict checkout problems and generate corresponding instructions. Furthermore, appropriate ML models and related data can be stored in a storage location in the cloud. Doing so allows users to access this information from any computing system attached to a network connected to the cloud (e.g., the Internet).

[0057]

[0063] While the forgoing subject matter is directed to one or more embodiments, other and further embodiments may be devised without departing from the basic scope thereof, which scope is determined by the following claims.

Claims

1. Identifying one or more images captured at a point-of-sale (POS) system of one or more items for purchase; predicting checkout problems related to the one or more items for purchase; the predicting includes determining the checkout problem using a trained machine learning (ML) model based on the one or more images; generating one or more instructions to a purchaser of the one or more items based on the predicted checkout problem; and presenting the one or more instructions in a user interface of the POS system.

2. determining the checkout problem using the trained ML model based on the one or more images, generating a first output in a computer vision ML model trained to recognize the item for purchase based on providing the one or more images to the computer vision ML model; and predicting the checkout problem by providing the first output to the trained ML model, wherein the trained ML model is different from the computer vision ML model.

3. generating one or more instructions to a purchaser of the one or more items based on the predicted checkout problem, 3. The method of claim 2, comprising determining the instructions using a natural language processing (NPL) ML model based on providing the predicted checkout problem to the NPL ML model.

4. The method of claim 3 , wherein the one or more instructions comprise instructions for improving accuracy in identifying the one or more items for purchase using the computer vision ML model.

5. determining the checkout problem using the trained ML model based on the one or more images, 2. The method of claim 1, comprising predicting the checkout problem by providing the one or more images to the trained ML model, wherein the trained ML model outputs the predicted checkout problem.

6. generating one or more instructions to a purchaser of the one or more items based on the predicted checkout problem, 6. The method of claim 5, comprising determining the instructions using a natural language processing (NLP) ML model based on providing the predicted checkout problem to the NLP ML model.

7. The method of claim 6 , wherein the one or more instructions comprise instructions for improving accuracy in identifying the one or more items for purchase using a computer vision ML model.

8. The method of claim 1 , wherein predicting the checkout problem for the one or more items for purchase is based on weight data captured using a scale in addition to the one or more images.

9. one or more non-transitory computer-readable media containing computer program code in any combination; The computer program code, when executed by any combination of one or more processors, Identifying one or more images captured at a point-of-sale (POS) system of one or more items for purchase; predicting checkout problems related to the one or more items for purchase; the predicting includes determining the checkout problem using a trained machine learning (ML) model based on the one or more images; generating one or more instructions to a purchaser of the one or more items based on the predicted checkout problem; and presenting the one or more instructions at a user interface of the POS system.

10. determining the checkout problem using the trained ML model based on the one or more images, generating a first output in a computer vision ML model trained to recognize the item for purchase based on providing the one or more images to the computer vision ML model; and predicting the checkout problem by providing the first output to the trained ML model, wherein the trained ML model is different from the computer vision ML model.

11. generating one or more instructions to a purchaser of the one or more items based on the predicted checkout problem, 11. The non-transitory computer program product of claim 10, comprising determining the instructions using a natural language processing (NLP) ML model based on providing the predicted checkout problem to the NLP ML model.

12. 12. The non-transitory computer program product of claim 11, wherein the one or more instructions comprise instructions for improving accuracy in identifying the one or more items for purchase using the computer vision ML model.

13. determining the checkout problem using the trained ML model based on the one or more images, 10. The non-transitory computer program product of claim 9, comprising predicting the checkout problem by providing the one or more images to the trained ML model, wherein the trained ML model outputs the predicted checkout problem.

14. generating one or more instructions to a purchaser of the one or more items based on the predicted checkout problem, determining the instructions using a natural language processing (NLP) ML model based on providing the predicted checkout problem to the NLP ML model; 14. The non-transitory computer program product of claim 13, wherein the one or more instructions comprise instructions for improving accuracy in identifying the one or more items for purchase using a computer vision ML model.

15. one or more processors; and one or more memories having stored thereon a program, the program, when executed by any combination of the one or more processors, performing operations, the operations including: Identifying one or more images captured at a point-of-sale (POS) system of one or more items for purchase; predicting checkout problems related to the one or more items for purchase; the predicting includes determining the checkout problem using a trained machine learning (ML) model based on the one or more images; generating one or more instructions to a purchaser of the one or more items based on the predicted checkout problem; and presenting the one or more instructions in a user interface of the POS system.

16. determining the checkout problem using the trained ML model based on the one or more images, generating a first output in a computer vision ML model trained to recognize the item for purchase based on providing the one or more images to the computer vision ML model; and predicting the checkout problem by providing the first output to the trained ML model, wherein the trained ML model is different from the computer vision ML model.

17. generating one or more instructions to a purchaser of the one or more items based on the predicted checkout problem, 17. The system of claim 16, comprising determining the instructions using a natural language processing (NLP) ML model based on providing the predicted checkout problem to the NLP ML model.

18. 20. The system of claim 17, wherein the one or more instructions comprise instructions for improving accuracy in identifying the one or more items for purchase using the computer vision ML model.

19. determining the checkout problem using the trained ML model based on the one or more images, 16. The system of claim 15, comprising predicting the checkout problem by providing the one or more images to the trained ML model, wherein the trained ML model outputs the predicted checkout problem.

20. generating one or more instructions to a purchaser of the one or more items based on the predicted checkout problem, determining the instructions using a natural language processing (NLP) ML model based on providing the predicted checkout problem to the NLP ML model; 20. The system of claim 19, wherein the one or more instructions comprise instructions for improving accuracy in identifying the one or more items for purchase using a computer vision ML model.