Barcode-aware object verification

Through machine learning-based classifiers and image processing technology, sub-images of the area of ​​interest of the item are generated, solving the problem of barcode fraud, achieving efficient item verification and preventing bill substitution.

CN120752642APending Publication Date: 2025-10-03ZEBRA TECHNOLOGIES CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480014726.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-27
Filing Date
2024-02-06
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively identifying and preventing barcode fraud, especially ticket swapping, which can result in high-value goods being incorrectly checked out.

Method used

A machine learning-based classifier, combined with object template data and image processing technology, generates sub-images corresponding to the area of ​​interest of the item, and trains the classifier model to identify the matching of barcodes and items, thereby preventing bill swapping.

Benefits of technology

It improves the accuracy of barcode verification and the ability to prevent bill substitution, ensuring the correctness and legality of commodity checkout.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120752642A_ABST
    Figure CN120752642A_ABST
Patent Text Reader

Abstract

A method comprising: capturing, by a scanner device comprising an image sensor, first image data representing at least a portion of a first item; decoding, by the scanner device, the first barcode presented in the first image data; determining a first article template associated with the first barcode, the first article template including first identifier data identifying the first article from other articles and first region of interest data specifying a first region of interest of the first article; generating second image data including a first region of interest of the first image data; determining, by a first machine learning model, that the second image data corresponds to first identifier data identifying the first item; and generating first data indicating that the first barcode matches the first item.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] A barcode represents data in a visual, machine-readable form. For example, a one-dimensional barcode represents data by varying the width and / or spacing of a series of parallel lines. Two-dimensional barcodes (sometimes called "matrix barcodes") are also used and, due to their two-dimensional structure, may have additional capabilities for encoding data relative to one-dimensional barcodes. A barcode scanner is a device that includes optical elements that can read or otherwise interpret a barcode. A barcode can be decoded using a scanner to produce a code that uniquely identifies the barcode (and / or an object associated with the barcode). Summary of the Invention

[0002] In various examples, a method for barcode-aware object authentication is generally described. In some examples, the method may include capturing, by a scanner device including an image sensor, first image data representing at least a portion of a first item. In some further examples, the method may include decoding, by the scanner device, a first barcode presented in the first image data. In still other examples, the method may include determining a first item template associated with the first barcode, wherein the first item template includes first identifier data that identifies the first item from other items. The first item template may also include first region of interest data that specifies a first region of interest of the first item. In some cases, the method may include generating second image data that includes the first region of interest of the first image data. In still other examples, the method may include determining, by a first machine learning model, that the second image data corresponds to first identifier data that identifies the first item. In some examples, the method may include generating first data indicating that the first barcode matches the first item.

[0003] In at least some further examples, the first item template may represent at least one of a contextual relationship or a geometric relationship between the first barcode and the first region of interest of the first item. In various examples, the first item template may include data indicating the barcode type of the first barcode. In various further cases, the first item template may include a template image of the first region of interest of the first item, while in other cases, the first item template may include the first region of interest data without the template image. In some examples where the first item template includes a template image, the first item template may also include at least one of coordinate data indicating the position of the first barcode in the template image, orientation data indicating the orientation of the first barcode in the template image, or size data indicating the size of the first barcode in the template image.

[0004] In various examples, barcode-aware object verification systems are generally described. In various examples, these systems may include an image sensor; at least one processor; and / or non-transitory computer-readable memory storing instructions. In various examples, the instructions, when executed by the at least one processor, may be effective to control the image sensor to capture first image data representing at least a portion of a first item. In various further examples, the instructions may be effective to decode a first barcode presented in the first image data. In various further examples, the instructions may be effective to determine a first item template associated with the first barcode. In various cases, the first item template may include first identifier data that identifies the first item from other items. In still other examples, the first item template may include first region of interest data that specifies a first region of interest of the first item. In at least some other examples, the instructions may be further effective to generate second image data that includes the first region of interest of the first image data. In at least some other examples, the instructions may be effective to use a first machine learning model to determine that the second image data corresponds to first identifier data that identifies the first item. In at least some examples, the instructions may be effective to generate first data indicating that the first barcode matches the first item.

[0005] In some other examples, other methods of barcode-aware object authentication may be described. In some examples, such other methods may include receiving first image data representing at least a portion of a first item. In various further examples, the other methods may include decoding a first barcode presented in the first image data. In some cases, the other methods may include determining a first item template associated with the first barcode. In various cases, the first item template may include first identifier data that identifies the first item from other items. In some examples, the first item template may include first region of interest data that specifies a first region of interest of the first item. In various cases, the other methods may also include generating second image data that includes the first region of interest of the first image data. In still other examples, the other methods may include determining, via a first machine learning model, that the second image data corresponds to first identifier data that identifies the first item. In various cases, the other methods may include generating first data indicating that the first barcode matches the first item. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The accompanying drawings, in which like reference numerals indicate identical or functionally similar elements throughout the separate views, together with the following detailed description, are incorporated herein and form a part of the specification and serve to further illustrate embodiments of the concepts comprising the claimed invention and to explain the various principles and advantages of those embodiments.

[0007] Figure 1is a diagram of a barcode-aware object authentication system in accordance with various aspects of the present disclosure.

[0008] Figure 2 Depicted are examples of object templates associated with barcodes according to various aspects of the present disclosure.

[0009] Figure 3 Depicted are example machine learning architectures that may be used to classify sub-images of objects in accordance with various aspects of the present disclosure.

[0010] Figures 4A to 4B Example image processing techniques that may be used to generate sub-images corresponding to regions of interest of an object in accordance with various aspects of the present disclosure are shown.

[0011] Figure 5 is a flow chart illustrating an example process 500 for barcode-aware object authentication in accordance with various aspects of the present disclosure.

[0012] Figures 6A to 6E Various examples of sub-image extraction based on regions of interest defined in object templates according to various aspects of the present disclosure are depicted.

[0013] Those skilled in the art will appreciate that the elements in the drawings are illustrated for simplicity and clarity and are not necessarily drawn to scale. For example, the sizes of some elements in the drawings may be exaggerated relative to other elements to help enhance understanding of the embodiments of the present invention.

[0014] Device and method configurations have been represented in the drawings by conventional symbols where appropriate, showing only those specific details that are relevant to an understanding of the embodiments of the invention so as not to obscure the disclosure with details that would be apparent to one of ordinary skill in the art having the benefit of the description herein. DETAILED DESCRIPTION

[0015] Barcodes can be used to quickly and accurately retrieve information about an object to which the barcode is attached (or to which the barcode is otherwise displayed or associated). For example, a barcode can be used in a point-of-sale system to determine the price of an object, thereby speeding up the checkout process. In other examples, barcodes can be used in inventory systems to provide information such as the quantity, category, price, location, etc. of an object. Typically, a barcode can be a visual code that can be scanned and decoded using a hardware and / or software-based barcode scanner (sometimes referred to as a barcode reader) to generate a code that identifies the barcode and, by extension, the object associated with the barcode. This code can be used as an index into a database that can store additional information about the object. The specific information that can be stored can vary depending on the desired implementation. For example, in a point-of-sale barcode system, a database entry associated with a barcode can include the price of the item (as well as other data).

[0016] Some barcode scanners (particularly those for one-dimensional barcodes) can use lasers (or other light) to scan barcodes. The decoded visual information can be used to find information associated with the barcode in a database. In some examples, the barcode scanner can use a camera including an image sensor (e.g., a complementary metal oxide semiconductor (CMOS) and / or a charge coupled device (CCD) image sensor) that can capture a frame of image data (e.g., a two-dimensional array including pixels, each pixel having a corresponding pixel value). The barcode presented in the image captured by such a camera-based barcode scanner can be detected and decoded using any desired object detection and / or barcode detection technology. Camera-based barcode scanners can effectively detect and decode two-dimensional matrix barcodes (such as two-dimensional (QR) codes). As with one-dimensional barcodes, the decoded barcode code (e.g., an alphanumeric string, a numeric code, etc.) can be used to find the corresponding entry for the object associated with the barcode.

[0017] A specific type of fraud, sometimes referred to as "bill switching," involves swapping a barcode on one item for one associated with a different item. For example, a high-value item such as a consumer electronics item may have a barcode label affixed to its packaging. A bad actor could obtain a barcode label from another, typically less expensive item and place that barcode over the barcode on the consumer electronics item. Consequently, during checkout, the decoded barcode would cause the bad actor to be charged the lower-priced item instead of the correct price for the consumer electronics item.

[0018] Various systems and techniques are described herein that can be used to verify that an object is correctly associated with a barcode scanned for the object. In various examples, a machine learning-based classifier can be deployed on a camera-based barcode reader (e.g., a barcode reader including an image sensor). Images of various objects containing their barcodes can be used as template images for object verification purposes. In addition to information about the barcode in the template for a given object (e.g., barcode type, identifier data (e.g., decoded code of the barcode), coordinates of the barcode in the template image, etc.), the object template data can also include information about the coordinates of a region of interest (ROI) in the template image. Thus, a close correlation between the size, position, and orientation of the object's ROI and the object's barcode can be established using the template data for a given object that is searched using a decoded barcode.

[0019] Object template data (also referred to as "item template data") may be stored in the memory of a scanner device and / or in the memory of another device configured to communicate with the scanner device (e.g., a point-of-sale computing device and / or a remote computing device). When the scanner device captures an image of an object including a barcode, the barcode may be detected and decoded, and the decoded barcode code may be used to look up the object template data in the memory. In some examples, the object template may include a template image of the object. Where the object template includes a template image, the object template may include information regarding the position, size, and / or orientation of the barcode in the template image (e.g., bounding box data for the barcode in the template image) and information regarding a region of interest of the item (e.g., a region of interest bounding box in the template image). For example, a first bounding box may identify the barcode in the template image, while a second bounding box in the template image may establish the region of interest of the first item. A sub-image corresponding to the region of interest may be extracted from the captured image of the object using the relationship between the barcode in the template image and the region of interest in the template image (and also by comparing the barcode in the captured image with the barcode in the template image). In some examples, geometric transformations can be performed based on the size, position, and / or orientation of the barcode in the captured image compared to the barcode in the template image. These transformations (sometimes referred to as "deskewing") can be used to provide a consistent sub-image that accurately reflects the region of interest of the item defined in the object template data, even though the captured image is closer, farther away, captures a different portion of the object, and / or is captured at a different orientation relative to the template image.

[0020] The extracted sub-image can be input into a classifier (e.g., a convolutional neural network, a model based on visual transformers, etc.), and the classifier can be trained to predict the object to which the input sub-image belongs. If the predicted object matches the object associated with the barcode, this indicates that the barcode and object are correctly associated with each other (e.g., no ticket swapping has occurred and the item can be sold). Conversely, if the predicted object does not match the object associated with the barcode, ticket swapping may have occurred. An error message or other output indicating the mismatch can be generated. Advantageously, providing a consistent region of interest (ROI) of the captured image, determined using the region of interest defined in the item template data, to the classifier network can result in more accurate performance of the classifier network. This is because captured images (e.g., images captured by a camera-based scanner) can capture different views and / or portions of an object. The views in the captured images can be at different ranges, and the object can be rotated at different angles. Furthermore, the various captured images can be under different lighting conditions. Such diverse views of different portions of the captured object can lead to less accurate classification by the image classifier network. Thus, the region of interest defined in the item template data can be used to extract relatively consistent sub-images from the captured image that conform to the region of interest defined in the item template. Since such images are more similar to images seen by the classifier network during training, the classifier network can better classify such images.

[0021] Machine learning techniques can be used to identify and / or classify objects in image data and / or generate encoded representations of inputs and / or generate predictions. In various examples, machine learning models can perform better than rule-based systems and can be more adaptable because they can be continuously improved over time by retraining the model as more data becomes available. Thus, machine learning techniques are often adaptable to changing conditions. Deep learning algorithms, such as neural networks, are often used to detect patterns in data and / or perform tasks.

[0022] Typically, in machine learning models such as neural networks, parameters control the activation of neurons (or nodes) within a layer of the machine learning model. The weighted sum of the activations of each neuron in the previous layer can be input to an activation function (e.g., a sigmoid function, a rectified linear unit (ReLu) function, etc.). The result determines the activation of the neurons in the next layer. In addition, a bias value can be used to shift the output of the activation function to the left or right on the x-axis, thereby biasing the neuron towards activation.

[0023] Typically, in a machine learning model such as a neural network, after initialization, annotated training data can be used to generate a cost or "loss" function that describes the difference between the expected output of the machine learning model and the actual output. The parameters of the machine learning model (e.g., weights and / or biases) can be updated to minimize (or maximize) the cost. For example, the machine learning model can use a gradient descent (or ascent) algorithm to incrementally adjust the weights to cause the fastest decrease (or increase) in the output of the loss function. The method of updating the parameters of a machine learning model is often called backpropagation.

[0024] Figure 1 is a diagram 100 of a barcode-aware object verification system in accordance with various aspects of the present disclosure. A scanner device 102 is depicted having various components that may be included in a camera-based barcode reader, although additional, fewer, and / or different components may be included in various implementations.

[0025] Scanner device 102 may include one or more processors (such as processor(s) 104), one or more memories (such as non-transitory computer-readable memory 103), a camera 108 (e.g., a CMOS-based camera and / or a CCD-based camera), and / or a light source 110. Light source 110 may effectively output light of any desired wavelength (e.g., infrared) to illuminate a scene for image capture by camera 108. In various examples, light source 110 may be optional. In various examples, memory 103 may store various item templates 116, as described in further detail below. Additionally, in some examples, scanner device 102 may perform the various barcode-aware object verification steps described herein, while in other examples, one or more remote devices (e.g., device(s) 126) may store one or more item templates 116 and / or may perform one or more of the various barcode-aware object verification steps. Device(s) 126 may be a point-of-sale device, a server device, and / or some combination of the two. Device(s) 126 may include non-transitory computer-readable memory 105 , which in some cases may store item templates 116 .

[0026] In various further examples, computer-readable instructions configured to execute a machine learning-based classifier (including parameters of the model) can be stored in memory 103 and / or memory 105. In various examples, the various processing techniques described herein can be performed locally on the scanner device 102 and / or a local point-of-sale computing device to provide low-latency barcode-aware object verification. In various other examples, one or more of the various processing techniques can be performed remotely on one or more computing devices configured to communicate with the scanner device 102 and / or device(s) 126 over a network.

[0027] Memory 103 and / or storage 105 may include one or more non-transitory storage media, such as, for example, volatile and / or non-volatile memory, which may be fixed or removable. Memory 103 may be configured to store information, data, applications, instructions, and the like for enabling scanner device 102 and / or device(s) 126 (and components thereof) to perform various functions according to the various examples described herein. For example, memory 103 may be configured to buffer input data for processing by processor(s) 104. Additionally or alternatively, memory may be configured to store instructions for execution by processor(s) 104. In some cases, memory 103 and / or storage 105 may be considered main memory and included in, for example, RAM or other forms of volatile storage that retains its contents only during operation, and / or memory 103 and / or storage 105 may be included in non-volatile storage, such as ROM, EPROM, EEPROM, FLASH, or other types of storage that retains memory contents independently of the power state of scanner device 102 and / or device(s) 126. Memory 103, memory 105 may also include auxiliary storage devices that store large amounts of data, such as external disk storage devices. In some embodiments, the disk storage devices may communicate with processor(s) 104 via a bus or other routing components using input / output components. The auxiliary storage may include a hard disk, an optical disk, a DVD, a memory card, or any other type of mass storage device known to those skilled in the art.

[0028] The scanner device 102 can be activated (e.g., via pulling a trigger or simply by placing an object in the field of view of the activated camera 108) to capture image data 124 of the physical object including the barcode 122. Figure 1 In the example depicted in , the object is a bottle that includes some text, graphics, and a barcode.

[0029] The memory 103 may store instructions for effectively detecting and decoding a barcode included in the captured image data 124 (at block 130). Detecting the barcode may be performed using an object detector that may determine a bounding box around the barcode in a frame of captured image data. Decoding the barcode may include generating a decoded code (e.g., an alphanumeric string, a numeric code, and / or some other data representation of the code) encoded by the barcode. At block 132, the code may be used to perform a lookup to determine an item template (e.g., among item templates 116) associated with the decoded barcode.

[0030] As described in further detail below, an item template may include identifier data (e.g., a decoded code from a barcode) that identifies a barcode from other barcodes. In some examples, an item template may include a template image. The template image may be an image of all or some portion of the object to which the particular item template belongs. Region of interest data in the item template may define a region of interest for the object. In some further examples, the item template may include data describing one or more of the size, orientation, and position of the barcode in the template image. Additionally, the item template may include data describing one or more of the size, orientation, and position of the region of interest in the template image (which may or may not include all or a portion of the barcode).

[0031] A region of interest associated with the item template can be determined at block 134. In various examples, the region of interest of the captured image data 124 can be determined based on a comparison of the barcode in the captured image data 124 with the barcode in the template image. As described in further detail below, the region of interest in the template image can be applied to the captured image data 124 based on a relationship between the region of interest in the template image and the barcode in the template image and the detected barcode in the captured image data 124. At block 136, a sub-image corresponding to the region of interest from the item template can be generated and applied to the captured image data 124 by cropping the captured image data 124 according to the region of interest defined in the item template. In some further examples, the item template 116 may not store the template image, but instead may store data identifying information regarding the location, size, and / or orientation of the barcode and / or region of interest identified in the template image.

[0032] After generating sub-images corresponding to the regions of interest defined in the item templates 116, the sub-images can be input into a classifier model (block 138). The classifier model can be trained to classify items using the region of interest image data defined in the various item templates 116. The training dataset can include training examples, each of which can include sub-images of items paired with ground truth labels (e.g., identifier data) that identify the item. Thus, a classifier can receive a sub-image as input and predict which item the sub-image belongs to. If the predicted item matches the item template associated with the decoded barcode (e.g., the barcode decoded from the captured image data 124 at block 130), the scanned barcode is considered correctly associated with the object / item (block 140). Conversely, if there is a mismatch between the classifier's predicted item and the decoded barcode, the scanned object can be considered a mismatch with respect to the barcode (block 140). Various actions can be taken in response to a detected mismatch, such as rejecting the transaction, generating an error code indicating an object / barcode mismatch, generating a buzzer sound, etc. The specific output actions depend on the desired implementation and / or user interface design and may vary accordingly.

[0033] Figure 2 Depicted are examples of object templates associated with barcodes according to various aspects of the present disclosure. Figure 2 206 . In some cases, template image 224 may be stored in the item template associated with decoded barcode 206 . However, in some examples, instead of storing template image 224 , data describing the position (e.g., coordinates in the image frame), dimensions (e.g., height and width in pixels), and orientation (e.g., relative to a vertical pixel axis, a horizontal pixel axis, or some other reference) of barcode bounding box 202 and / or data describing region of interest data 208 may be included without storing any image data. This can reduce memory requirements and enable resource-constrained devices (e.g., barcode scanners or other devices) to store more templates. The item template associated with decoded barcode 206 may also store data representing a spatial and / or geometric relationship between the barcode in the item template and the region of interest in the item template. This relationship between the template barcode and the region of interest can be used to determine where to crop the captured image based on the position / size / or orientation of the barcode detected in the captured image to generate a sub-image corresponding to the region of interest of the item.

[0034] like Figure 2As shown, the item template associated with the decoded barcode 206 may include region of interest data 208, which may define a geometric and / or contextual relationship between the barcode of the template image and the region of interest of the template image. Thus, the region of interest data 208 may include data representing the position, size, and / or orientation of the barcode bounding box 202 in the template image. In some further examples, the region of interest data 208 may include data representing the type of barcode (e.g., QR code, UPC code, etc.). Additionally, the region of interest data 208 may include data representing the position, size, and / or orientation of the region of interest in the template image. The region of interest of the template image may or may not include all or a portion of the barcode. For example, in Figure 2 In the illustrated template image 224, the bounding box for the region of interest data 208 includes the barcode bounding box 202. However, in some instances, it may be beneficial for the region of interest defined in the item template to not include the barcode. This is because this may better enable the image classifier network to accurately classify the object presented in the captured image, regardless of whether a ticket swap (e.g., a mismatched barcode) is attached to the object. In some other examples where the region of interest includes a barcode, pixels within the barcode bounding box may be ignored by the classifier network and / or a sub-image may replace all pixels within the bounding box with predefined pixel values ​​to mask the barcode prior to classification. As used herein, the term "bounding box" may refer to data that defines a perimeter around a region of interest (or barcode) in image data. The perimeter defined by the bounding box may be any desired shape (e.g., any polygon such as a square, rectangle, hexagon, etc.). In some cases, the bounding box may be user-defined (e.g., a perimeter drawn by the user around the region of interest). In some examples, the region of interest may be the entire image frame.

[0035] The region of interest data 208 for a given item template can be generated in any desired manner. For example, the region of interest data 208 can be automatically determined (e.g., using image segmentation techniques), can be automatically defined relative to the location of the barcode, or can be manually selected by a user.

[0036] The region of interest data 208 may also describe the spatial and / or geometric relationships of the region of interest. Figure 2As shown, the region of interest data 208 can specify the coordinates of the bounding box of the barcode in the template image and / or the coordinates of the bounding box of the region of interest in the template image. Thus, when a barcode is detected in a captured image, an appropriate region of interest for the captured image can be determined based on the position, size, and / or orientation of the barcode in the captured image and the relationship between the barcode bounding box 202 and the region of interest data 208 in the item template. Thus, a sub-image of the captured image can be generated by cropping the captured image data to generate a portion of the captured image data (e.g., a sub-image) corresponding to the region of interest specified in the item template. Various cropping techniques can be used depending on the desired implementation.

[0037] The item template may also include any other desired information (e.g., price, stock quantity, other identifier codes, etc.). Figure 2 In the example depicted in , an item template may include identifier data 210 that identifies the item template from other item templates. Identifier data 210 may be a decoded code (e.g., an alphanumeric string or a numeric code), such as a barcode, that uniquely identifies the item or item type from other items.

[0038] Figure 3 Depicts an example machine learning architecture (e.g., a convolutional neural network classifier model) that can be used to classify sub-images of objects in accordance with various aspects of the present disclosure. Note that Figure 3 The example machine learning architecture in is just one example of a machine learning based image classifier architecture, and any other desired image classifier may be used to provide barcode-aware object verification in accordance with the various techniques described herein.

[0039] As previously described, the input to the classifier model may be a sub-image 310 that has been generated using the captured image based on the region of interest defined in the item template data for the barcode being decoded. Figure 3 In the example depicted in , sub-image 310 includes a barcode. However, as previously mentioned, in other examples, the region of interest defined in the item template may not include a barcode or may include only a portion of a barcode. In still other examples, the barcode may be masked to improve the accuracy of the classifier model (even in the face of bill swaps). Figure 3 In the example of , the classifier model includes a convolutional neural network (CNN). However, as known to those skilled in the art, other image classification networks can be used. For example, an image classifier based on a recurrent neural network (RNN), an image classifier based on a transformer (e.g., a visual transformer classifier using a visual transformer (ViT)), etc. can be used. Therefore, Figure 3The architecture depicted in is only intended to illustrate one possible implementation.

[0040] The sub-image 310 can be a frame of image data that includes a two-dimensional grid of pixel values. The input sub-image 310 can include additional data, such as a histogram representing the tonal distribution of the image. In some examples, a series of convolution filters can be applied to the image to generate a feature map 312. The convolution operation applies a sliding window filter kernel of a given size (e.g., 3×3, 5×5 in units of pixel height and pixel width) to the image and calculates the dot product of the filter kernel with the pixel values. The output feature map 312 for a single convolution kernel represents the features detected by the kernel at different spatial locations within the input frame of image data. Zero padding can be used at the boundaries of the input image data to allow the convolution operation to calculate values ​​for columns and rows at the edges of the image frame

[0041] Downsampling can be used to reduce the size of the feature map 312. For example, max pooling can be used to downsample the feature map 312 to generate a feature map 314 of reduced size (a modified feature map relative to the feature map 312). Alternatively, other pooling techniques can be used to downsample the feature map 312 and generate the feature map 314. Typically, pooling involves a sliding window filter on the feature map 312. For example, using a 2×2 max pooling filter, the maximum value in the feature map 312 in a given window (at a given frame position) can be used to represent the portion of the feature map 312 in the downsampled feature map 314. Max pooling uses the features that have the greatest impact on a given window and reduces the processing time of subsequent operations. Although in Figure 3 Not shown in FIG, an activation function may be applied to the reduced size feature map 314 after the pooling operation. For example, a rectified linear unit (ReLU) activation function or a sigmoid function may be applied to prevent gradient reduction during training.

[0042] Figure 3 Only a single convolution stage and a single pooling stage are depicted. However, any number of convolution and pooling operations may be used, depending on the desired implementation. Once the convolution and pooling stages are completed, the classifier model may optionally generate a column vector 316 from the resulting feature map by converting the two-dimensional feature map (e.g., an array) into a one-dimensional vector. In various examples, the column vector 316 may be a dense feature representation of the input sub-image 310 (e.g., an embedding representing the sub-image 310).

[0043] The one-dimensional column vector 316 (representing one or all of the feature maps 314, depending on the implementation) can be input into a classifier network for predicting a classifier output 320, which can be a prediction of the object corresponding to the input sub-image 310. In some examples, the classifier network can be a fully connected network (FCN) 318 (e.g., a neural network, a multi-layer perceptron, etc.). However, any other classifier can be used depending on the desired implementation. For example, a random forest classifier, a regression-based classifier, a deep learning-based classifier, etc. can be used. Figure 3 In an example of , the FCN 318 may take as input a one-dimensional column vector 316 representing an input sub-image 310 and may be trained to predict the object to which the input sub-image belongs. In various examples, the FCN 318 may be trained to generate an item identifier (e.g., a decoded barcode code) that identifies a specific item. For example, a classifier model (including a CNN and FCN 318) may be trained using training images that include sub-images (such as predefined regions of interest of a given item defined by an item template) paired with correct identifier data (e.g., decoded barcode codes) as labels for training instances. In some cases, the classifier model may be trained using both positive samples (where the sub-image and barcode are correctly matched) and negative samples (where the sub-image and barcode are not matched). In such examples, the training instances may include data indicating whether a given training instance represents a positive sample or a negative sample. During training of the classifier model (and / or its encoder backbone), a binary cross-entropy loss or any other desired loss function may be used.

[0044] Depending on the implementation, FCN 318 can include any number of hidden layers. In some examples, FCN 318 can be trained in an end-to-end fashion with a convolutional neural network (CNN) to classify an input image as belonging to one of a plurality of items. In at least some other examples, a pre-trained CNN can be used to generate an embedding (e.g., column vector 316) that can be used as input to FCN 318 or other classifiers. In such examples, FCN 318 or other classifiers can be trained without retraining the CNN.

[0045] A normalized softmax layer can be used as part of the FCN 318. The softmax layer can include a node for each object / item that the FCN 318 has been trained to classify. Thus, the classifier output 320 vector can have n dimensions, where n is the number of distinct items. The value of each dimension can be a score for that image class, where all scores in the classifier output 320 vector sum to 1. The element of the classifier output 320 vector with the highest score can be selected as the predicted object for the input sub-image 310.

[0046] At act 330, logic may be implemented (e.g., using computer-executable instructions executed by a barcode scanner, point-of-sale device, and / or any other computing device on which all or part of a barcode-aware object authentication system is deployed) and a determination may be made as to whether the classifier prediction matches the identifier data from the item template / barcode. For example, the classifier model may output predicted identifier data for the item (e.g., a predicted decoded barcode code for the item). This predicted identifier data may be compared to identifier data from the decoded barcode of the captured image (e.g., the actual decoded barcode code from the scanned object). If a match exists, processing may proceed to act 334, where output data indicating the match may be generated. In the case of a point-of-sale transaction, the price of the item may be displayed and checkout may proceed. Conversely, if a mismatch exists, processing may proceed to act 332, where output data indicating the mismatch may be generated. In this case, the transaction may be blocked (because a ticket exchange may have occurred), an error message may be displayed, etc. The specific actions taken in response to a match and / or mismatch depend on the specific implementation and can vary as needed.

[0047] In various examples, instead of or in addition to providing a classification head (e.g., FCN 318), a feature representation of input sub-image 310 (e.g., column vector 316 or another encoded representation of input sub-image 310) can be used as a feature vector and compared to other feature vectors stored in a data repository representing encoded regions of interest for various items (e.g., item vectors representing regions of interest for various different items). For example, a k-nearest neighbor and / or another clustering-based algorithm can be used to determine the vectors (corresponding to other items) that are most similar to the feature vector. In another example, a distance metric can be used to compare the feature vector of input sub-image 310 (e.g., column vector 316) to a database of feature vectors for all relevant items (e.g., all items in inventory). Each of the feature vectors can correspond to corresponding identifier data (e.g., a decoded barcode for the item), allowing the feature vector to be used for the item classification task. The predicted item can be the item having the closest feature vector to the feature vector generated for input sub-image 310. Depending on the desired implementation, various distance metrics (e.g., cosine similarity, cosine distance, Euclidean distance, etc.) can be used.

[0048] In the example using eigenvectors, Figure 3 The CNN depicted in

[15] can be used as an image encoder. Alternatively, various other image encoders besides CNN-based image encoders can be used. For example, an encoder-decoder architecture (e.g., autoencoder, variational autoencoder (VAE), adversarial network, etc.) can be used. In such cases, a reconstruction loss can be used instead of or in addition to the classification loss during training.

[0049] Figures 4A to 4B Example image processing techniques that may be used to generate sub-images corresponding to regions of interest of an object in accordance with various aspects of the present disclosure are shown. Figure 4A In the example, the item may have been scanned using a camera-based barcode reader to generate captured image data. A barcode may be detected in the captured image data. Thus, Figure 4A A captured image is depicted with a detected bounding box around the barcode 406. Once the detected barcode is decoded, the item template data 411 (e.g., identifier data for the barcode) associated with the decoded barcode code may be determined (e.g., using a lookup). The item template data 411 defines the region of interest 402 and the template barcode bounding box 404. In some examples, this information may be stored in the item template data 411 as coordinate data (and / or other data describing the region of interest 402 and the template barcode bounding box 404). In some cases, the template image (e.g., Figure 4AThe barcode image (shown in FIG. 4 ) may be stored in the item template data 411 , while in other scenarios, the template image may not be stored. The region of interest data 440 defines the geometric relationship and / or contextual relationship between the barcode in the template image and the region of interest of the item.

[0050] like Figure 4A As shown, both the orientation and size of the barcode in the captured image are significantly different relative to the barcode in the template image (e.g., the image associated with item template data 411). Additionally, the amount of bottles presented in the captured image, which provides a closer framing of the barcode, is significantly different than that in the template image.

[0051] At act 408, the position (e.g., pixel coordinates of the bounding box), size (e.g., width and height in pixels), and orientation (e.g., rotation angle about a selected axis) of the detected barcode bounding box may be determined. Processing may continue to Figure 4B Act 410. At act 410, the position, size, and orientation of the item template barcode bounding box can be determined. This data can be stored in item template data 411 and / or can be determined using the barcode in the template image. Processing can continue to act 412, where a geometric transformation can be determined so that the detected bounding box (around the barcode 406 in the captured image) corresponds to the template bounding box (e.g., in one or more of position, size, and / or orientation). For example, the ratio of the size of the barcode in the captured image to the barcode in the template image can be determined. The angle and direction of rotation for rotating the barcode in the template image so that it corresponds to the barcode in the captured image can be determined. A translation of the barcode within the image frame can be determined so that the resized and / or reoriented barcode appears in the same position in the frame (relative to the barcode in the captured image).

[0052] At action 414, these geometric transformations may be applied to the region of interest bounding box of the item template data 411 so that the region of interest of the item in the captured image may be determined. Figure 4A 、 4B As shown, the bounding box of the region of interest 402 defined in the template has been rotated and resized to capture a corresponding region of interest (e.g., region of interest 417) of the bottle in the captured image 406. The transformed region of interest bounding box can be applied to the captured image (box 416) as region of interest 417 relative to the position of the barcode bounding box in the captured image 406. Thereafter, the captured image can be cropped to the region of interest to generate a sub-image 418. Sub-image 418 now corresponds to the region of interest of the item defined in the item template data 411 and can be input into a classifier network (e.g., Figure 3 ).

[0053] Figure 5 is a flow chart illustrating an example process 500 for barcode-aware object verification according to various aspects of the present disclosure. Figure 5 The flowchart shown in the figure describes an example process 500, but it should be understood that many other methods of performing the actions associated with process 500 can be used. For example, the order of some of the blocks can be changed, certain blocks can be combined with other blocks, blocks can be repeated, and some of the blocks described can be optional. Process 500 can be performed by processing logic that can include hardware (circuitry, dedicated logic, etc.), software, or a combination of both. In some examples, the actions described in the blocks of process 500 can represent a series of instructions including computer-readable machine code that can be executed by one or more processing units of one or more computing devices. In various examples, the computer-readable machine code can be composed of instructions selected from the native instruction set and / or operating system (or multiple systems) of one or more computing devices.

[0054] Processing may begin at act 510, where a scanner device (e.g., a camera-based scanner device) may capture first image data representing at least a portion of a first item. For example, the scanner device may be used to capture an image of the first item that includes a barcode attached to or otherwise associated with the first item.

[0055] Processing can continue at act 520 where the first barcode in the captured image data can be detected and decoded to generate a decoded barcode code. The decoded barcode code can be an alphanumeric string, a numeric code, and / or any other desired code associated with a barcode.

[0056] Processing can continue at act 530, where a first item template associated with the first barcode can be determined. For example, the decoded barcode code determined at act 520 can be used to perform a lookup in a data structure to determine the first item template associated with the first barcode. The first item template can include first identifier data (e.g., the decoded barcode code) and first region of interest data. The first region of interest data can define a region of interest for the first item. At least a portion of the first region of interest can be a non-barcode portion of the item. In some examples, the first item template can further include a template image. In some cases, the first item template can include data identifying the coordinates of the barcode in the template image and / or the coordinates of the region of interest in the template image. In some examples, the first item template can include size and / or orientation information describing the size and / or orientation of the barcode and / or region of interest in the template image. In some examples, the first item template can define the type of barcode associated with the item (e.g., a UPC1 code, a matrix code, etc.). The first item template can define a geometric and / or contextual relationship between the barcode in the template image and the region of interest in the template image.

[0057] Processing can continue at act 540, where second image data including the first region of interest of the first image data can be generated. The second image data can be a sub-image cropped from the captured first image data to represent the region of interest of the first item presented in the captured first image data. The first region of interest of the first image data can be determined based at least in part on a comparison of one or more of the size, position, and / or orientation of the barcode in the captured first image data with one or more of the corresponding size, position, and / or orientation of the barcode defined by the first item template. A geometric transformation can be determined based on a comparison of the captured barcode and the barcode in the template image. Thereafter, these geometric transformations can be applied to the region of interest in the template image (e.g., defining a bounding box of the region of interest in the template image), and the transformed region of interest bounding box can be applied to the captured image data to determine a region in the captured image data corresponding to the region of interest defined in the template image. The captured image can then be cropped to generate a sub-image corresponding to the region of interest of the first item.

[0058] Processing can continue at act 550, where a first machine learning model (e.g., an image classifier model) can generate predicted item identifier data for the second image data (e.g., a sub-image corresponding to the first region of interest of the first item). At act 560, a determination can be made as to whether the predicted item identifier data output by the first machine learning model matches the first identifier data (e.g., the decoded barcode encoding of the first item). If so, processing can continue to act 570, where output data can be generated indicating that the first barcode matches the first item. This can indicate that no ticket swapping occurred. Thus, if barcode-aware object verification is used as part of the checkout system, the transaction can be allowed to proceed.

[0059] Conversely, if the predicted item identifier data output by the first machine learning model does not match the first identifier data, processing can continue to act 580 where output data can be generated indicating that the first barcode does not match the first item. In the transaction example, if a mismatch is determined, the transaction can be blocked and / or an alert can be generated.

[0060] Figures 6A to 6E Various examples of sub-image extraction based on regions of interest defined in object templates according to various aspects of the present disclosure are depicted. Figure 6A An example captured image 602 captured by a barcode scanner is depicted. In the example image, a barcode 603 has been detected. Furthermore, a region of interest 604 of the captured image 602 has been automatically determined using an object template, wherein the object template is identified using the barcode 603. Figure 6B 606 depicts an example of cropping the captured image 602 to the boundaries of the region of interest 604. Note that image 606 is not yet an extracted sub-image, it merely illustrates an example image cropping technique. Any desired image cropping technique may be used in accordance with various aspects of the present disclosure.

[0061] Figure 6C An example extracted sub-image 610a is depicted (wherein Figure 6B The sub-image is extracted using the cropping technique of the image 606 in FIG. Figure 6D Another example extracted sub-image 610b is depicted, where Figure 6B The cropping method of the image 606 in FIG. 6 is used, and the background of the sub-image has been filled with a predefined color (eg, white).

[0062] Figure 6E Describes the Figure 6B6 and then de-skewing the cropped image according to the object template (e.g., by rotating and / or resizing the image).

[0063] Among other potential benefits, the various barcode-aware object verification systems described herein can enable computing devices and / or barcode scanner devices to use computer vision techniques to detect document swaps. In various examples, the image classifier models described herein can be deployed locally relative to the scanner device (e.g., on the scanner device and / or at a point of sale or other local computing device), which can avoid transmitting image data to back-end systems - a bandwidth and latency intensive operation. In addition, the various techniques described herein enable determining standardized regions of interest for scanned items based on item templates that establish geometric and contextual relationships between the item's barcode and the item's region of interest. Classifying the standardized regions of interest can greatly enhance classification accuracy because the types of images used during inference can be more similar to the types of images seen in the training data.

[0064] In the foregoing description, specific embodiments have been described. However, those skilled in the art will appreciate that various modifications and changes may be made without departing from the scope of the present invention as set forth in the following claims. Accordingly, the description and drawings are to be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of the present teachings.

[0065] These benefits, advantages, solutions to problems, and any element(s) that enable or enhance any benefit, advantage, or solution are not to be construed as critical, required, or essential features or elements of any or all the claims. The invention is defined solely by the appended claims including any amendments made during the pendency of this application and all equivalents of those claims as issued.

[0066] In addition, in this document, relational terms such as first and second, top and bottom, etc. may be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. The terms "comprises," "comprising," "has," "having," "includes," "including," "contains," "containing," or any other variations thereof are intended to encompass a non-exclusive inclusion such that a process, method, article, or apparatus that comprises, has, includes, or contains a list of elements includes not only those elements but also may include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element preceded by "comprises," "has," "includes," or "contains" does not, in the absence of more constraints, preclude the presence of additional identical elements in the process, method, article, or apparatus that comprises, has, includes, or contains the element. The terms "a" and "an" are defined as one or more, unless expressly stated otherwise herein. The terms "substantially," "approximately," "approximately," "about," or any other version of these terms are defined as close as understood by one of ordinary skill in the art, and in one non-limiting embodiment, these terms are defined as within 10%, in another embodiment within 5%, in another embodiment within 1%, and in another embodiment within 0.5%. The term "coupled," as used herein, is defined as connected, although not necessarily directly and not necessarily mechanically. A device or structure that is "configured" in a certain manner is configured in at least that manner, but may also be configured in ways not listed.

[0067] Certain expressions may be used herein to list combinations of elements. Examples of such expressions include: "at least one of A, B, and C"; "one or more of A, B, and C"; "at least one of A, B, or C"; "one or more of A, B, or C." Unless expressly stated otherwise, the above expressions encompass any combination of A and / or B and / or C.

[0068] It will be understood that some embodiments may include one or more special-purpose processors (or "processing devices"), such as microprocessors, digital signal processors, custom processors, and field programmable gate arrays (FPGAs), and uniquely stored program instructions (including both software and firmware) that control the one or more processors to implement some, most, or all of the functions of the methods and / or apparatus described herein, along with certain non-processor circuits. Alternatively, some or all of the functions may be implemented by a state machine without stored program instructions, or in one or more application-specific integrated circuits (ASICs), where each function or some combination of some of the functions is implemented as custom logic. Of course, a combination of these two approaches may also be used.

[0069] In addition, embodiments can be implemented as a computer-readable storage medium having computer-readable code stored thereon for programming a computer (e.g., including a processor) to perform the methods as described and claimed herein. Examples of such computer-readable storage media include, but are not limited to, hard disks, CD-ROMs, optical storage devices, magnetic storage devices, ROMs (read-only memories), PROMs (programmable read-only memories), EPROMs (erasable programmable read-only memories), EEPROMs (electrically erasable programmable read-only memories), and flash memory. Furthermore, it is expected that one of ordinary skill in the art, while making potentially significant efforts and numerous design choices driven by, for example, available time, current technology, and economic considerations, will readily be able to generate such software instructions and programs and ICs with minimal experimentation when guided by the concepts and principles disclosed herein.

[0070] The abstract of the present disclosure is provided to allow the reader to quickly ascertain the nature of the present technical disclosure. This abstract is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the above-mentioned detailed description, it can be seen that various features are grouped together in various embodiments for the purpose of integrating the present disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed embodiments require more features than expressly recited in the various claims. On the contrary, as reflected in the following claims, the utility model subject matter lies in less than all the features of a single disclosed embodiment. Therefore, the following claims are hereby incorporated into the detailed description, with each claim representing itself as a separately claimed subject matter.

Claims

1. A method comprising: capturing, by a scanner device including an image sensor, first image data representing at least a portion of a first item; decoding, by the scanner device, a first barcode present in the first image data; determining a first item template associated with the first barcode, the first item template comprising first identifier data identifying the first item from other items and first region of interest data specifying a first region of interest of the first item; generating second image data of the first region of interest including the first image data; determining, by a first machine learning model, that the second image data corresponds to the first identifier data identifying the first item; as well as First data is generated indicating that the first barcode matches the first item.

2. The method according to claim 1, wherein The first machine learning model includes a convolutional neural network classifier or a visual transformer classifier trained to classify a given item based on images of a predefined region of interest of the given item.

3. The method of claim 1, further comprising: generating, by the first machine learning model, a first vector representing the second image data; comparing the first vector to a plurality of item vectors stored in a data warehouse; determining a second vector from the plurality of item vectors based at least in part on a first distance metric for determining a distance between the first vector and the second vector; as well as The second vector is determined to be associated with the first identifier data in the first item template, wherein the determination that the second image data corresponds to the first identifier data is made at least in part based on associating the second vector with the first identifier data.

4. The method of claim 1, further comprising: determining a first bounding box around the first barcode in the first image data using an object detector; determining a first size of the first bounding box; determining a first orientation of the first bounding box; determining a second size of a second barcode associated with the first region of interest data of the first item template; as well as A ratio between the first size and the second size is determined.

5. The method of claim 4 , further comprising determining the first region of interest of the first image data based at least in part on: resizing a second bounding box corresponding to the first region of interest of the first item in the first item template using the ratio; and The resized second bounding box is applied to the first image data.

6. The method of claim 1, further comprising: capturing, by the scanner device, third image data representing at least a portion of a second article; decoding, by the scanner device, a second barcode present in the third image data; determining a second item template associated with the second barcode, the second item template comprising second identifier data identifying the second item from other items and second region of interest data specifying a second region of interest of the second item, the second region of interest comprising the second barcode and a second non-barcode portion of the second item; generating fourth image data of the second region of interest including the third image data; determining, by the first machine learning model, that the fourth image data does not match the second barcode; and First output data is generated indicating that the second barcode does not match the second item.

7. The method of claim 1, further comprising: generating third image data representing a second region of interest of a second item, the second region of interest representing a second barcode of the second item and at least a second non-barcode portion of the second item; generating second identifier data that identifies the second item from other items; generating a first training instance comprising the third image data and the second identifier data; as well as The first machine learning model is trained to classify items using a training dataset including the first training examples.

8. The method of claim 1, wherein: The first region of interest of the first item includes the first barcode and a non-barcode portion of the first item.

9. The method of claim 1, wherein: The first item template represents at least one of a contextual relationship or a geometric relationship between the first barcode and the first region of interest of the first item.

10. The method of claim 1, wherein: The first item template further includes data representing a barcode type of the first barcode.

11. The method of claim 1, wherein: The first item template further includes: a template image of the first region of interest of the first object; and At least one of coordinate data indicating a position of the first barcode in the template image, orientation data indicating an orientation of the first barcode in the template image, or size data indicating a size of the first barcode in the template image.

12. A system comprising: Image sensor; at least one processor; as well as a non-transitory computer-readable memory storing instructions that, when executed by the at least one processor, are effective to: controlling the image sensor to capture first image data representing at least a portion of a first item; decoding a first barcode present in the first image data; determining a first item template associated with the first barcode, the first item template comprising first identifier data identifying the first item from other items and first region of interest data specifying a first region of interest of the first item; generating second image data of the first region of interest including the first image data; determining, using a first machine learning model, that the second image data corresponds to the first identifier data identifying the first item; as well as First data is generated indicating that the first barcode matches the first item.

13. The system of claim 12, wherein: The first machine learning model includes a convolutional neural network classifier or a visual transformer classifier trained to classify a given item based on images of a predefined region of interest of the given item.

14. The system of claim 12, the non-transitory computer readable memory storing further instructions that, when executed by the at least one processor, are further effective to: generating, by the first machine learning model, a first vector representing the second image data; comparing the first vector to a plurality of item vectors stored in a data warehouse; determining a second vector from the plurality of item vectors based at least in part on a first distance metric for determining a distance between the first vector and the second vector; as well as The second vector is determined to be associated with the first identifier data in the first item template, wherein the determination that the second image data corresponds to the first identifier data is made at least in part based on associating the second vector with the first identifier data.

15. The system of claim 12, the non-transitory computer readable memory storing further instructions that, when executed by the at least one processor, are further effective to: determining a first bounding box around the first barcode in the first image data using an object detector; determining a first orientation of the first bounding box; determining a second size of the barcode associated with the first region of interest data of the first item template; as well as An amount of rotation between the first orientation and the second orientation is determined.

16. The system of claim 15, the non-transitory computer readable memory storing further instructions that, when executed by the at least one processor, are further effective to: reorienting a second bounding box corresponding to the first region of interest of the first item in the first item template based on the rotation amount; and The reoriented second bounding box is applied to the first image data.

17. The system of claim 12, the non-transitory computer readable memory storing further instructions that, when executed by the at least one processor, are further effective to: controlling the image sensor to capture third image data representing at least a portion of a second item; decoding a second barcode present in the third image data; determining a second item template associated with the second barcode, the second item template comprising second identifier data identifying the second item from other items and second region of interest data specifying a second region of interest of the second item, the second region of interest comprising the second barcode and a second non-barcode portion of the second item; generating fourth image data of the second region of interest including the third image data; determining, by the first machine learning model, that the fourth image data does not match the second barcode; and First output data is generated indicating that the second barcode does not match the second item.

18. The system of claim 12, the non-transitory computer readable memory storing further instructions that, when executed by the at least one processor, are further effective to: generating third image data representing a second region of interest of a second item, the second region of interest representing a second barcode of the second item and at least a second non-barcode portion of the second item; generating second identifier data that identifies the second item from other items; generating a first training instance comprising the third image data and the second identifier data; as well as The first machine learning model is trained to classify items using a training dataset including the first training examples.

19. A method comprising: receiving first image data representing at least a portion of a first item; decoding a first barcode represented in the first image data; determining a first item template associated with the first barcode, the first item template comprising first identifier data identifying the first item from other items and first region of interest data specifying a first region of interest of the first item; generating second image data of the first region of interest including the first image data; determining, by a first machine learning model, that the second image data corresponds to the first identifier data identifying the first item; as well as First data is generated indicating that the first barcode matches the first item.

20. The method of claim 19, further comprising: generating, by the first machine learning model, a first vector representing the second image data; comparing the first vector to a plurality of item vectors stored in a data warehouse; determining a second vector from the plurality of item vectors based at least in part on a first distance metric for determining a distance between the first vector and the second vector; as well as The second vector is determined to be associated with the first identifier data in the first item template, wherein the determination that the second image data corresponds to the first identifier data is made at least in part based on associating the second vector with the first identifier data.