Product authentication using packaging
Machine learning models analyze packaging images in real-time to enhance security and reliability in product authentication, addressing the challenge of counterfeit products by reducing false positives and optimizing computational resources.
Patent Information
- Application Number
- PCT/US2024/062081
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-27
- Publication Date
- 2025-07-03
AI Technical Summary
Counterfeit products are often packaged in authentic-looking packaging, making it difficult to authenticate the product's authenticity through traditional methods that rely on physical markers, which can be easily replicated.
A method using machine learning models to analyze packaging images in real-time, ensuring high accuracy in image capture and processing by identifying packaging faces, assessing capture conditions, and providing real-time feedback to enhance security without additional physical components.
This approach significantly reduces false positives in authenticity determinations, optimizes computational efficiency, and ensures reliable product authentication by leveraging dual-layered machine learning models for image analysis and feedback.
Smart Images

Figure US2024062081_03072025_PF_FP_ABST
Abstract
Description
PRODUCT AUTHENTICATION USING PACKAGINGCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of the filing date of U.S. Application No. 18 / 400,874, filed on December 29, 2023, the entirety of which is incorporated herein by reference.FIELD OF THE DISCLOSURE
[0002] Technologies are described for authenticating products.BACKGROUND
[0003] Counterfeit products may be packaged in packaging that differs from the packaging used for authentic products despite an attempt to make the counterfeit packaging look authentic. Analysis of the packaging can thus indicate an authenticity of the product within.SUMMARY
[0004] Some aspects of this disclosure describe a method. The method includes capturing, at a mobile device, an image using a camera of the mobile device; processing, at the mobile device, the image using one or more machine learning models, wherein the one or more machine learning models have been trained to identify a face of first packaging in the image, and determine whether the first packaging in the image satisfies one or more capture conditions; providing, at the mobile device, feedback for image capture based on a first output of the one or more machine learning models relating to the one or more capture conditions; and in response to output of the one or more machine learning models indicating that the one or more capture conditions are satisfied, and in response to the output of the one or more machine learning models indicating that the face of the first packaging is present in the image, sending the image for authentication of the first packaging.
[0005] This and other methods described herein can have one or more of at least the following characteristics.
[0006] In some implementations, the one or more machine learning models have been trained to determine a face type of the face of the first packaging.
[0007] In some implementations, the face type includes a front face or a rear face.
[0008] In some implementations, the method includes determining whether the first packaging is authentic. Determining whether the first packaging is authentic includes: selecting, from two or more faces of second packaging, a first face based on the face type of the face of the first packaging matching a face type of the first face of the second packaging; determining at least one similarity between the first face and the face of the first packaging; and selecting, from among a plurality of images of packaging, an image of the second packaging as a reference image based on the at least one similarity between the first face and the face of the first packaging.
[0009] In some implementations, the at least one similarity includes a textual similarity between text included on the face of the first packaging and text included on the first face of the second packaging, and a graphical similarity between the reference image and the image.
[0010] In some implementations, determining whether the first packaging is authentic includes, in response to selecting the image of the second packaging as the reference image, determining whether the first packaging is authentic based on a comparison between the first packaging in the image and the second packaging in the reference image. In some implementations, the method includes determining whether the first packaging is authentic. Determining whether the first packaging is authentic includes: determining whether the image includes a data-encoding symbol; in response to determining that the image includes the data-encoding symbol, decoding data encoded by the data-encoding symbol, and determining a reference image based on the data, or in response to determining that the image does not include the data-encoding symbol, determining the reference image based on a graphical comparison between the image and the reference image.[OU] In some implementations, the method includes determining whether the first packaging is authentic. Determining whether the first packaging is authentic includes: receiving the image from the mobile device, and processing the image using a machine learning model distinct from a first machine learning model, of the one or more machinelearning models, that has been trained to identify the face of the first packaging in the image.[0121 In some implementations, the method includes determining whether the first packaging is authentic. Determining whether the first packaging is authentic includes: determining a textual similarity between text in the image and text in a reference image; determining a graphical similarity between the image and the reference image; determining, based on at least one of the textual similarity or the graphical similarity, that the first packaging is not authentic; and determining, based on the image, a packaging of which the first packaging is a counterfeit.
[0013] In some implementations, the one or more capture conditions are based on at least one of an orientation of the first packaging in the image or a level of corruption in the image.
[0014] In some implementations, the feedback for image capture includes at least one of: a graphical bound for placement of the first packaging during image capture, the graphical bound being moved to different locations on a display of the mobile device over capture of multiple images, an indication of whether an orientation of the first packaging satisfies an orientation condition, or a progress indicator that progresses based on satisfaction of the one or more capture conditions.
[0015] In some implementations, the feedback for image capture includes an indicator of a location of corruption in the image.
[0016] In some implementations, the method includes training the one or more machine learning models. Training the one or more machine learning models includes: obtaining an image of reference packaging; generating a plurality of images by modifying at least one of orientation, background, or contrast of the image of the reference packaging; and training the one or more machine learning models using the plurality of images as training data.
[0017] In some implementations, the method includes determining whether the first packaging is authentic. Determining whether the first packaging is authentic includes: comparing at least one feature of the first packaging to a digital blueprint of reference packaging, the digital blueprint including a label indicating a face type of a face of thereference packaging, a graphical representation of the face of the reference packaging, and text included on the face of the reference packaging.[0181 In some implementations, the method includes generating the digital blueprint. Generating the digital blueprint includes: processing an image of the reference packaging using a machine learning model that has been trained to determine the face type of the face of the reference packaging; and generating the digital blueprint based on an output of the machine learning model that has been trained to determine the face type of the face of the reference packaging.
[0019] In some implementations, the method includes training the one or more machine learning models using as training data, images of faces of a plurality of packaging, and as labels for the training data, data indicative of types of faces of the plurality of packaging portrayed in the images.
[0020] In some implementations, the one or more machine learning models include a first machine learning model that has been trained to identify the face of the first packaging in the image, and a second machine learning model that has been trained to determine whether the first packaging in the image satisfies the one or more capture conditions.
[0021] In some implementations, the method includes training the one or more machine learning models. Training the one or more machine learning models includes: providing, in a user interface, a display of an image of reference packaging captured by a second mobile device; processing the image of the reference packaging using a machine learning model that has been trained to identify a face of the reference packaging in the image of the reference packaging, to obtain, as an output, an auto-annotation indicative of at least one of text included in the face of the reference packaging, or a face type of the face of the reference packaging; providing, in the user interface, one or more tools usable to manually alter the auto-annotation to obtain a modified annotation; and training the one or more machine learning models using, as training data, the image of the reference packaging and the modified annotation.
[0022] Some aspects of this disclosure relate to another method that includes: obtaining a first image of first packaging; identifying a first plurality of regions in the first image based on content in the plurality of regions; determining similarities between the firstplurality of regions and a second plurality of regions in a second image of second packaging; and determining, based on the similarities, whether the first packaging is authentic.
[0023] This and other methods described herein can have one or more of at least the following characteristics.
[0024] In some implementations, the method includes identifying the plurality of regions using a trained machine learning model.
[0025] In some implementations, determining the similarities between the first plurality of regions and the second plurality of regions includes determining a displacement between a first region in the first plurality of regions and a corresponding second region in the second plurality of regions.
[0026] In some implementations, determining the similarities between the first plurality of regions and the second plurality of regions includes determining a graphical similarity between a first region in the first plurality of regions and a corresponding second region in the second plurality of regions.
[0027] In some implementations, determining the similarities between the first plurality of regions and the second plurality of regions includes determining a text morphology similarity between a first region in the first plurality of regions and a corresponding second region in the second plurality of regions.
[0028] In some implementations, determining the similarities between the first plurality of regions and the second plurality of regions includes determining a text content similarity between a first region in the first plurality of regions and a corresponding second region in the second plurality of regions.
[0029] In some implementations, the method includes determining whether the packaging is authentic by: identifying variable data in the first image; and determining whether the variable data is consistent.
[0030] Some aspects of this disclosure relate to another method that includes: obtaining a first image of first inauthentic packaging, the first image annotated with at least one region including inauthentic content; projecting the at least one region onto a second image of authentic packaging, to annotate the second image of authentic packaging with at least one region including authentic content; training a machine learning model toidentify inauthentic packaging using, as training data, the annotated first image and annotated second image; providing a test image as input to the machine learning model; and obtaining, as output of the machine learning model, a determination of authenticity of packaging in the test image.
[0031] The described methods can be associated at least with corresponding systems, processes, devices, and / or instructions stored on non-transitory computer-readable media. For example, some aspects of this disclosure describe a non-transitory computer-readable medium tangibly encoding a computer program operable to cause a data processing apparatus to perform operations of the foregoing method and / or other methods described herein. Further, some aspects of this disclosure describe a system including one or more computers programmed to authenticate images of packaging; and a mobile device communicatively coupled with the one or more computers, the mobile device being programmed to perform operations of the foregoing method and / or other methods described herein and to send the images of packaging to the one or more computers.
[0032] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other aspects, features and advantages will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0033] FIG. 1 is a diagram showing an example of image processing.
[0034] FIG. 2 is a diagram showing an example of a system associated with image processing.
[0035] FIG. 3 is a diagram showing an example of a packaging authentication process.
[0036] FIG. 4 is a diagram showing an example of a system associated with packaging authentication.
[0037] FIG. 5 is a diagram showing an example of an image processing and machine learning model training process.
[0038] FIG. 6 is a diagram showing an example of corruption-related processing.
[0039] FIGS. 7A-7B are examples of screen displays associated with image capture.
[0040] FIG. 8 is a diagram showing examples of screen displays associated with image capture.
[0041] FIGS. 9A-9C are examples of user interfaces associated with annotation.
[0042] FIG. 10 is a diagram showing examples of generated images.
[0043] FIG. 11 is an example of a user interface associated with annotation.
[0044] FIGS. 12A-12E are examples of screen displays associated with image capture.
[0045] FIG. 13A is a diagram showing an example of a packaging authentication process.
[0046] FIG. 13B is a diagram showing an example of a comparison according to the process of FIG. 13A.
[0047] FIG. 14A is a diagram showing an example of a packaging authentication process.
[0048] FIG. 14B is a diagram showing examples of regions associated with the process of FIG. 14A.
[0049] FIG. 15 is a diagram showing an example of a system associated with packaging authentication.
[0050] FIG. 16 is a diagram showing examples of images of packaging and corresponding outputs of a registration module.
[0051] FIG. 17 is a diagram showing examples of images of packaging and corresponding outputs of a text morphology module.
[0052] FIG. 18 is a diagram showing examples of images of packaging and corresponding outputs of a text accuracy module.
[0053] FIG. 19 is a diagram showing an example of an image of packaging corresponding outputs of a variable data module.
[0054] FIG. 20A is a diagram showing an example of a packaging authentication process.
[0055] FIG. 20B is a diagram showing examples of regions associated with the process of FIG. 20 A.
[0056] FIG. 21 is a diagram showing an example of a user interface associated with suspect packaging annotation.
[0057] FIG. 22 is a diagram showing an example of a user interface associated with forensic analysis.DETAILED DESCRIPTION
[0058] This disclosure relates to capturing and processing images of packaging for product authentication. Traditional methods of product authentication often rely on physical markers or tags, which can be easily replicated by counterfeiters. This disclosure addresses these limitations by, in part, leveraging advanced machine learning models to analyze packaging images, providing a non-intrusive, scalable solution that enhances security without the need for additional physical components. This disclosure describes integration of machine learning models that not only identify packaging faces but also dynamically assess capture conditions in real-time. This dual-layered approach ensures high accuracy in image capture and processing, significantly reducing false positives in authenticity determinations. For example, images of packaging can be processed using one or more machine learning models to facilitate capture of suitable images and to identify packaging in the images. A machine learning model trained to identify packaging can ensure that only images that include packaging undergo authentication testing, reducing computational errors and unnecessary bandwidth usage. In addition, a machine learning model trained to identify a face type (e.g., front packaging face or rear packaging face) can significantly reduce the search space for image comparison, leading to faster authentication and reducing usage of computational resources for authentication. Region-based analyses can provide more-accurate packaging authentication and useful outputs for forensic analyses. Other machine learning models and processes described herein allow for the fast and efficient processing of images of packaging to create “digital blueprints” against which future images of packaging can be compared for authentication testing. Users can upload images of packaging, have the images auto-annotated using machine learning (in some cases with the option for manual annotation), and use the annotated images to train machine learning models without requiring technical, machine learning-specific tasks on the part of the users. In some implementations, the system’s architecture is designed to optimize computational efficiency by distributing tasks between mobile devices and cloud-based servers. This hybrid processing model allows for real-time feedback during image capture while offloading intensive computations to cloud resources, thereby enhancing user experience and system scalability.
[0059] FIG. 1 shows an example of a process 100 according to some implementations of this disclosure. In some implementations, the process 100 can be performed by a mobile device, e.g., a smartphone, tablet, wearable device, laptop, or another type of mobile device. Performing the process 100 on a mobile device (e.g., a mobile device that captures images for processing) allows for real-time user feedback and correction, improving user experience. Moreover, as discussed in further detail below, the most computationally-intensive aspects of the processes discussed herein (e.g., reference image identification and / or image comparison) can be performed at a computer system remote from the mobile device (e.g., a cloud computing system), such that the mobile device can perform real-time tasks associated with image capture, while the remote computer system with more computational resources (e.g., storage and / or processing resources) can perform more computationally-intensive tasks, providing for overall efficient authentication. However, the process 100 need not be performed by a mobile device; elements of the process 100, and of other processes discussed herein, such as process 300 and process 500, can in various implementations be performed by a mobile device, by a remote computer system, or by the two in combination.
[0060] FIG. 2 shows an example of a system 200 associated with process 100. For example, elements of the system 200 can be configured to perform process 100. For example, the system 200 can include one or more computers 201 (e g., a data processing apparatus) and a non-transitory computer-readable medium 203 tangibly encoding a computer program operable to cause the one or more computers 201 to perform process 100 and / or associated operations described herein. The one or more computers 201 and / or the non-transitory computer-readable medium 203 can implement at least some of the modules of FIG. 2.
[0061] The system 200 includes a mobile device camera 202, a machine learning module 204 implementing one or more machine learning model (in this example, a face identification model 206 and one or more capture condition models 208), an image suitability module 214, a feedback module 210, a transmission module 216, and a mobile device display 212. The mobile device camera 202 and the mobile device display 212 can be included in the same mobile device, e.g., a mobile device that also includes / implements the modules 204, 210, 214, 216.
[0062] The modules 204, 210, 214, 216 can be hardware and / or software modules, e.g., implemented by one or more computer systems. For example, the modules 204, 210, 214, 216 can include one or more hardware processors and one or more computer- readable mediums encoding instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations such as portions of process 100 and / or other processes described herein. The modules 204, 210, 214, 216 can be software modules executed by one or more hardware processors based on instructions encoded in one or more computer-readable mediums. In some implementations, the modules 204, 210, 214, 216 are modules of a mobile device, such that process 100 and associated operations can advantageously be performed largely or entirely on a mobile device, in some cases reducing processing latency compared to processes that require execution on a remote system. The modules 204, 210, 214, 216 need not be separate but, rather, can be at least partially integrated together as one or more combined modules or fully integrated together as a single program on the mobile device.
[0063] The process 100 includes capturing an image (sometimes referred to as a “test image”) using a camera of a mobile device (102), e.g., mobile device camera 202. For example, the test image can be an image of product packaging that a user would like to test to determine the authenticity of the product. For example, upon receiving a shipment of product, a merchant can capture an image of the product on their smartphone and quickly test the authenticity of the product. More specifically, processes described herein are used to test the authenticity of packaging, and, in many cases, it can be presumed that the authenticity of the product within the packaging follows the authenticity of the packaging. For example, some implementations according to this disclosure can be used to test the authenticity of medicine packaging, and counterfeit medicine may be presumed to be packaged in counterfeit packaging.
[0064] Capturing an image can include obtaining signals and / or data representative of the image (e.g., signals from photodetectors of the mobile device camera 202) and can include one or more processing steps performed on the image, such as downscaling / upscaling, size and / or resolution adjustment, computational distortioncompensation, brightness adjustment (e.g., to compensate for underexposure / overexposure), and / or one or more other image adjustment operations. [0651 The image can be captured as an individual image or in a sequence of images, e.g., in a sequence of images that form a video or live image. The capture of multiple images can allow a user to alter image capture in real-time (e.g., in response to feedback from the feedback module 210, as discussed in further detail below),
[0066] The process 100 further includes processing the test image using one or more machine learning models (104). The model(s) have been trained to (i) identify a face of first packaging in the test image, and ((ii) to determine whether the first packaging in the test image satisfies one or more capture conditions.
[0067] For example, in the system 200 of FIG. 2, a machine learning module 204 is configured to process the test image using a face identification model 206 and one or more capture condition models 208. The face identification model 206 is configured to identify a face of packaging in the test image, and the one or more capture condition models 208 are configured to determine whether capture conditions of the test image are satisfied. These processes are discussed in further detail below. However, the machine learning model(s) need not be divided in this manner. For example, in some implementations, a single machine learning model is configured to both identify a face of packaging (e g., and / or perform other tasks described herein in relation to the face identification model 206) and to determine whether capture conditions are satisfied (e.g., and / or perform other tasks described herein in relation to the capture condition model(s) 208).
[0068] The face identification model 206 is configured to identify a face of packaging in the test image, e.g., to determine whether a packaging face is present in the image and, in some implementations, to determine location(s) of the face (e.g., a bounding box for the face). For example, the face identification model can be trained as discussed with respect to FIG. 5, or by another method. For example, the face identification model 206 can be trained using, as training data, (i) images that include faces of packaging (and, in some cases, images that do not include faces of packaging), and (ii) as labels of the images, an indicator of whether a packaging face is included and / or location(s) of thepackaging face, in some cases with further label(s) such as a face type, as discussed in further detail below.[0691 Face identification can serve as a prerequisite for further processing, to reduce wasted computational resources and user time. For example, it may be undesirable to perform a full authentication process on images without packaging, because such processing may (i) consume a significant amount of computational resources, and / or (ii) introduce error into datasets / machine learning models that are based on the images, e.g., because the datasets are intended to include only images with packaging or because the machine learning models are intended to be trained on images with packaging. Moreover, performing a “check” for a packaging face can allow a user to adjust image capture if no packaging face is detected, e.g., by adjusting focus settings or by bringing the mobile device camera 202 closer to the packaging. In some implementations, feedback provided to the user can include an indication of whether a packaging face was detected, as discussed in further detail below with respect to the feedback module 210.
[0070] The face identification model 206 can be an image classification model (e.g., to determine whether or not a packaging face is included in the test image) or an object detection model (e.g., to determine a location of the packaging face in the test image). In some implementations, one or more of the machine learning models discussed herein, including the face identification model 206, is configured specifically for use on mobile devices, e.g., can be a TensorFlow Lite model. As such, the model(s) can be deployed as edge models on mobile devices, e.g., so that captured images can be processed locally without having to be sent to a remote server for processing. As such, in some implementations, the bandwidth associated with image transmission can be reduced.
[0071] In some implementations, the face identification model 206 is configured to determine a face type of a packaging face in the test image. A “face type” can be, for example, a specific orientation of the face with respect to the packaging, for example, a front face, a rear face, or a side face. A “face” need not be flat (e.g., the packaging need not be cuboid) but, rather, can be curved. As discussed in further detail below with respect to FIGS. 3-4, the identification of a face type in the test image can reduce the search space for image-matching and facilitate faster, more computationally efficient authentication.
[0072] The capture condition model(s) 208 are configured to determine whether capture conditions are satisfied, and / or to provide information that can be used to make such a determination. For example, even in cases in which a packaging face is present (e.g., as determined by / using the face identification model 206), the test image may be unsuitable for authentication, such as because the packaging face is not sufficiently infocus, because corruption (e.g., shadows / glare) is present in the image, and / or because the packaging face is not correctly oriented (e.g., tilted or inverted). Based on an output of the capture condition model(s) 208, corrective feedback can be provided to a user. Training of the capture condition model (s) 208 is discussed below with respect to FIG. 5.
[0073] Non-limiting examples of outputs of the capture condition model(s) 208 include: location(s) and / or level(s) of corruption (e.g., glare and / or shadow) in the test image; an orientation of a packaging face in the test image; a focus level of the packaging face in the test image; and / or a capture proportion of the packaging face in the test image (e.g., whether the entire face is imaged or whether a portion of the face is cut-off in the image frame). In some implementations, one or more different models are trained to provide one or more of these different outputs. In some implementations, a single model is trained to provide all outputs of the capture condition model(s) 208. Accordingly, hereafter the capture condition model(s) 208 are referred to as a capture condition model 208, with the understanding that two or more models can be trained as described and used to performed the described functions.
[0074] For example, in some implementations, the capture condition model 208 includes (i) an orientation detection model trained to detect inverted, tilted, and correct packaging orientations, (ii) a warp / out-of-plane model trained to detect warpage and / or out-of-plane tilts in the packing, and (iii) a corruption detection model trained to detect one or more forms of corruption, such as bright spots, diffused glare, spot glare, blooming, and / or shadow. In some implementations, the orientation detection model is trained on images of packaging, e.g., the same or similar training data as is used to train the face identification model 206. This is discussed in further detail with respect to FIG.5. The warp / out-of-plane model and the corruption model can be trained using that or similar training data, and / or trained using a generic pool of images (e.g., smartphone- captured images) labeled with suitable labels such as “warped,” “out of plane,” and / orcorruption labels. As such, the warp / out-of-plane model and the corruption model can be trained agnostic to any particular packaging features and stock keeping unit (SKU). For purposes of this disclosure, it has been recognized that the wrap / out-of-plane model is particularly useful in the context of image capture by mobile devices, in view of typical imaging patterns by users. As such, the inclusion of this model, and / or its associated operations related to confirming satisfactory capture conditions, can provide significant benefits for analysis efficiency and user experience. The incorporation of warp / out-of- plane detection models represents a significant technical improvement over existing systems. By identifying and correcting potential distortions during image capture, the system ensures that only high-quality images are processed, thereby improving the reliability of subsequent authentication analyses.
[0075] In some implementations, the improved image processing techniques for cellphone images utilize a Brute-Force (BF) matcher for feature extraction. The use of BF matcher in feature extraction enhances alignment accuracy between captured images and reference templates. Additionally, in some implementations warp-detection models are integrated to guide users in capturing high-quality images by detecting and correcting potential distortions in real-time. In some implementations, integration of warp-detection models in edge devices prevents poor image captures by providing real-time feedback on image alignment. Referring again to FIG. 1, feedback for image capture is provided based on a first output of the one or more machine learning models relating to the one or more capture conditions (106). For example, the capture condition model 208 can output information relating to the capture conditions, and the feedback module 210 can generate and display (e.g., on the mobile device display 212) feedback based on the information relating to the capture conditions. In some implementations, providing feedback is conditional upon one or more of the capture conditions being unsatisfied, e.g., on a determination by the image suitability module 214 that feedback is required to improve image capture conditions. In some implementations, providing feedback can be performed even when image capture is satisfactory, e.g., to indicate that the capture is satisfactory and / or to provide a level of progress of image capture.
[0076] FIG. 6 illustrates an example of a captured test image 602; an image representing a cormption segmentation map 604 that can be output by someimplementations of the capture condition model 208 or determined based on an output of the capture condition model 208; and an overlaid image 606 displayed by the feedback module 210 based on the corruption segmentation map. The test image 602 includes a packaging face 608. The segmentation map 604 identifies areas exhibiting one or more types of corruption in the test image 602. In this case, the segmentation map 604 identifies a diffused glare region 612 and a spot glare region 610. For example, the capture condition model 208 can be trained to identify (and output) types of corruption in images, and the regions exhibiting the corruption. The overlaid image 606 includes the packaging face 608 overlaid by overlays 616, 614 indicating the diffused glare region 612 and the spot glare region 610, respectively. For example, feedback including the overlaid image 606 can be displayed alongside a message instructing the user to adjust image capture to reduce the prominence of the glare, e.g., by changing an image capture angle and / or by changing lighting conditions. Although shown as opaque in FIG. 6, in some implementations an overlay can be partially transparent to show the full test image underneath the overlay. For example, different-colored transparent overlays can be used to illustrate different types of corruption (e.g., spot glare, diffused glare, and / or shadow). In some implementations, the colors can match feedback provided by the feedback module 210. For example, the feedback can instruct a user to “decrease glare in red region,” and a red overlay can be displayed over a region exhibiting glare.
[0077] In some implementations, the capture condition model 208 is trained to output a proportion of the test image and / or of the packaging face in the test image that is obscured by corruption. This proportion can be compared to a threshold proportion, e.g., as discussed below in reference to the image suitability module 214.
[0078] FIG. 7A illustrates examples of displays 702, 704 that can be provided by the feedback module 210 based on an output of the capture condition model 208. In this example, the capture condition model 208 has been trained to identify an orientation of a packaging face 706. The feedback includes a target box 708 representing a target positioning of the packaging face 706. The inclusion of a target box 708 (or other indicator of the target positioning) can aid in the capture of multi-pose test images, in which multiple test images are captured so that corrupted region(s) of the packaging face in any one of the images may be uncorrupted in other images, such that the entirepackaging face can be analyzed in aggregate. For example, in some implementations the target box 708 is moved between different positions of the mobile device display 212 between capture of different images, to encourage moving the packaging and aid in capturing the packaging under a wider variety of conditions.
[0079] For detecting the orientation of the packaging face, the capture condition model 208 can be trained to determine whether the packaging face is tilted with respect to a target orientation and, in some implementations, a degree and / or characteristic of the tilt. For example, the target orientation can be a right-side-up orientation in which a top side of the packaging face faces a top of the image frame and a bottom side of the packaging face faces a bottom of the image frame (e.g., a roll angle of zero), a ninetydegree-rotated orientation in which the top side of the packaging face faces the right side or the left side of the imaging frame and the bottom side of the packaging face faces the left side or the right side, respectively, and / or an orientation in which the packaging face is imaged head-on (e.g., pitch and / or yaw angles of zero). In some implementations, the capture condition model 208 is trained to determine a quantity associated with the orientation, such as an angular degree of roll, pitch, and / or yaw. In some implementations, the capture condition model 208 is trained to classify the orientation, for example, into “correct orientation,” “tilt,” or “inverted,” where “correct orientation” can be, for example, a tilt within a predetermined difference from the target orientation, and where “inverted” can be an up-side down image or a left-right-inverted. “Tilt” can include a tilted orientation in one or more dimensions, e.g., pitch, yaw, and / or roll.
[0080] In the case of the display 702, the packaging face 706 is captured with a tilted orientation. In the case of the display 704, the packaging face 706 is captured with an inverted orientation. The displays 702, 704 can include one or more feedback elements to indicate the incorrect orientation to the user. For example, in some implementations the target box 708 can be displayed in a particular color to indicate the incorrect orientation (e g., red), and / or a progress bar 710 can be made to not advance and / or be displayed in a particular color to indicate the incorrect orientation. Based on the feedback, the user can adjust the orientation. In some implementations, the capture condition model 208 is trained to determine whether the packaging face is in the target box 708, and the packaging face being substantially out of the target box 708 (e.g., a proportion of thepackaging face out of the target box 708 being above a threshold) can correspond to an incorrect orientation.[0811 FIG. 7B illustrates examples of displays 720, 722 that can be provided by the feedback module 210 in response to the capture condition model 208 indicating a correct orientation. In the displays 720, 722, the target boxes 724, 726 and the progress bars 728, 730 are changed to green (from red) to indicate the correct orientation. In addition, the progress bars 728, 730 progress from a first state in display 720 to a complete state in display 722, indicating progress in image capture based on the correct orientation. In this example, image capture results in capture of multiple multi-pose images: display 720 shows a single image 732, while display 722, corresponding to later in the image capture process, shows multiple images 734, one or more of which can be used for packaging authentication. As shown in FIG. 7B, target box 726 is in a different location from target box 724, so as to cause movement of the packaging face 706 and different capture conditions between the multiple images 734. The test image discussed with respect to process 100 can be an image captured without a multi-pose image process, a single one of multiple images in a multi-pose image process (such as a single one of the multiple images 734), or a composite image obtained by combining multiple images in a multipose image process, in various implementations. The use of multi-pose image enables comprehensive analysis across various capture conditions. This approach not only enhances the robustness of authenticity determinations but also provides a more detailed forensic analysis capability.[0821 Insome implementations, the user interface guides the user to capture image(s) of packaging faces that need not be used for authentication, for example, side faces. Even if not used for authentication, these faces can be added to a database to aid in future forensics and analysis.
[0083] Referring again to FIG. 1, in response to output of the one or more machine learning models indicating that the one or more capture conditions are satisfied, and in response to the output of the one or more machine learning models indicating that the face of the first packaging is present in the image, the image is sent for authentication of the first packaging (108). For example, as shown in FIG. 2, the image suitability module 214 can receive outputs from the face identification model 206 and the capture conditionmodel 208. For example, the output from the face identification model 206 can include an indication of whether a packaging face is present in the test image. In some implementations, the output from the face identification model 206 includes a face type of the packaging face present in the test image (e.g., front face or rear face). The output from the capture condition model 208 can include information relating to one or more capture conditions, e.g., a level of corruption in the test image, an orientation of the packaging face in the test image, whether any additional images need to be captured for a multi-pose image set, and / or other information.
[0084] Based on the outputs from the machine learning module 204, the image suitability module 214 can determine whether the conditions for authentication are satisfied, e.g., (i) whether a packaging face is present in the image and (ii) whether one or more capture conditions are satisfied. In some implementations, determining whether the conditions for authentication are satisfied includes comparing one or more of the outputs to a threshold level. For example, if the level of corruption (e.g., glare and / or shadow) in the image is above a threshold, the capture conditions can be determined to be not satisfied; otherwise, the capture conditions can be determined to be satisfied.
[0085] In some implementations, if one or more conditions are not satisfied, the image suitability module 214 can cause the feedback module 210 to output feedback to correct image capture and cause future images to satisfy the conditions. For example, the feedback can include a display stating “no packaging visible, please include packaging in the image,” “there is too much shadow, please increase lighting for image capture,” or “the packaging is currently in an inverted orientation, please face the packaging to the right instead of the left.”
[0086] If the one or more conditions are satisfied (e.g., as set forth for operation 108 in process 100), the image suitability module 214 can trigger the transmission module 216 to send the test image for authentication. For example, the transmission module 216 can include a network interface (e.g., an Internet interface), and the transmission module 216 can send the test image over the network interface to another computer system for authentication. In some implementations, the transmission module 216 sends the test image to a remote computer system, e.g., a cloud computer system. For example, the remote computer system can include elements of system 400 of FIG. 4. In someimplementations, other information is sent in addition to the test image. For example, a face type of the packaging face in the test image (as output by the face identification model 206) can be sent by the transmission module 216, to aid in authentication.
[0087] FIG. 3 shows an example of an image authentication process 300. The process 300 can be performed by a computer system. For example, in some implementations the process 300 is performed by a computer system, such as a cloud computing system, that is remote from a mobile device that captured the test image discussed with respect to FIG. 1. The process 300 can be performed on a test image that was processed according to the process 100 and / or associated operations discussed herein.
[0088] FIG. 4 shows an example of a system 400 associated with process 300. For example, elements of the system 400 can be configured to perform process 300. For example, a reference image selection module 404, an image comparison module 406, and an optional symbol detection module 410 can be configured to perform the process 300 using a test image received from a mobile device 402. The modules 404, 406, 410, and other modules of the system 400, can be hardware and / or software modules, e.g., implemented by one or more computer systems. For example, the modules can include one or more hardware processors and one or more computer-readable mediums encoding instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations such as portions of process 300 and / or other processes described herein. The modules can be software modules executed by one or more hardware processors based on instructions encoded in one or more computer-readable mediums. In some implementations, the modules are modules of a computing system such as a cloud computing system, such that process 300 and associated operations can be performed by a computer system with significant processing resources, so as to deliver authentication determinations rapidly to improve user experience. The modules need not be separate but, rather, can be at least partially integrated together as one or more combined modules or fully integrated together as a single program.
[0089] The system 400 can include one or more computers 401 (e.g., a data processing apparatus) and a non-transitory computer-readable medium 403 tangibly encoding a computer program operable to cause the one or more computers 401 toperform process 100 and / or associated operations described herein. The one or more computers 401 and / or the non-transitory computer-readable medium 403 can implement at least some of the modules of FIG. 4.
[0090] The process 300 includes receiving an image from a mobile device (302). For example, the image can be the test image discussed with respect to FIGS. 1-2. In some implementations, based on the process 100, the test image is received with the presumption that the test image satisfies capture conditions and includes a packaging face, based on the processing performed by the machine learning module 204. In some implementations, a face type of the packaging face is received from the mobile device.
[0091] The process 300 further includes comparing the face of the packaging in the image to faces, in candidate images, having the same face type as the face in the image (304). As noted above, in some implementations the face type of the packaging in the test image is determined by the face identification model 206 and received from the mobile device 402. In some implementations, the face type is not received from the mobile device 402 but, rather, is determined by the system receiving the test image from the mobile device 402. For example, the system receiving the test image from the mobile device 402 can implement the face identification model 206 or another model trained to determine the face type.
[0092] The face type is used to reduce the search space for identifying a reference image for comparison. Because robust image-to-image comparisons to make a final authentication decision may be computationally intensive, it can be beneficial to perform such comparisons only on one or more specifically-identified reference images that are determined to be relevant to the test image, e.g., that are at least somewhat similar to the test image. The identification of one or more reference images can serve as a relatively fast, relatively computationally-efficient first-pass filter for ruling out packaging that is unlikely to match the packaging in the test image.
[0093] As shown in FIG. 4, the reference image selection module 404 includes (or accesses) a candidate image database 408. The candidate image database 408 includes records of authentic packaging for comparison. For example, in some implementations the candidate image database 408 includes, for each of a plurality of candidate packaging, one or more images corresponding to one or more faces of the candidate packaging. Insome implementations, the candidate image database 408 includes, for each of a plurality of candidate packaging, a digital blueprint of the candidate packaging. The digital blueprint can be generated in a process such as process 500 described with respect to FIG. 5. For example, as further discussed in reference to FIG. 5, the digital blueprint for each candidate packaging can include, for each of one or more faces of the candidate packaging, (i) a label of the face type of the face, (ii) an image of the face, and (iii) text (if any) on the face.
[0094] The reference image selection module 404 is configured to compare feature(s) of the test image to corresponding feature(s) of faces of packaging depicted in candidate images in the candidate image database 408, in order to identify one or more reference images that satisfy a threshold level of similarity. In some implementations, the faces of the packaging to which the test image is compared have the same face type as the face in the test image. For example, if the test image depicts a rear face of packaging, then the reference image selection module 404 compares that rear face to rear faces of packaging in the candidate image database 408 (e.g., and not to front faces of packaging in the candidate image database 408). This ensures that only relevant packaging faces are used for comparison, providing significant improvement in comparison speed. For example, in a case where each packaging has two faces depicted in the candidate image database 408, this method can reduce the total search time (for identifying reference image(s)) by 50%, compared to performing comparisons for all faces.
[0095] The comparison between the test image and each candidate image can include (i) a textual comparison between text on the packaging face in the test image and text on the packaging face in the candidate image (e.g., as indicated by a digital blueprint having the text as a data element), (ii) a graphical comparison between the packaging face in the test image and the packaging face in the candidate image, or (iii) both (i) and (ii). The textual comparison can use any suitable text comparison algorithm, such as Jaccard similarity and / or cosine similarity, to provide a textual similarity. The graphical comparison can use any suitable graphical comparison algorithm, such as cosine similarity or Euclidean similarity, to provide a graphical similarity. In some implementations, this graphical comparison is performed using a simple / computational non-intensive graphical comparison algorithm, e g., compared to an algorithm used tomake a final authentication decision as described with respect to operation 308 below. This can reduce the computational cost of identifying the reference image without compromising the final authentication decision.
[0096] Referring again to FIG. 3, the process 300 includes selecting a reference image from the candidate images based on the comparison (306). For example, the reference image selection module 404 can select one or more reference images, where the reference images are candidate images whose similarit(ies) with the test image satisfy a threshold condition, e.g., is above a threshold value. For example, in some implementations, the textual similarity and the graphical similarity are combined into a joint similarity, and the joint similarity is used to select the one or more reference images. In some implementations, the candidate image having the highest similarity is selected as the reference image.
[0097] In some implementations, the selection of a reference image may fail. For example, if the similarities with all searched candidate images are below a threshold value (e.g., the threshold value for selecting a reference image), then it can be determined that no reference image exists in the candidate image database 408. In some implementations, a notification can be provided to a user (e.g., a user of the mobile device 402) responsive to this determination. For example, in some implementations, an indication that the packaging in the test image is inauthentic is displayed (e.g., on the mobile device display 212), under the presumption that, if the packaging were authentic, a matching reference image would have been found. In some implementations, an indication that reference image selection has failed is displayed, e.g., because in some cases the lack of a reference image may be indicative of a missing entry in the candidate image database 408, rather than inauthentic packaging in the test image.
[0098] In some implementations, a data-encoding symbol can optionally be used to identify the reference image. The symbol detection module 410 can determine whether a data-encoding symbol is present in the test image and, if so, reference image identification is performed using the data-encoding symbol. The data-encoding symbol can include, for example, a barcode, a QR (quick response) code, or another onedimensional or two-dimensional symbol that encodes data. The symbol detection module 410 can decode the symbol to obtain the encoded data, and can search for packaginghaving corresponding or matching data. For example, the encoded data can directly indicate a particular product, and, in response, an image of packaging of the particular product (from the candidate image database 408) can be selected as the reference image. This process can speed up selection of the reference image. However, a data-encoding symbol may not be available in all cases. In some implementations, if no data-encoding symbol is found or if the decoded data does not indicate a particular candidate image, another method, such as a textual and / or graphical comparison, can be used to identify the reference image.
[0099] In some implementations, the mobile device 402 sends an indication of whether a data-encoding symbol is included in the test image, e.g., based on a selection by a user of the mobile device 402, as shown in FIG. 8.
[0100] If at least one reference image is identified, the image (e g., the test image) is compared to the reference image to determine the authenticity of the packaging in the image (308). For example, as shown in FIG. 4, the reference image selection module 404 can provide the test image and the one or more reference images to the image comparison module 406. The image comparison module 406 is configured to determine a graphical similarity between the packaging face in the test image and the packaging face in each reference image. In some implementations, this graphical comparison uses a more computationally-intensive method than can be used by the reference image selection module 404, in order to provide more accurate comparison results. For example, the image comparison module 406 can determine one or more similarities using one or more of structural similarity index measure (SSIM), scale invariant feature transformation (SIFT), or Oriented FAST and Rotated BRIEF (ORB). Because the image comparison module 406 performs a relatively small number of comparisons (using the pre-selected reference images), the use of such methods will not consume excessive computational resources or take an excessively long time.
[0101] One or more of various methods can be used to determine the similarity between the test image and the one or more reference images. In some implementations, the similarity is determined as described above in reference to operation 304 and the reference image selection module 404. For example, the similarity can be a graphical similarity and / or a textual similarity as described above.
[0102] Three similarity determination processes will now be described, any one or more of which can be used to determine a similarity between a test image and a reference image, for example, in operation 308. These processes can be performed, for example, by the image comparison module 406, which may be a module of a remote computing system such as a cloud computing system.
[0103] As shown in FIGS. 13A-13B, a process 1300 of determining image similarity can be referred to as a “grid-based model.” In the process 1300, the test image and the reference image are segmented into multiple geometrically-defined regions (1302). The regions are “geometrically-defined” in that the regions need not be defined based on any characteristics of the packaging depicted in the regions. For example, the test image and the reference image can each be segmented into a grid (e.g., an 18x18 grid, which has been found to provide effective results), and each geometrically-defined region can be a square or rectangle of the grid. It will be understood that “grid” here includes various patterns according to which an image can be divided into multiple regions and includes, for example, triangular regions and regions of other shapes, which may be equal in size to one another or may not be. The segmentations of the test image and the reference image can match one another, e.g., such that the same or similar grid is formed for both segmented images.
[0104] Corresponding regions between the test image and the reference image are compared to obtain region-wise similarities (1304). For example, as shown in FIG. 13B, regions 1310a and 1310b, which are in the same position in the segmented grids of the reference image and the test image, respectively, are compared to one another. The comparison can be an image-similarity comparison, for example, based on luminance, contrast, and structure of the pairs of regions, and / or based on a cosine similarity. For example, the similarity can be a combination of one or more of: a structural similarity index measure (SSIM) in grayscale; SSIM in color (e.g., RGB); and cosine similarity in color. When multiple measures contribute to the similarity, the measures can be weighted. In some implementations, the SSIM weight(s) are larger than the cosine similarity weight.
[0105] Based on the region-wise similarities, an authenticity of the packaging in the test image is determined (1306). For example, in some implementations, each region orregion pair (e.g., pair of grid cells) is classified as “authentic” or “suspect” based on whether the corresponding region-wise similarity satisfies a threshold condition. For example, if a similarity (out of a maximum similarity of 1) is 0.6 or less, the region can be classified as “suspect;” otherwise, the region is classified as “authentic.” If a number of “suspect” regions (or cells) is less than a threshold number, the packaging can be determined to be authentic; otherwise, the packaging can be determined to be suspect or inauthentic. In some implementations, the threshold number is fewer than a total number of regions, which can allow the authenticity process to withstand the presence of anomalous grids (e.g., due to corruption such as glare, shadows, etc.).
[0106] Although the grid-based model has been found to provide effective results, in some cases the grid-based model may be less accurate at identifying inauthentic packaging and / or more sensitive to corruption, warping, and / or out-of-plane tilting. Further, in some cases, the grid-based model may be more effective when the reference image is a scanned packaging image, e.g., in .pdf form, compared to when the reference image is an image captured by a mobile device.
[0107] Two other methods may provide improved performance in assessing test images in comparison to mobile device-captured reference images and / or in the context of corruption and the like. As shown in FIGS. 14A-14B, a process 1400 of determining image similarity can be referred to as a “region-based model.” In the case of the regionbased model, regions are defined based on content in the regions, e.g., rather than resulting from a grid-based division as in some implementations of the process 1300. The process 1400 can be performed as, or as part of, operation 308. Region-based models as described herein have been found to provide high levels of accuracy in identifying authentic / inauthentic packaging. Moreover, as described below, region-based models can provide detailed feedback to users as to which element(s) may be inauthentic, facilitating subsequent analysis and adaptation of packaging to make counterfeiting more difficult.
[0108] In some implementations, the process 1400 is performed by a system 1500 shown in FIG. 15. As described in reference to system 400, the system 1500 includes various modules and models that can be implemented by software, hardware, or a combination thereof, and that can be combined with one another and / or distinct, in various implementations. The models and modules of the system 1500 can beimplemented by one or more computing systems, which can have characteristics as described for the system 400. In some implementations, the modules and models of the system 1500 are part of a computing system such as a cloud computing system, which can receive test images from mobile devices over networks (e.g., the Internet). One or more outputs of the system 1500 (e.g., data indicative of differences between test images and reference images) can be sent by the system 1500 to a user device, such as the mobile device 402, and an application of the mobile device 402 can present corresponding user interfaces based on data from the system 1500. The system 1500 can be included in, or can implement, the image comparison module 406.
[0109] The process 1400 includes segmenting a test image 1502 and a reference image 1504 into regions based on respective content in the regions (1402). In some implementations, the reference image 1504 has been pre-segmented, e g., when the reference image 1504 was stored prior to receiving the test image 1502. As such, in some implementations, operation 1402 includes segmenting the test image 1502 and accessing a stored segmentation of the reference image 1504.
[0110] The content-based regions are determined to correspond to particular graphical and / or textual features of the packaging in the test image 1502 and the reference image 1504. For example, in some implementations, the content-based regions correspond to predefined categories / types such as “text,” “variable data,” “product name,” “brand name,” “dosage,” “logo,” and “product-info.” In the example of FIG. 14B, regions 1410a, 1410b are “logo” regions; regions 1412a, 1412b are “dosage” regions; regions 1414a, 1414b, 1420a, 1420b, 1422a, 1422b are “product name,” “brand name,” or “logo” regions; and regions 1416a, 1416b, 1418a, 1418b are “text” or “product-info” regions. It will be understood that these categorizations are provided as examples, and that many different suitable combinations of categories for content-based segmentation can be used based on, for example, product-specific considerations, packaging design, and the like. In some implementations, portions of the packaging that are not assign to a particular defined region are assigned to “remainder” region (including, for example, background elements), so that the entirety of the packaging can be used for authentication analysis.
[0111] Segmentation can be performed automatically by one or more trained region identification machine learning models 1506, and / or manually by users (1508). The trained region identification machine learning models 1506 can be trained using training data including (i) images (e.g., scans and / or captured photographs) of packaging and (ii) labels for regions (e.g., the region types noted above) in the images. As such, the one or more models 1506 can be trained to identify regions of various types in captured images of packaging. The one or more models 1506 can be trained to identify (i) a region type and / or (ii) region boundaries, for each identified region. In some implementations, the training data is generated as described in reference to process 500.
[0112] In some implementations, a distinct model 1506 is trained for each region type. For example, a first region identification machine learning model 1506 can be trained to identify logo region(s), and a second region identification machine learning model 1506 can be trained to identify text region(s).
[0113] Instead, or additionally, the models 1506 can be trained in a product-specific manner. For each product, a corresponding set of one or more region identification machine learning models 1506 is trained to identify types and / or boundaries in the packaging for that particular product. To be product-specific, in some implementations, the models 1506 are specific to a stock keeping unit (SKU) and / or universal product code (UPC). For example, the models 1506 for a given SKU can be trained using labeled training images of packaging corresponding to that SKU. Accordingly, the models 1506 can be trained in a specialized manner to accurately identify region types and boundaries for each product. When region-specific models 1506 are used in operation 1402, the region- specific models 1506 can correspond to a product depicted in the selected reference image 1504. For example, the selected reference image 1504 can depict a particular packaging size and style for a product. The size, style, and product correspond to a specific SKU. One or more region identification machine learning models 1506 that were trained using images of that same SKU packaging (e.g., captured and / or generated images, as described in reference to FIG. 5) are then executed to identify regions in the test image 1502.
[0114] In the case of manual identification (1508), a UI can be presented on the mobile device that captures the test image, the UI including features that allow a user to select region boundaries and region types in the test image.
[0115] Referring again to FIG. 14A, after the test image 1502 (and the reference image 1504, if not already performed) have been segmented into content-based regions, corresponding regions are compared to obtain region-wise similarities (1404). The region-wise comparison may take one or more of various forms, as described with respect to a region-based comparison module 1510 shown in FIG. 15. The region-based comparison module 1510 is configured to compare corresponding regions to determine corresponding similarities. To that end, the region-based comparison module 1510 can perform one or more methods, corresponding to one or more modules of the region-based comparison module 1510, as shown in FIG. 15. The subsequent description of these modules corresponds to processes that can be performed as part of operation 1404.
[0116] The combination of multiple comparison types, corresponding to use of multiple modules of the region-based comparison module 1510, can provide a high degree of accuracy in determining authenticity. Counterfeits that satisfy one or several tests may fail to satisfy all of the tests, such that the use of multiple tests can provide more reliable authentication. For example, a counterfeit that accurately reproduces descriptive labels on packaging may fail to accurately reproduce the positioning and / or morphology of the labels. Based on the analyses set forth below, even subtle differences between reference packaging and suspect packaging can be identified.
[0117] In some implementations, regions having high levels of corruption in the test image 1502 can be ignored in the comparison analysis, e.g., can have regions weights equal to zero if corruption in the region is above a threshold value.
[0118] In some implementations, the region-based comparison module 1510 includes a registration module 1512 that is configured to compare region locations. For a given product, the region locations should match between the test image 1502 and the reference image 1504. For example, the logo should be in the same location on the packaging in the two images 1502, 1504. Deviations in the location may be indicative of inauthentic packaging.
[0119] As an example of analysis performed by the region-based comparison module 1510, a center (e.g., centroid) of each region can be determined. For corresponding pairs of regions in the test image 1502 and the reference image 1504, a distance (e.g., Euclidean distance) between locations of the centers is determined. A percentage deviation is determined based on the distance, and higher deviations correspond to more- suspect packaging in the test image 1502. The deviation can be determined as an absolute value and / or as a proportional deviation.
[0120] FIG. 16 illustrates examples of reference images (for both front and rear face), test images, and registration module 1512 output. Both front and rear faces appear generally similar. However, the registration model analysis indicates that several regions 1602 are shifted down and to the right in packaging in the test image, relative to the packaging in the reference image, as indicated by the arrows and corresponding superimposed numerical values (which indicate an amount of center deviation).
[0121] The registration module 1512 can output region-wise similarities based on the determined distances and / or deviations. In some implementations, the similarities are continuous, e.g., can have various numerical values where higher distance / deviation corresponds to lower similarity. Instead, or in addition, output similarities can be binary, e.g., a “flag” (indicating suspect or inauthentic) for each region if the deviation for the region is greater than a threshold value. In some implementations, if a region is present in one of the images 1502, 1504 and absent in the other, indicating a significant different in packaging, the registration module’s output indicates that difference as a significant dissimilarity.
[0122] In some implementations, the region-based comparison module 1510 includes a graphics morphology module 1514 that is configured to compare graphical structural features between corresponding regions in the test image 1502 and the reference image 1504. The graphical comparison can include, for example, any one or more of the methods described with reference to the image comparison module 406 and / or operation 1304. For example, for each pair of corresponding regions, the graphics morphology module can compute SSIM in grayscale, SSIM in color (e.g., RGB), and / or cosine similarity in color, and determine a region-wise similarity based on those computations (e g., as a weighted average).
[0123] In some implementations, the graphics morphology module 1514 executes a two-step process. In a first step, region-wise SSIMs between the test image 1502 and the reference image 1504 are determined, e.g., using SSIM in color and / or grayscale. If the SSIM for a region is at least a threshold value (e.g., 90%), the SSIM is determined as the region-wise similarity output by the graphics morphology module 1514. If the SSIM for a region is less than the threshold value, pixel-wise intensity differences for the region, between the test image 1502 and the reference image 1504, are determined. In that case, an aggregate measure of the pixel-wise intensity differences is output as the region-wise similarity by the graphics morphology module 1514. This two-tiered or layered approach that incorporates SSIM in conjunction with pixel intensity differences has been found to provide improved accuracy in identifying suspect regions, e.g., compared to the use of only SSIM. In some implementations, the pixel-wise intensity differences are determined based on grayscale versions of the images 1502, 1504. In some implementations, the region-wise similarity is, or includes, a binary indicator of authentic or anomalous / suspect based on the pixel-wise intensity differences. For example, a region can be identified as suspect by the graphics morphology module if a set of n consecutive or adjacent pixels has intensity differences above a threshold, where n can be a tunable parameter.
[0124] In some implementations, a user interface on the mobile device 402 presents an output of the graphics morphology module 1514. For example, the user interface can present a graphical representation of graphical morphology differences. For example, the user interface can indicate portions of regions in the test image 1502 for which there is a notable graphical morphological difference with corresponding regions in the reference image 1504.
[0125] In some implementations, the region-based comparison module 1510 includes a text morphology module 1516 that is configured to compare text graphics between corresponding regions in the test image 1502 and the reference image 1504. The comparison of text graphics can be based on shapes, colors, and / or the like of the text, as opposed to the content of the text. The text morphology module 1516 outputs one or more region-wise similarities indicating similarities between the morphology of text in corresponding pairs of regions.
[0126] In some implementations, the text morphology module 1516 uses a contourbased technique to detect text in the test image 1502 and the reference image 1504. The text in one of the images 1502, 1504 is then projected onto the other image 1502, 1504. The projection (or, e.g., overlay) can be for text in the images 1502, 1504 as a whole, or for specific regions, e.g., text in a region of the test image 1502 can be projected onto the corresponding region in the reference image 1504. Moreover, comparison methods besides projection are within the scope of this disclosure.
[0127] Based on the detection of the text, one or more graphical similarity metrics are applied to graphically / morphologically compare the text between the images 1502, 1504. For example, in some implementations, a combination (e.g., a weighted combination) of one or more of grayscale SSIM, color SSIM, and cosine similarity is applied. In some implementations, if the resulting similarity for a region is below a threshold value, the region is flagged as anomalous / suspect, and an indicator of the flagging is output by the text morphology module 1516. Alternatively, or in addition, one or more calculated region-wise similarities can be output.
[0128] FIG. 17 illustrates examples of a logo region 1702 in a reference image, a logo region 1704 in a test image, and an overlay 1706 of text in the regions 1702, 1704. Boxes in the overlay 1706, such as boxes 1708, indicate the presence of substantial morphological differences. In some implementations, a user interface presented on the mobile device 402 includes presentation similar to that of FIG. 17, e.g., can present side- by-side images of corresponding regions from the images 1502, 1504, along with a graphical indicator of morphological differences therebetween.
[0129] In some implementations, the region-based comparison module 1510 includes a text accuracy module 1518 that is configured to compare text context between the test image 1502 and the reference image 1504. For example, the region-based comparison module 1510 can determine similarities between the letters and / or words in corresponding regions of the images 1502, 1504.
[0130] In some implementations, the text accuracy module 1518 extracts text from identified regions in the test image 1502 and the reference image 1504. For example, the text accuracy module 1518 can execute an optical character recognition (OCR) engine. The text accuracy module 1518 can then perform one or more suitable text comparisonmethods, such as determining a Jaccard similarity, cosine similarity, and / or Levenshtein distance between extracted text from corresponding regions. Instead, or in addition, the text accuracy module 1518 can identify insertions and / or deletions between corresponding text, and determine region-wise similarities based on the insertions / deletions. The text accuracy module 1518 can output the similarities and / or a derivative thereof.
[0131] In some implementations, a user interface on the mobile device 402 can indicate identified insertions and / or deletions. FIG. 18 illustrates an example of a such a user interface 1800. The user interface 1800 includes the test and reference images (or portions thereof), in which text regions 1808 are identified visually by boundary boxes. As noted above, the text regions 1808 can be identified by one or more trained region identification models 1506. The user interface 1800 further includes indicators 1806 (e g., boundary boxes) of compared text between the text regions 1808. The indicated text is shown in a user interface region 1802. A set of user interface elements 1804 indicates text that has been added / deleted in the test image 1502 compared to the reference image 1504, and controls among the user interface elements 1804 allow a user to select which added / deleted text should be indicated in the right side of the user interface 1800. For example, the indicators 1806 can be a first color to indicate text insertions and a second color to indicate text deletions.
[0132] In some implementations, the region-based comparison module 1510 includes a color module 1520 configured to compare color between corresponding regions in the test image 1502 and the reference image 1504, to determine one or more similarities of the colors. The one or more similarities, or a derivative thereof, can be output by the color module 1520. In some implementations, prior to comparing colors, the color module 1520 white-balances the test image 1502 so that the test image 1502 is more similar to the reference image 1504. In some implementations, a heat map of color differences in one or more regions is displayed on the mobile device 402, for review by users. One or more suitable color comparison methods can be used. For example, a distance (e.g., Euclidean or Manhattan) in a suitable color space, such as RGB and / or LAB, can be computed between the regions, where smaller distances correspond to higher similarities.
[0133] In some implementations, the region-based comparison module 1510 includes a variable data module 1522 that is configured to verify variable data in the test image 1502 and / or compare variable data between the test image 1502 and the reference image 1504. Variable data, in the context of packaging, refers to packaging features (e.g., text, images, graphics, and the like) that are printed onto packaging materials in a personalized and / or customized manner. Variable data can include product information, pricing, barcodes, logos, and / or other design elements that are printed using digital technology onto pre-printed packaging substrates. Variable data can be are specific to each piece of packaging. For example, every box of an over-the-counter drug may be printed with a unique identifier. Counterfeiters may fail to accurately include variable data, e.g., by failing to follow industry standard rules that dictate form / content of variable data. Counterfeiters may also fail to accurately match symbols indicating variable data with text indicating variable data.
[0134] As such, in some implementations, the variable data module 1522 is configured to identify variable data in the test image 1502 and (i) verify the authenticity of the variable data and / or (ii) compare the variable data between the test image 1502 and the reference image 1504. As an example of a process performed by the variable data module 1522, with reference to FIG. 19, the variable data module 1522 can identify a data-encoding symbol 1902 (e.g., a QR code) in a test image. The data-encoding symbol 1902 can be identified using any suitable machine learning and / or computer vision method, e.g., a Viola-Jones method. The variable data module 1522 can decode encoded text 1904 from the data-encoding symbol 1902. The variable data module 1522 can also identify and extract text 1906, e.g., a serial number (SN), an expiration date, a lot number, and / or a global trade item number (GTIN). The text 1906 can be extracted, for example, using any suitable OCR method. The variable data module 1522 can then compare the extracted text 1906 with the decoded text 1904 and identify a similarity therebetween. Differences between the extracted text 1906 and the decoded text 1904 can be indicative of suspect / counterfeit packaging in the test image 1502.
[0135] Instead, or in addition, the extracted text 1906 and / or the decoded text 1904 can be compared to extracted and / or decoded text in a corresponding region in the reference image 1504. While the text itself, as variable data, may be different, in someimplementations, a format or structure of the text can be compared to determine a similarity. Variable data analysis involves decoding data-encoding symbols like QR codes and comparing them with OCR-extracted text from packaging. Discrepancies between decoded data and extracted text are indicative of counterfeit packaging. It will be understood that this comparison is optional and that, in some implementations, operations of the variable data module 1522 are limited to analysis of the test image 1502, without necessarily performing comparison with the reference image 1504.
[0136] Referring again to FIG. 15, based on the operations of one or more of the modules of the region-based comparison module 1510, the region-based comparison module 1510 outputs one or more similarities 1524 that are indicative of an authenticity of the test image 1502. The one or more similarities 1524 can be individual output(s) of module(s) of the region-based comparison module 1510, and / or can include an aggregate measure thereof, e g., a weighted average of similarities output by multiple of the modules of the region-based comparison module 1510. In some implementations, the one or more similarities 1524 is a lowest similarity output by the modules of the region-based comparison module 1510, on the basis that the “worst-case” comparison should be taken as indicative, even if other comparisons indicate a high degree of similarity. Other forms of output of the region-based comparison module 1510, based on the above-described operations of the modules thereof, are also within the scope of this disclosure.
[0137] As noted above, the process 1300 of FIG. 13A corresponds to a “grid-based model,” and the process 1400 of FIG. 14A corresponds to a “region-based model.” The processes 1300 and / or 1400 can be performed in operation 308 to determine the authenticity of packaging. In some implementations, instead, or in addition, a “suspect packaging” model is applied. The “suspect packaging model” is based on one or more machine learning models that have been trained on examples of suspect / inauthentic packaging.
[0138] As shown in FIG. 20A, a process 2000 for performing authentication includes training a suspect packaging model (one or more machine learning models) based on annotated images of suspect / inauthentic packaging (2002). The suspect packaging model can be an object detection model. The images can be captured and annotated as described with respect to FIG. 21. For example, the annotations can include indications of particularregion(s) in the images that are thought or known to include suspect features such as incorrect graphics, inconsistent text, and / or the like. In some implementations, the suspect packaging model is trained specific to a particular product / packaging (e.g., SKU) or class of products / packaging, by being provided with training data depicting inauthentic examples of that particular product / packaging or class of products / packaging. In some implementations, the suspect packaging model is trained agnostic to products.
[0139] In some implementations, the suspect packaging model need not be (e.g., is not) trained on human-annotated images of authentic packaging. The process 2000 can include projecting annotated suspect regions onto images of authentic packaging, as authentic regions (2004), thereby obtaining training data of authentic regions (authentic text, authentic logos, etc.) without requiring user annotation. This can streamline the model training process. For example, as shown in FIG. 20B, a logo region 2014 (indicated by a boundary box around the logo “SecureSample”), annotated and labeled by a user in an inauthentic image 2010, can be projected onto an authentic image 2012 as an authentic logo region 2016. The projection can be based on, for example, sizing the images 2010, 2012 to the same size, so that corresponding regions can be mapped between them. The authentic images annotated with authentic regions can then be used as training data in operation 2002, alongside the inauthentic images annotated with inauthentic regions.
[0140] Based on the training data (e.g., annotated inauthentic images and, in some implementations, authentic images annotated based on the annotated inauthentic images), the suspect packaging model is trained to identify suspect regions in newly -presented images such as the test image 1502. The authentication process 2000 can include providing the test image 1502 to the suspect packaging model to obtain an authenticity result (2006). For example, if the suspect packaging model identifies any inauthentic regions in the test image 1502, the test image 1502 can be determined to be inauthentic; otherwise, the test image 1502 can be determined to be authentic.
[0141] In some implementations, to provide training data for the suspect packaging model, training data of inauthentic packaging can be provided. Users can upload images of inauthentic packaging and label suspect element(s) of the packaging, for use as training data to assist the suspect packaging model in learning to identify suspectfeatures. For example, a user device can be provided with a user interface like user interface 2100 shown in FIG. 21. The user interface 2100 includes a presentation of an image 2102 of inauthentic packaging, and user interface elements 2106 usable to identify particular regions that include suspect graphics, text, and / or the like, such as indicated region 2104. The image 2102 can be captured by a mobile device, and the same mobile device can be provided with the user interface 2100. Labeled images of inauthentic packaging can be used as training data. In some implementations, the training data is augmented by generated images, as described in reference to operation 510 of FIG. 5. In some implementations, the upload and annotation of images of inauthentic packaging can be performed as described in reference to FIG. 5, with operation 514 then including the training of the suspect packaging model. For example, a user can upload an image; label the packaging in the image as “authentic” or “suspect”; and identify regions in the image (e g., text, logo, etc.), optionally with one or more regions additionally identified as including suspect elements.
[0142] It has been found that, in some cases, the suspect packaging model provides superior performance to the region-based model, which provides superior performance to the grid-based model. Accordingly, in some implementations, the suspect packaging model is executed as a first choice, if a trained suspect packaging model is available. As a second choice, the region-based model is executed, if region data (e.g., training data labeled with regions, region identification model(s) 1506, reference images labeled with regions, and / or the like) is available. If neither the suspect packaging model nor the region-based model can be used, the grid-based model is used. In some implementations, the method for selecting authentication models prioritizes Al-based models over Region- Based Models (RBM) and Math-Models due to their higher accuracy in determining authenticity. This selection process dynamically adapts based on model availability and computational efficiency.
[0143] Based on the similarity or similarities as determined by the image comparison module 406 (for example, by one or more of the grid-based model, the region-based model, or the suspect packaging model), the image comparison module 406 determines an authenticity of the packaging in the test image. In process 1300 shown in FIG. 13A, this determination corresponds to operation 1306. In process 1400 shown in FIG. 14A,this determination corresponds to operation 1406. For example, if a similarity between the test image and any of the reference images is above a threshold value, it can be determined that the packaging is authentic. As another example, if a region-wise similarity between a region of the test image and a most-similar corresponding region of a reference image is below a threshold value, it can be determined that the packaging is inauthentic. In some implementations, the image comparison module 406 can classify the packaging into a category based on pre-defined similarity thresholds, e.g., “suspect / inauthentic,” “needs further review,” and “genuine,” each of those categories corresponding to an increasing similarity. In some implementations, if multiple reference images have been identified by the reference image selection module 404, the determination can be made based on the highest-similarity reference image (as determined by the image comparison module 406).
[0144] In process 2000 shown in FIG. 20A, there need not be an explicit comparison to a reference image. Rather, for example, the trained suspect packaging model, provided with a test image as input, can output a determination indicative of suspect or authentic packaging based on its previous training, without having to perform a separate, explicit comparison step. As such, in the case of the process 2000, the reference image selection module 404 can be omitted, and the image comparison module 406 can execute the suspect packaging model to determine whether the test image is authentic or inauthentic. As another example, the reference image selection module 404 can be used to determine a product or packaging corresponding to the test image (e.g., by identifying a reference image and determining the product / packaging of the reference image as the product corresponding to the test image). Then, a trained suspect packaging model specific to that product / packaging can be applied by the image comparison module 406, using the test image as input.
[0145] Correspondingly, in implementations of the process 300 in which the trained suspect packaging model is used, operation 308 can include “provide image as input to trained suspect packaging model, where the trained suspect packaging model has been trained specific to product / packaging depicted in the reference image selected in operation 306.”
[0146] In some implementations, an output is provided to a user based on the determination by the image comparison module 406. For example, if the graphical similarity is above a threshold, the mobile device 402 can display a notification that “the product is authentic.” If the graphical similarity is below a threshold, the mobile device 402 can display a notification that “the product is inauthentic” or “the product may be inauthentic.” The notification can be displayed based on data sent from the image comparison module 406 to the mobile device 402, e.g., from the remote computing system performing process 300 to the mobile device 402 over one or more networks, such as the Internet.
[0147] In some implementations, in at least some cases in which the packaging in the test image is inauthentic, the image comparison module 406 can be configured to provide an identity of packaging / product of which the packaging in the test image is a counterfeit. For example, if the graphical similarity with a reference image (as determined by the image comparison module 406) is below a first threshold corresponding to authentic packaging but above a second threshold, it can be determined that the packaging in the test image is a counterfeit of the packaging in the reference image, because the two packaging are somewhat similar without being exact matches for one another. In some implementations, the identity of the packaging / product of which the packaging in the test image is a counterfeit is the identity of the packaging / product in the single identified reference image or the packaging / product in the reference image having the highest similarity from among multiple identified reference images.
[0148] In some implementations, prior to performing image comparison, the reference image selection module 404 and / or the image comparison module 406 aligns the test image with the reference image, e.g., causes the packaging face in the test image to have the same orientation as the packaging face in the reference image (e.g., by rotation of one or both images). This can ensure that comparisons are performed on aligned images to obtain accurate results. To perform alignment, an Oriented FAST and Rotated BRIEF (ORB) method can be applied, in which parameters that maximize the number of key point-features across the test image and reference image are selected. In some implementations, the reference image selection module 404 and / or the image comparison module 406 removes whitespace from the images.
[0149] In some implementations, comparison accuracy is aided by the use of multiple test images. For example, test images corresponding to two or more different face types can be used. The mobile device 402 can send a first test image showing a first face type (e.g., front face) and a second test image showing a second face type (e.g., rear face). The reference image selection module can select “reference packaging” based on one or both of the test images, where the candidate image database 408 stores multiple images of the reference packaging, respective ones of the multiple images depicting to the first face type and the second face type.
[0150] For example, the reference image selection module 404 can use the first test image to identify a first reference image that depicts a first face of particular packaging, the first face having the first face type (e.g., front face). The candidate image database 408, in addition to the first reference image, also includes a second reference image that shows a different face (e.g., rear face) of the same particular packaging, the different face having the second face type. Both of the reference images can be selected by the reference image selection module 404 and provided to the image comparison module 406. The image comparison module 406 can then compare the first test image to the first reference image and the second test image to the second reference image to obtain two similarities, both of which can be used to determine an authenticity of the packaging in the test images. For example, the two similarities can be averaged or otherwise combined, and the combined similarity can be compared to a threshold value to determine the authenticity. As a result, authenticity determination can be performed more reliably, e g., to identify counterfeits that may have one accurate face and one inaccurate face.
[0151] In some implementations, capture of multiple faces is used in response to an indeterminate authenticity determination, e.g., the “needs further review” determination discussed above. For example, in response to such a determination by the image comparison module 406 based on a first face of packaging, a user can be instructed (e.g., by instructions provided on the mobile device display 212 by the feedback module 210) to capture an image of a second face of the packaging. The image of the second face can then be processed to determine a graphical similarity between the image of the second face and a reference image of the second face. In some such cases, authenticity may bedetermined based on the second face even if the graphical similarity for the first face was such that authenticity could not be determined.
[0152] As another example of the use of multiple test images, the mobile device 402 can provide multiple test images to the reference image selection module 404, the multiple test images depicting the same packaging face but with different packaging positionings, capture conditions, etc. For example, the multiple test images can be images of a multi-pose set captured as described with respect to FIGS. 1-2. Each of the multiple test images can then be compared to the reference image by the image comparison module 406, and a mean similarity (or other aggregate / joint similarity) can be used to determine the authenticity of the packaging in the test image. By using multiple images of the same face, the net effect of environmental noise such as glare, shadow, mechanical / surface scratches on the packaging, and / or other forms of minor degradation can be reduced.
[0153] In some implementations, both multi-face-type comparisons and multi-pose set comparisons are performed, and the various similarities that result are combined into an aggregate similarity measure that the image comparison module 406 uses to assess the authenticity of the packaging in the test images.
[0154] In some implementations, one or more of the grid-based, region-based, and suspect packaging models is usable in a customized manner, e.g., to provide different levels of detail and / or different processing speeds. For example, operations can be divided into a “fast response” routine and a “forensic” routine. In a fast response routine, one or more of the three models is used to quickly determine whether a test image depicts inauthentic packaging, and / or to quickly identify particular packaging features that are suspect. For example, the fast response routine may implement the grid-based model as a rough determination of whether a test image depicts authentic packaging, and / or may implement only a subset of available modules of the region-based comparison module 1510 in the case of the region-based model or the suspect packaging model. For example, the fast response routine may implement only the graphics morphology module 1514. The fast response routine may provide some information about particular likely- inauthentic elements, e.g., may flag one or more regions as inauthentic in a suitable userinterface. The fast response routine may be able to provide results quickly (e g., in about thirty seconds) to permit rapid testing of potentially-suspect packaging.
[0155] The forensic routine may implement more computationally-intensive processes, more processing modules, and / or the like. For example, FIG. 22 illustrates an example of a user interface 2200 associated with the forensic routine. The user interface 2200 includes user interface elements 2202 by which a user can select which module(s) (referred to in FIG. 22 as “labels”) of the region-based comparison module 1510 will be utilized. Further, the user interface elements 2202 can allow a user to switch between user interfaces / presentations corresponding to the different modules. For example, user interface 2200 can first present an interface showing differences in text content based on output by the text accuracy module 1518, and can subsequently, based on selection by the user, present an interface showing differences in color based on output by the color module 1520. In the configuration shown in FIG. 22, the user interface 2200 includes presentation of a reference image 2204 and a corresponding test image 2206 in which overlays 2208 (e.g., boundary boxes, shading, and / or the like) indicate elements with notable differences compared to the reference image 2204. A user can use the user interface 2200 for detailed analyses of packaging in order to detect and counter counterfeiting.
[0156] Referring again to FIG. 4, in some implementations, the system 400 includes an overlay detection module 420. The overlay detection module 420 is configured to detect overlays in provided test images. Overlays can include, for example, holograms, price tags, stickers, and / or the like, which may be provided on packaging for rebranding in certain regions, by retailers, etc. These overlays may not be included in reference images, and, therefore, may result in false determination of inauthentic packaging if not accounted-for. The overlay detection module 420 can include a machine learning model trained to detect overlays and exclude them from authentication analyses, e.g., by the image comparison module 406. The machine learning model can be trained based on labeled images of packaging with overlays, the overlays labeled as, for example, “mandatory” or “optional.” In some implementations, overlays are excluded from the analysis based on being identified as optional by the machine learning model of the overlay detection module 420. Mandatory overlays (which should be present on allauthentic packaging) can be included in the analysis even if identified as overlays by the machine learning model.
[0157] FIG. 8 illustrates examples of user interfaces 802, 804, 806 that can be displayed during performance of the processes 100, 300. The user interfaces 802, 804, 806 can be displayed, for example, on the mobile device display 212, e.g., of the mobile device 402. User interface 802 provides controls allowing a user to select between scanning a serial number code (e.g., a QR code), capturing a test image of a front of packaging, or scanning a barcode (e.g., a one-dimensional barcode). User interface 804 shows a live image of capture by the mobile device camera 202, in this case for scanning a QR code 808. An interface element 810 triggers the sending of the QR code 808 (or data encoded by the QR code 808), along with a test image 812 of packaging, for authentication. User interface 806 can be displayed while a remote system performs authentication, e g., performs process 300. The remote system uses the QR code 808 to identify one or more reference images, as discussed with respect to the symbol detection module 410. At the conclusion of authentication, another user interface (not shown) can display an authentication result.
[0158] FIGS. 12A-12E illustrate further examples of user interfaces that can be displayed during performance of the processes 100, 300. The illustrated user interfaces can be displayed, for example, on the mobile device display 212, e.g., of the mobile device 402. In some implementations, the user interfaces of FIGS. 12A-12E are displayed during, or in association with, image capture in operation 102.
[0159] As shown in FIG. 12A, user interface 1200 indicates that a captured image of packaging is warped (e.g., due to improper handling of the packaging by a user) and requests that the image be recaptured. This indication can be based on, for example, an output of the capture condition model 208, which can execute on the mobile device that has captured the image and is displaying the user interface 1200. A bounding box or frame (shown in FIGS. 12A-12D as four superimposed comers) can indicate a detected face of the packaging. As shown in FIG. 12B, user interface 1210 indicates that an image has been captured successfully, e.g., satisfying requirements for the capture conditions. User interface 1210 further indicates that the image depicts a packaging front face.
[0160] As shown in FIG. 12C, user interface 1220 indicates detection of an invalid face. For example, the face identification model 206 can, in real time, determine that the captured image shown in FIG. 12C depicts the front face (which has already been captured), whereas a rear face has been requested. The user interface 1220 includes a request that the packaging be rotated to permit imaging of the rear face. As shown in FIG. 12D, user interface 1230 indicates that a captured image of the packaging (in this case, the rear face) includes light corruption, e.g., glare, based on an output of the capture condition model 208. As shown in FIG. 12E, when suitable image(s) have been captured (e.g., both front and rear face images), user interface 1240 can indicate that the image is being sent to a remote computing system for further processing, e.g., for authentication using process 300.
[0161] FIG. 5 illustrates a process 500 for (i) generating digital blueprints for comparison to identify reference images and perform packaging authentication, and (ii) training machine learning models described herein, such as the face identification model 206 and the capture condition model 208. The process 500 can be performed, for example, by a mobile device or local computer system, by a remote computer system such as a cloud computer system, or by both, as discussed in further detail below. For example, image capture and / or upload, and manual annotation, can be performed using a mobile device or local computer system, while auto-annotation, model training, and / or digital blueprint generation can be performed using a remote computer system.
[0162] The process 500 includes obtaining an image of packaging (502). While processes 100 and 300 related to testing the authenticity of packaging, the packaging discussed with respect to FIG. 5 is known to be authentic. Accordingly, information about the packaging can be added to the candidate image database 408 for use in subsequent authentications. Obtaining the image of the packaging can include, for example, obtaining an uploaded image of the packaging or capturing an image of the packaging. The image can be captured, for example, by a camera of a mobile device.
[0163] In some implementations, the image is checked for duplication (503) and / or for properly showing packaging (505). In de-duplication, the image can be compared to previously-obtained (e.g., uploaded) packaging images using any suitable file comparison method, e.g., an MD5 hash comparison. For example, the image can be compared topackaging images in the candidate image database 408. If the image is found to be a duplicate, the image can be rejected and the process 500 terminated.
[0164] Further, in some implementations, the image is checked to determine whether the image includes an acceptable template that qualifies as packaging graphics. A template can be artwork content that is common across a product / packaging or group of products / packaging, e.g., an SKU family or group. If the image is not found to include such a template, the image can be rejected and the process 500 terminated. This latter check can reject extraneous, non-packaging images from being used for training.
[0165] In some implementations, the process 500 includes automatically correcting the orientation of the packaging for standardization across all uploaded and stored images. For example, the orientation detection model discussed with respect to FIG. 4 can be applied to determine whether the image portrays tilted or inverted packaging. If so, the process 500 can include adjusting the orientation (e.g., by varying a tilt angle of the obtained image or a portion thereof) until the orientation is determined to be correct.
[0166] The image is processed using a machine learning model (sometimes referred to as an “auto-annotation machine learning model”) trained to determine a face type of a face of the packaging (504). For example, the auto-annotation machine learning model can be trained to identify whether the image portrays a front face or a rear face of the packaging. In some implementations, the auto-annotation machine learning model has been trained to extract text from the packaging in the image for auto-annotation purposes. In some implementations, the auto-annotation machine learning model is an object detection model trained on a set of packaging images to delineate and identify different faces such as front and rear faces, the packaging images labeled with (i) locations of the packaging faces in the images and (ii) an indicator of face type for each packaging face. The auto-annotation machine learning model can output an auto-annotation that includes a face type of the packaging face in the image, a location of the packaging face in the image, and / or text included in the packaging face in the image.
[0167] In some implementations, the obtained image (see operation 502) portrays multiple packaging faces, e.g., when the image is an uploaded image that includes an “unfolded” packaging, for example, as shown in FIG. 9A. As such, processing by the auto-annotation machine learning model can include the identification of multiple faces,determination of the face types of the multiple faces, and auto-annotation of the multiple faces.
[0168] The auto-annotation machine learning model used in operation 504 of the process 500 can be different from the machine learning model(s) of the machine learning module 204. For example, in some implementations, the machine learning module 204 is a module of a mobile device, while the auto-annotation machine learning model is implemented in a module of a remote computer system, e.g., a server. For example, after capture or upload of the image at a mobile device (502), the mobile device can send the image to the remote computer system for further processing (e.g., auto-annotation), including processing by the auto-annotation machine learning model (504).
[0169] FIGS. 9A-9C illustrate examples of user interfaces 902, 904, 920 associated with the process 500, e g., that can be displayed to a user of a mobile device or other computing device during performance of the process 500. As shown in FIG. 9A, user interface 902 includes an element 905 used to select an image of product packaging (502). In this example, a PDF fde containing an image 906 is selected. The image 906 includes multiple packaging faces, such as a front face 908 and a rear face 910. The image 906 can be sent from the mobile device or other computing device to a remote system for processing, e.g., for auto-annotation (504)
[0170] User interface 904 illustrates at least partial results of auto-annotation. The auto-annotation machine learning model has identified and extracted images of each of the front face 908 and the rear face 910. In addition, the auto-annotation model has performed text recognition and extraction on the faces 908, 910, as indicated by an “OCR Extraction” progress bar 912 that refers to optical character recognition (OCR).
[0171] In some implementations, instead of, or in addition to, operation 504 in which face types are determined, the process 500 includes processing the image using an autoannotation region identification model (507). The auto-annotation machine learning model (which can include one or more distinct machine learning models) is trained to identify region types and region boundaries in the obtained image, for example, “logo” region(s) and / or “product name” region(s), as discussed in reference to the ’’region-based model” for FIGS. 14A-14B and 15. The auto-annotation region identification model of operation 507 can output an auto-annotation that includes boundaries of one or moreregions in the image, and corresponding region types of the regions. The auto-annotation region identification model can be the same as, or different from, the one or more region identification machine learning models 1506.
[0172] Referring again to FIG. 5, in some implementations, a user of the mobile device (e.g., the mobile device that provides the image in operation 502) is provided with an interface usable to manually alter the auto-annotation and / or to manually annotate (506), e.g., to correct any errors in the auto-annotation. The manual annotation can include identification of packaging faces and face types, entry of text on the packaging faces, selection of locations of the packaging faces, and / or other annotation operations. For example, FIG. 9C illustrates a user interface 920 for manual annotation. User interface 920 includes, among other features, the uploaded / captured image 906, a selection of bounding shapes 922 selectable by a user to bound packaging faces in the image 906, and an attributes menu 924 usable to label face(s) portrayed in the image 906, e.g., to label with a face type and / or with text included on the face. In this case, autoannotated bounding boxes 928, 930 bound the front and rear faces 908, 910. The user can adjust the bounding boxes 928, 930 if the auto-annotation machine learning model has incorrectly identified the bounds of the faces 908, 910.
[0173] FIG. 11 illustrates another example of a user interface 1100 configured for manual annotation (e.g., in operation 504). The user interface 1100 includes a presentation of an uploaded packaging image 1102; controls 1104 usable by a user to label front / rear faces and input information about the text orientation on those faces; and a list of selectable annotation templates 1106 that can be used to speed-up manual annotation. The annotation templates 1106 can correspond, for example, to previously- annotated packaging and / or to standard packaging layouts for a given product or supplier. When an annotation template 1106 is selected, the faces of the annotation template 1106 (e.g., a selection of area(s) with corresponding front / rear label(s), in some cases with corresponding text orientation labels) are populated to the current packaging image 1102. The user can make further adjustments from the template information, if desired.
[0174] In some implementations, the user interface 1100 includes controls usable to label regions in the packaging image 1102. For example, the controls can be usable by a user to draw boundaries (e.g., boundary boxes) around regions and label the regions withregion-types, as discussed below for the “region-based model” in reference to FIGS. 14A-14B and 15. In some implementations, the controls enable a user to modify autoannotated regions output by operation 507, e.g., to correct errors in the auto-annotation. The labeled regions can then be used to train region identification machine learning model(s) 1506.
[0175] As a result of machine learning-based auto-annotation (504) with optional manual annotation (506), a digital blueprint is generated (508). The digital blueprint represents attributes of the packaging for subsequent comparison during authentication testing, e.g., during process 300. For example, the digital blueprint can be included in the candidate image database 408 for use in selecting a reference image.
[0176] An example of a digital blueprint 512 is shown in FIG. 5. The digital blueprint 512 is a data object representing attributes of one or more faces of the packaging. In this example the digital blueprint 512 includes, for two faces of the packaging, an indicator of the face type (“front_face” or “rear_face”), an image of the face (e.g., based on autoannotated and / or manually annotated bounding boxes, in this case saved in .jpg form), and text included on the face (in the example of FIG. 5, a table cell that includes text of the front face has been selected and so obscures text extracted from the rear face). The digital blueprint 512 can be used, for example, to determine a textual similarity and / or a graphical similarity during reference image selection, and / or to determine a graphical similarity during final packaging authentication.
[0177] In some implementations, the digital blueprint 512 includes region data, e.g., region boundaries and region types, as determined in operations 506 and / or 507.
[0178] In some implementations, the image of each face (e.g., in the digital blueprint 512) is rotated (if necessary) to be in a particular orientation, e.g., corresponding to a “correct” orientation.
[0179] In some implementations, the packaging annotations are used for machine learning model training. This can allow users without technical machine learning experience to train authentication models (such as the face identification model 206, the capture condition model 208, the region identification model 1506, and / or the autoannotation machine learning model and / or auto-annotation region identification modelthemselves) through user interfaces, e.g., without having to perform coding or other technical tasks.
[0180] In some implementations, model training is performed based on an augmented dataset of packaging images. For example, as shown in FIG. 5, the process 500 can include generating further images differing in one or more parameters (510). The further images can be based on image(s) extracted in the auto- and / or manual annotation 504, 506, and the parameters can include one or more of background, orientation, image contrast, a noise level (e.g., by adding Gaussian noise), a blur level (e.g., by adding simulated blur, such as motion blur), and / or corruption. The backgrounds can include backgrounds of various colors, textures, and / or images, e.g., to emulate typical environments in which users are likely to capture test images (e.g., a factory floor, an office, etc ). The generated images can vary from one another in their background, orientation, image contrast, and / or corruption so as to provide a diverse set of training data.
[0181] Because the further images are generated, each further image has known ground-truth label(s) (e.g., labeling the orientation of the packaging face, the location of the packaging face in the image, locations and types of regions in the image, corruption in the image, etc.), and the labeled images can be used as training data for one or more of the machine learning models described herein. For example, generated images having various different capture orientations can be used to train the orientation detection model; generated images in which the packaging is warped and / or out-of-plane can be used to train the warp / out-of-plane model; and generated images having different levels and / or types of corruption (which, as noted above, can be artificial corruption provided during the image generation process) can be used to train the corruption detection model. Further, the generated images can be used to train the face identification model 206 and the region identification model 1506 so that the models 206, 1506 can effectively identify packaging faces for a wide range of capture conditions.
[0182] For example, FIG. 10 illustrates examples of images generated based on the extracted front face 908 of the packaging shown in FIGS. 9A-9C. The images differ from one another in background (e.g., backgrounds 1002, 1004) and orientation (e.g., labeled with “correct,” “inverted,” and “tiltfed] Implicit in FIG. 10 is that each image is alsolabeled with a location of the packaging face in the image. Moreover, in some implementations each image is labeled with the face type of the face, known based on the prior annotation. Not shown in FIG. 10, but within the scope of this disclosure, is the addition of simulated / generated corruption, such as shadow and / or glare, in the generated images.
[0183] Accordingly, using the generated images and their labels as training data, an object detection model (such as the face identification model 206, the capture condition model 208, the region identification model 1506, the auto-annotation machine learning model, and / or the auto-annotation region identification model) can be trained (514) to identify packaging faces in images (including determining the locations of the packaging faces), determine the face types of the faces, identify boundaries and types of regions in images, and / or determine capture condition information such as orientation and glare level / location. This training can be performed without requiring arduous capture of many images of the packaging from many different orientations and in different capture conditions. Rather, in some implementations, a single image of a face of packaging can be used to generate many training images, improving model accuracy and reliability while improving the user experience. Model training or re-training can be performed easily through a graphical user interface, such as the interfaces 902, 904, 920 of FIGS. 9A-9C. For example, a user can select button 932 in FIG. 9C to trigger (i) generation of additional images based on the annotations of FIG. 9C and (ii) training based on the additional images, so that the machine learning model(s) will be retrained with a new example of packaging. This user-friendly approach to annotation and model training can make it more likely for users to participate (e.g., to upload packaging), increasing the dataset available for training and improving model accuracy and / or reliability.
[0184] Types of machine learning models within the scope of this disclosure include, for example, machine learning models that implement supervised, semi-supervised, unsupervised and / or reinforcement learning; neural networks, including deep neural networks, autoencoders, convolution neural networks, multi-layer perceptron networks, and recurrent neural networks; classification models; large language models (LLM); and regression models. The machine learning models described herein can be configured with one or more approaches, such as back-propagation, gradient boosted trees, decision trees,support vector machines, reinforcement learning, partially observable Markov decision processes (POMDP), and / or table-based approximation, to provide several non-limiting examples. Based on the type of machine learning model, the training can include adjustment of one or more parameters. For example, in the case of a regression-based model, the training can include adjusting one or more coefficients of the regression so as to minimize a loss function such as a least-squares loss function. In the case of a neural network, the training can include adjusting weights, biases, number of epochs, batch size, number of layers, and / or number of nodes in each layer of the neural network, so as to minimize a loss function. Because each machine learning model is defined by its parameters (e.g., coefficients, weights, layer count, etc.), and because the parameters are based on the training data used to train the model, machine learning models trained based on different data differ from one another structurally and provide different outputs.
[0185] In some implementations, the face identification model 206 is additionally trained using images of inauthentic packaging and / or images without packaging. The use of this training data can improve the ability of the face identification model 206 to determine whether a packaging face is present in test images. For example, images of counterfeit packaging can be uploaded and annotated in a manner similar to the process 500, with suitable modifications to ensure that the images of counterfeit packaging are not treated as authentic.
[0186] Images processed according to the methods described herein (e.g., test images and images uploaded / captured to model training and digital blueprint generation) can be collected and labeled for various analyses. For example, images of packaging can be labeled with location to allow for location-aware authentication testing. For example, if tested packaging in a first country matches authentic packaging that is only used in a second country, the tested packaging may be determined to be inauthentic, or the tested packaging can be marked for further investigation. As another example, images of packaging can be labeled with timestamps to facilitate user-friendly tracking of changes in packaging over time, e.g., to view product relaunches over time.
[0187] Although the foregoing description refers to user interfaces and visualizations that can be displayed on the mobile device 402, it will be understood that at least some of the user interfaces and visualizations are not limited to display on the particular mobiledevice 402. For example, user interfaces that illustrate differences between a test image 1502 and a reference image 1504, as described with respect to the region-based comparison module 1510, can be provided on the mobile device 402 and on other user devices, for example, to allow later review of archived test images 1502 by administrators, packaging designers, anti-counterfeiters, and the like.
[0188] Accordingly, based on the foregoing systems and processes, machine learning models can be trained on auto-annotated images of packaging, speeding up data intake and improving user experience. Manual annotation can be used to supplement the autoannotation to improve the accuracy of training data and improve the outputs of trained machine learning models. Generation of multiple images of packaging with different characteristics can provide expanded sets of training data to improve machine learning model accuracy and / or reliability. The use of training data labeled with a face type results in machine learning models that differ from other machine learning models, allowing the machine learning models to detect face types in test images and / or for auto-annotation. The use of face type as a model output allows packaging faces having matching faces to be compared to one another, significantly reducing the search space for packaging authentication and reducing the computational resources used for authentication. The addition of a reference image selection process prior to final image authentication can also improve the efficiency of authentication.
[0189] Some features described herein, such as the systems 200, 400, 1500, and elements thereof, may be implemented in digital and / or analog electronic circuitry or in computer hardware, firmware, software, or in combinations of them. Some features may be implemented in a computer program product tangibly embodied in an information carrier, e.g., in a machine-readable storage device, for execution by a programmable processor. Method steps may be performed by a programmable processor executing a program of instructions to perform functions of the described implementations by operating on input data and generating output, by discrete circuitry performing analog and / or digital circuit operations, or by a combination thereof.
[0190] Some described features may be implemented advantageously in one or more computer programs that are executable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmitdata and instructions to, a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that may be used, directly or indirectly, in a computer to perform a certain activity or bring about a certain result. A computer program may be written in any form of programming language (e.g., Objective- C, Java, Python, JavaScript, Swift), including compiled or interpreted languages, and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0191] Suitable processors for the execution of a program of instructions include, by way of example, both general and special purpose microprocessors, and the sole processor or one of multiple processors or cores, of any kind of computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer may communicate with mass storage devices for storing data fdes. These mass storage devices may include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including by way of example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, ASICs (application-specific integrated circuits). To provide for interaction with a user the features may be implemented on a computer having a display device such as a CRT (cathode ray tube), LED (light emitting diode) or LCD (liquid crystal display) display or monitor for displaying information to the author, a keyboard and a pointing device, such as a mouse or a trackball by which the author may provide input to the computer.
[0192] Examples according to the present disclosure include, but are not limited to, the following.
[0193] Example 1: A method, including: capturing, at a mobile device, an image using a camera of the mobile device; processing, at the mobile device, the image using one or more machine learning models, wherein the one or more machine learning modelshave been trained to identify a face of first packaging in the image, and determine whether the first packaging in the image satisfies one or more capture conditions; providing, at the mobile device, feedback for image capture based on a first output of the one or more machine learning models relating to the one or more capture conditions; and in response to output of the one or more machine learning models indicating that the one or more capture conditions are satisfied, and in response to the output of the one or more machine learning models indicating that the face of the first packaging is present in the image, sending the image for authentication of the first packaging.
[0194] Example 2: The method of Example 1, wherein the one or more machine learning models have been trained to determine a face type of the face of the first packaging.
[0195] Example 3: The method of Example 2, wherein the face type includes a front face or a rear face.
[0196] Example 4: The method of Example 2 or 3, including determining whether the first packaging is authentic, wherein determining whether the first packaging is authentic includes: selecting, from two or more faces of second packaging, a first face based on the face type of the face of the first packaging matching a face type of the first face of the second packaging; determining at least one similarity between the first face and the face of the first packaging; and selecting, from among a plurality of images of packaging, an image of the second packaging as a reference image based on the at least one similarity between the first face and the face of the first packaging.
[0197] Example 5: The method of Example 4, wherein the at least one similarity includes: a textual similarity between text included on the face of the first packaging and text included on the first face of the second packaging, and a graphical similarity between the reference image and the image.
[0198] Example 6: The method of Example 4 or 5, wherein determining whether the first packaging is authentic includes: in response to selecting the image of the second packaging as the reference image, determining whether the first packaging is authentic based on a comparison between the first packaging in the image and the second packaging in the reference image.
[0199] Example 7: The method of any preceding Example, including determining whether the first packaging is authentic, wherein determining whether the first packaging is authentic includes: determining whether the image includes a data-encoding symbol; in response to determining that the image includes the data-encoding symbol, decoding data encoded by the data-encoding symbol, and determining a reference image based on the data, or, in response to determining that the image does not include the data-encoding symbol, determining the reference image based on a graphical comparison between the image and the reference image.
[0200] Example 8: The method of any preceding Example, including determining whether the first packaging is authentic, wherein including whether the first packaging is authentic includes: receiving the image from the mobile device, and processing the image using a machine learning model distinct from a first machine learning model, of the one or more machine learning models, that has been trained to identify the face of the first packaging in the image.
[0201] Example 9: The method of any preceding Example, including determining whether the first packaging is authentic, wherein determining whether the first packaging is authentic includes: determining a textual similarity between text in the image and text in a reference image; determining a graphical similarity between the image and the reference image; determining, based on at least one of the textual similarity or the graphical similarity, that the first packaging is not authentic; and determining, based on the image, a packaging of which the first packaging is a counterfeit.
[0202] Example 10: The method of any preceding Example, wherein the one or more capture conditions are based on at least one of an orientation of the first packaging in the image or a level of corruption in the image.
[0203] Example 11 : The method of any preceding Example, wherein the feedback for image capture includes at least one of: a graphical bound for placement of the first packaging during image capture, the graphical bound being moved to different locations on a display of the mobile device over capture of multiple images, an indication of whether an orientation of the first packaging satisfies an orientation condition, or a progress indicator that progresses based on satisfaction of the one or more capture conditions.
[0204] Example 12: The method of any preceding Example, wherein the feedback for image capture includes an indicator of a location of corruption in the image.
[0205] Example 13: The method of any preceding Example, including training the one or more machine learning models, wherein training the one or more machine learning models includes: obtaining an image of reference packaging; generating a plurality of images by modifying at least one of orientation, background, or contrast of the image of the reference packaging; and training the one or more machine learning models using the plurality of images as training data.
[0206] Example 14: The method of any preceding Example, including determining whether the first packaging is authentic, wherein determining whether the first packaging is authentic includes: comparing at least one feature of the first packaging to a digital blueprint of reference packaging, the digital blueprint including: a label indicating a face type of a face of the reference packaging, a graphical representation of the face of the reference packaging, and text included on the face of the reference packaging.
[0207] Example 15: The method of Example 14, including generating the digital blueprint, wherein generating the digital blueprint includes: processing an image of the reference packaging using a machine learning model that has been trained to determine the face type of the face of the reference packaging; and generating the digital blueprint based on an output of the machine learning model that has been trained to determine the face type of the face of the reference packaging.
[0208] Example 16: The method of any preceding Example, including training the one or more machine learning models using, as training data, images of faces of a plurality of packaging, and, as labels for the training data, data indicative of types of faces of the plurality of packaging portrayed in the images.
[0209]
[0210] Example 17: The method of any preceding Example, wherein the one or more machine learning models include a first machine learning model that has been trained to identify the face of the first packaging in the image, and a second machine learning model that has been trained to determine whether the first packaging in the image satisfies the one or more capture conditions.
[0211] Example 18: The method of any preceding Example including training the one or more machine learning models, wherein training the one or more machine learning models includes: providing, in a user interface, a display of an image of reference packaging captured by a second mobile device; processing the image of the reference packaging using a machine learning model that has been trained to identify a face of the reference packaging in the image of the reference packaging, to obtain, as an output, an auto-annotation indicative of at least one of text included in the face of the reference packaging, or a face type of the face of the reference packaging; providing, in the user interface, one or more tools usable to manually alter the auto-annotation to obtain a modified annotation; and training the one or more machine learning models using, as training data, the image of the reference packaging and the modified annotation.
[0212] Example 19: A non-transitoiy computer-readable medium tangibly encoding a computer program operable to cause a data processing apparatus to perform the operations of any of Examples 1-18 or 21-28.
[0213] Example 20: A system including: one or more computers programmed to authenticate packaging; and a mobile device communicatively coupled with the one or more computers, the mobile device being programmed to perform the operations of any of Examples 1-18 or 21-28.
[0214] Example 21 : A method, including: obtaining a first image of first packaging; identifying a first plurality of regions in the first image based on content in the plurality of regions; determining similarities between the first plurality of regions and a second plurality of regions in a second image of second packaging; and determining, based on the similarities, whether the first packaging is authentic.
[0215] Example 22: The method of Example 21, including identifying the plurality of regions using a trained machine learning model.
[0216] Example 23: The method of Example 21 or 22, wherein determining the similarities between the first plurality of regions and the second plurality of regions includes determining a displacement between a first region in the first plurality of regions and a corresponding second region in the second plurality of regions.
[0217] Example 24: The method of any one of Examples 21-23, wherein determining the similarities between the first plurality of regions and the second plurality of regionsincludes determining a graphical similarity between a first region in the first plurality of regions and a corresponding second region in the second plurality of regions.
[0218] Example 25: The method of any one of Examples 21-24, wherein determining the similarities between the first plurality of regions and the second plurality of regions includes determining a text morphology similarity between a first region in the first plurality of regions and a corresponding second region in the second plurality of regions.
[0219] Example 26: The method of any one of Examples 21-25, wherein determining the similarities between the first plurality of regions and the second plurality of regions includes determining a text content similarity between a first region in the first plurality of regions and a corresponding second region in the second plurality of regions.
[0220] Example 27: The method of any one of Examples 1-26, including determining whether the packaging is authentic by: identifying variable data in the first image; and determining whether the variable data is consistent.
[0221] Example 28: A method including: obtaining a first image of first inauthentic packaging, the first image annotated with at least one region including inauthentic content; projecting the at least one region onto a second image of authentic packaging, to annotate the second image of authentic packaging with at least one region including authentic content; training a machine learning model to identify inauthentic packaging using, as training data, the annotated first image and annotated second image; providing a test image as input to the machine learning model; and obtaining, as output of the machine learning model, a determination of authenticity of packaging in the test image.
[0222] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made. Elements of one or more implementations may be combined, deleted, modified, or supplemented to form further implementations. In yet another example, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.
Claims
What is claimed is:
1. A method, comprising: capturing, at a mobile device, an image using a camera of the mobile device; processing, at the mobile device, the image using one or more machine learning models, wherein the one or more machine learning models have been trained to identify a face of first packaging in the image, and determine whether the first packaging in the image satisfies one or more capture conditions; providing, at the mobile device, feedback for image capture based on a first output of the one or more machine learning models relating to the one or more capture conditions; and in response to output of the one or more machine learning models indicating that the one or more capture conditions are satisfied, and in response to the output of the one or more machine learning models indicating that the face of the first packaging is present in the image, sending the image for authentication of the first packaging.
2. The method of claim 1, wherein the one or more machine learning models have been trained to determine a face type of the face of the first packaging.
3. The method of claim 2, wherein the face type comprises a front face or a rear face.
4. The method of claim 2, comprising determining whether the first packaging is authentic, wherein determining whether the first packaging is authentic comprises: selecting, from two or more faces of second packaging, a first face based on the face type of the face of the first packaging matching a face type of the first face of the second packaging; determining at least one similarity between the first face and the face of the first packaging; andselecting, from among a plurality of images of packaging, an image of the second packaging as a reference image based on the at least one similarity between the first face and the face of the first packaging.
5. The method of claim 4, wherein the at least one similarity comprises: a textual similarity between text included on the face of the first packaging and text included on the first face of the second packaging, and a graphical similarity between the reference image and the image.
6. The method of claim 4, wherein determining whether the first packaging is authentic comprises: in response to selecting the image of the second packaging as the reference image, determining whether the first packaging is authentic based on a comparison between the first packaging in the image and the second packaging in the reference image.
7. The method of claim 1, comprising determining whether the first packaging is authentic, wherein determining whether the first packaging is authentic comprises: determining whether the image includes a data-encoding symbol; in response to determining that the image includes the data-encoding symbol, decoding data encoded by the data-encoding symbol, and determining a reference image based on the data, or in response to determining that the image does not include the data-encoding symbol, determining the reference image based on a graphical comparison between the image and the reference image.
8. The method of claim 1, comprising determining whether the first packaging is authentic, wherein determining whether the first packaging is authentic comprises: receiving the image from the mobile device, and processing the image using a machine learning model distinct from a first machine learning model, of the one or more machine learning models, that has been trained to identify the face of the first packaging in the image.
9. The method of claim 1, comprising determining whether the first packaging is authentic, wherein determining whether the first packaging is authentic comprises: determining a textual similarity between text in the image and text in a reference image; determining a graphical similarity between the image and the reference image; determining, based on at least one of the textual similarity or the graphical similarity, that the first packaging is not authentic; and determining, based on the image, a packaging of which the first packaging is a counterfeit.
10. The method of claim 1, wherein the one or more capture conditions are based on at least one of an orientation of the first packaging in the image or a level of corruption in the image.
11. The method of claim 1, wherein the feedback for image capture comprises at least one of: a graphical bound for placement of the first packaging during image capture, the graphical bound being moved to different locations on a display of the mobile device over capture of multiple images, an indication of whether an orientation of the first packaging satisfies an orientation condition, or a progress indicator that progresses based on satisfaction of the one or more capture conditions.
12. The method of claim 1, wherein the feedback for image capture comprises an indicator of a location of corruption in the image.
13. The method of claim 1, comprising training the one or more machine learning models, wherein training the one or more machine learning models comprises: obtaining an image of reference packaging;generating a plurality of images by modifying at least one of orientation, background, or contrast of the image of the reference packaging; and training the one or more machine learning models using the plurality of images as training data.
14. The method of claim 1, comprising determining whether the first packaging is authentic, wherein determining whether the first packaging is authentic comprises: comparing at least one feature of the first packaging to a digital blueprint of reference packaging, the digital blueprint comprising: a label indicating a face type of a face of the reference packaging, a graphical representation of the face of the reference packaging, and text included on the face of the reference packaging.
15. The method of claim 14, comprising generating the digital blueprint, wherein generating the digital blueprint comprises: processing an image of the reference packaging using a machine learning model that has been trained to determine the face type of the face of the reference packaging; and generating the digital blueprint based on an output of the machine learning model that has been trained to determine the face type of the face of the reference packaging.
16. The method of claim 1, comprising training the one or more machine learning models using as training data, images of faces of a plurality of packaging, and as labels for the training data, data indicative of types of faces of the plurality of packaging portrayed in the images.
17. The method of claim 1, wherein the one or more machine learning models comprise a first machine learning model that has been trained to identify the face of the first packaging in the image, and a second machine learning model that has been trainedto determine whether the first packaging in the image satisfies the one or more capture conditions.
18. The method of claim 1, comprising training the one or more machine learning models, wherein training the one or more machine learning models comprises: providing, in a user interface, a display of an image of reference packaging captured by a second mobile device; processing the image of the reference packaging using a machine learning model that has been trained to identify a face of the reference packaging in the image of the reference packaging, to obtain, as an output, an auto-annotation indicative of at least one of text included in the face of the reference packaging, or a face type of the face of the reference packaging; providing, in the user interface, one or more tools usable to manually alter the auto-annotation to obtain a modified annotation; and training the one or more machine learning models using, as training data, the image of the reference packaging and the modified annotation.
19. A non-transitory computer-readable medium tangibly encoding a computer program operable to cause a data processing apparatus to perform operations comprising: capturing, at a mobile device, an image using a camera of the mobile device; processing, at the mobile device, the image using one or more machine learning models, wherein the one or more machine learning models have been trained to identify a face of first packaging in the image, and determine whether the first packaging in the image satisfies one or more capture conditions; providing, at the mobile device, feedback for image capture based on a first output of the one or more machine learning models relating to the one or more capture conditions; and in response to output of the one or more machine learning models indicating that the one or more capture conditions are satisfied, and in response to the output of the oneor more machine learning models indicating that the face of the first packaging is present in the image, sending the image for authentication of the first packaging.
20. A system comprising: one or more computers programmed to authenticate packaging; and a mobile device communicatively coupled with the one or more computers, the mobile device being programmed to perform operations comprising: capturing, at a mobile device, an image using a camera of the mobile device; processing, at the mobile device, the image using one or more machine learning models, wherein the one or more machine learning models have been trained to identify a face of first packaging in the image, and determine whether the first packaging in the image satisfies one or more capture conditions; providing, at the mobile device, feedback for image capture based on a first output of the one or more machine learning models relating to the one or more capture conditions; and in response to output of the one or more machine learning models indicating that the one or more capture conditions are satisfied, and in response to the output of the one or more machine learning models indicating that the face of the first packaging is present in the image, sending the image to the one or more computers for authentication of the first packaging.
21. A method, comprising: obtaining a first image of first packaging; identifying a first plurality of regions in the first image based on content in the plurality of regions; determining similarities between the first plurality of regions and a second plurality of regions in a second image of second packaging; and determining, based on the similarities, whether the first packaging is authentic.
22. The method of claim 21, comprising identifying the plurality of regions using a trained machine learning model.
23. The method of claim 21, wherein determining the similarities between the first plurality of regions and the second plurality of regions comprises determining a displacement between a first region in the first plurality of regions and a corresponding second region in the second plurality of regions.
24. The method of claim 21, wherein determining the similarities between the first plurality of regions and the second plurality of regions comprises determining a graphical similarity between a first region in the first plurality of regions and a corresponding second region in the second plurality of regions.
25. The method of claim 21, wherein determining the similarities between the first plurality of regions and the second plurality of regions comprises determining a text morphology similarity between a first region in the first plurality of regions and a corresponding second region in the second plurality of regions.
26. The method of claim 21, wherein determining the similarities between the first plurality of regions and the second plurality of regions comprises determining a text content similarity between a first region in the first plurality of regions and a corresponding second region in the second plurality of regions.
27. The method of claim 21, comprising determining whether the packaging is authentic by: identifying variable data in the first image; and determining whether the variable data is consistent.
28. A method comprising: obtaining a first image of first inauthentic packaging, the first image annotated with at least one region including inauthentic content;projecting the at least one region onto a second image of authentic packaging, to annotate the second image of authentic packaging with at least one region including authentic content; training a machine learning model to identify inauthentic packaging using, as training data, the annotated first image and annotated second image; providing a test image as input to the machine learning model; and obtaining, as output of the machine learning model, a determination of authenticity of packaging in the test image.
Citation Information
Patent Citations
Training and identification method and system for identification model, equipment and medium
CN112651410A
Methods, devices, electronic equipment, and storage media for identifying genuine and counterfeit cigarettes.
CN114494765B
Machine learning based imaging method of determining authenticity of a consumer good
US20210192340A1
Method and system for identifying authenticity of an object
US20210374476A1
System and method for detecting the authenticity of products
US20220130159A1