Mark capture for signature generation
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- SYS TECH SOLUTIONS INC
- Filing Date
- 2024-08-14
- Publication Date
- 2026-05-27
AI Technical Summary
Existing technologies face challenges in generating reliable electronic signatures from images of marks, such as printed labels, due to issues with orientation and corruption in the images.
A method that uses a mobile device to capture multiple images of a candidate mark, process them using machine learning models trained for orientation and corruption identification, and adjust image capture based on model outputs to generate a high-resolution electronic signature.
This approach enables the generation of accurate and reliable electronic signatures by improving image quality through real-time analysis and adjustment, reducing corruption, and ensuring suitable orientation for signature generation.
Smart Images

Figure US2024042351_20022025_PF_FP_ABST
Abstract
Description
MARK CAPTURE FOR SIGNATURE GENERATIONCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 532,635 filed August 14, 2023, the disclosure of which is incorporated herein by reference in its entirety.FIELD OF THE DISCLOSURE
[0002] Technologies are described for generating electronic signatures based on marks such as printed labels.BACKGROUND
[0003] Aspects of a printed label can be used to generate electronic signatures for identification, authentication, and other purposes.SUMMARY
[0004] Some aspects of this disclosure describe a method. The method includes capturing, at a mobile device, multiple images using a camera of the mobile device. The method includes processing, at the mobile device, a first subset of the multiple images at a first resolution using two or more machine learning models. A first of the two or more machine learning models has been trained to identify an orientation of a candidate mark in the first subset of the multiple images, and a second of the two or more machine learning models has been trained to identify an amount of corruption in an image of the candidate mark. The method includes providing, at the mobile device, one or more instructions to modify the capturing of a second subset of the multiple images based on a first output of the first model trained to identify the orientation of the candidate mark, and based on a second output of the second model trained to identify the amount of corruption in the image of the candidate mark. The method includes generating, at the mobile device, an electronic signature of the candidate mark from two or more images of the second subset of the multiple images, at a second resolution that is higher than the first resolution, in response to both the first output indicating the orientation is suitable for electronic signature generation and the second output indicating the amount ofcorruption is satisfies a threshold condition. The method includes sending the electronic signature of the candidate mark for authentication.[0051 This and other methods described herein can have one or more of at least the following characteristics.
[0006] In some implementations, the first subset of the multiple images is captured at the first resolution, and the second subset of the multiple images is captured at the second resolution.
[0007] In some implementations, the two or more machine learning models include three machine learning models including the first model trained to identify the orientation of the candidate mark, the second model trained to identify the amount of corruption in the image of the candidate mark, and a third model trained to identify spatial dimensions of the candidate mark.
[0008] In some implementations, the providing includes displaying, on the mobile device, at least one instruction to a user of the mobile device to adjust a distance between the mobile device and the candidate mark, based on the identified spatial dimensions of the candidate mark.
[0009] In some implementations, the providing includes displaying, on the mobile device, at least one instruction to a user of the mobile device regarding how to move the mobile device in three-dimensional space to improve image suitability, reduce image corruption, or both.
[0010] In some implementations, the providing includes sending, to the camera of the mobile device, at least one instruction that adjusts at least one setting of the camera to improve image suitability, reduce image corruption, or both.[OH] In some implementations, processing the first subset of the multiple images using the two or more machine learning models includes identifying, using the first model, an area of a first image of the first subset of the multiple images, the area corresponding to the candidate mark; cropping the first image to obtain a cropped image that includes the identified area and excludes at least another area of the first image; and providing, as input to the second model, the cropped image.
[0012] In some implementations, generating the electronic signature includes processing a first image of the two or more images using the second model to generate acorruption mask defining a corrupted area in the first image; and producing the electronic signature using a masked version of the first image based on the corruption mask, the masked version excluding the corrupted area.
[0013] In some implementations, the second of the two or more machine learning models has been trained to identify at least one of the following in the first subset of the multiple images: an area exhibiting glare, or an area that is shadowed.
[0014] In some implementations, the first of the two or more machine learning models has been trained to determine whether the candidate mark is tilted with respect to a target orientation.
[0015] In some implementations, the second of the two or more machine learning models has been trained to identify a corrupted area in the image, and providing the one or more instructions includes displaying an overlay on a live image of the candidate mark, the overlay covering a corrupted area in the live image.
[0016] In some implementations, a color of the overlay is indicative of a type of corruption in the corrupted area.
[0017] In some implementations, the candidate mark includes a barcode.
[0018] In some implementations, generating the electronic signature of the candidate mark is based on printing artifacts in the candidate mark.
[0019] In some implementations, the method includes training the two or more machine learning models using, as training data, a plurality of images of different printed marks, the plurality of images captured under different lighting conditions, from different perspectives, and with different levels of image corruption.
[0020] In some implementations, the threshold condition includes that the level of corruption is less than a threshold, and the method includes: using the second model to determine levels of corruption in a plurality of images; generating electronic signatures of candidate marks in the plurality of images; and based on a level of success in generating the electronic signatures, and based on the levels of corruption in the plurality of images, adjusting the threshold.
[0021] In some implementations, the amount of corruption includes a proportion of an area of the candidate mark that is corrupted.
[0022] In some implementations, the method includes: receiving, at the mobile device, an authentication response for the candidate mark; determining, by the mobile device, based on the authentication response, whether the candidate mark is authentic; and displaying, by the mobile device, an indication of the authenticity of the candidate mark.
[0023] The described methods can be associated at least with corresponding systems, processes, devices, and / or instructions stored on non-transitory computer-readable media. For example, some aspects of this disclosure describe a non-transitory computer-readable medium tangibly encoding a computer program operable to cause data processing apparatus to perform operations of the foregoing method and / or other methods described herein. Further, some aspects of this disclosure describe a system including one more computers programmed to authenticate electronic signatures of candidate marks; and a mobile device communicatively coupled with the one or more computers, the mobile device being programmed to perform operations of the foregoing method and / or other methods described herein.
[0024] Implementations described herein, such as the foregoing method, can provide various advantages. For example, image capture can be improved by use of real-time analysis and capture adjustment, e.g., using instructions provided to a user, camera settings adjustment, or both. Image processing for this adjustment can be performed locally and in real-time, resulting in low-latency, live-updating instructions. Moreover, instructions can be provided in the form of intuitive graphical displays such as augmented reality overlays, facilitating accurate and rapid corrections based on the instructions. Accordingly, higher-quality images can be captured, e.g., images with improved focus and reduced corruption. And, corruption that remains in captured images can have its effect on signature generation reduced or eliminated, e.g., by excluding corrupted areas from the signature generation process. Based on these and other processes described in this disclosure, electronic signatures that authenticate and / or identify items can be determined accurately, reliably, and quickly.
[0025] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other aspects, features and advantages will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0026] FIG. l is a diagram showing an example of a mark.
[0027] FIGS. 2A-2B are diagrams showing an example of a signature generation process.
[0028] FIGS. 3A-3B are diagrams showing an example of a system associated with signature generation.
[0029] FIG. 4 is photograph showing an example of a processed image.
[0030] FIG. 5 is a diagram showing an example of corruption-related processing.
[0031] FIGS. 6A-6B are examples of screen displays.
[0032] FIG. 7 is an example of a screen display.
[0033] FIG. 8 is a diagram showing an example of a candidate mark analysis process.
[0034] FIG. 9 is a diagram showing an example of a machine learning training process.DETAILED DESCRIPTION
[0035] This disclosure relates to capturing and processing images of marks (e.g., barcodes) for electronic signature generation. For example, images can be processed using various machine learning models, including an orientation model trained to identify the orientation of a mark in an image, and a corruption model trained to identify corruption (such as over-bright spots, shadow, etc.) in an image. Based on outputs of the models, image capture can be adjusted, such as by providing instructions to a user and / or by altering camera settings. Images captured based on this adjustment can then be used to generate electronic signatures with high reliability and accuracy.
[0036] FIG. 1 shows an example of a mark 100 that can be imaged and processed according to the methods discussed in this disclosure. The mark 100 encodes data that can be processed by visually scanning and analyzing the mark 100. In this example, the mark 100 is a linear barcode. However, other types of marks (sometimes referred to as symbologies) are also within the scope of this disclosure, such as other barcode types, including, for example, matrix barcodes such as QR (Quick Response) codes, and other data-encoding visual markers. The mark 100 can be, for example, a label printed on an item to identify a type (e g., model) of the item.
[0037] Besides the data encoded intentionally in the mark 100, printing artifacts in the mark 100 vary on a mark-by-mark (e.g., label-by-label) basis. For example, the mark 100 exhibits variations 102 in the width of or spacing between bars; variations 104 in the average color, pigmentation and / or intensity of bars or portions thereof; voids 106 in black bars (or black spots in white stripes); and irregularities 108 in the shape of the edges of the bars. While the same mark printed on different items (e.g., different units of a single type of item) will encode the same data (e.g., a product identifier), these and other printing artifacts are expected to vary each time the mark is printed, or at least when the mark is printed by different printing devices.
[0038] Accordingly, an encoding of the artifacts in the form of an electronic signature can be used to track individual items and to verify authenticity of the items. For example, a mark printed on a potentially-counterfeit item can be imaged and its artifacts processed to obtain a first electronic signature, and the first electronic signature can be compared to a stored electronic signature that is known to authentically characterize the mark printed on the authentic item. If the signatures match, the potentially-counterfeit item is confirmed to be authentic. Further details on electronic signature generation and associated processes can be found in U.S. Patent No. 9,519,942 and in PCT Application No. PCT / US2023 / 019484, each of which is incorporated herein by reference in its entirety.
[0039] Because the printing artifacts based on which the electronic signatures are generated may be subtle, it may be desirable to capture high-quality images of marks so that the electronic signatures can be generated more reliably. For example, it may be desirable to capture images that are high-resolution and low in corruption, such as imaging shadows and glare. Moreover, to provide for efficient analysis and a good user experience, it may be desirable to perform the image capture (and, if necessary, capture correction) process quickly and with low latency. Implementations according to this disclosure provide efficient image capture and image processing operations to obtain these high-quality images and utilize them to generate accurate signatures.
[0040] FIGS. 2A-2B show an example of a process 200 according to some implementations of this disclosure, and FIGS. 3A-3B show an example of a system 300associated with process 200. For example, elements of the system 300 can be configured to perform process 200.[0411 The system 300 includes a mobile device camera 302, a mobile device display 304, a machine learning module 316 implementing machine learning models 306, 308, 310, an instruction module 312, and a signature generation module 314. The mobile device camera 302 and the mobile device display 304 can be included in the same mobile device, e.g., a smartphone, a tablet, a laptop, a wearable device, or another type of mobile device.
[0042] The modules 312, 314, 316 can be hardware and / or software modules, e.g., implemented by one or more computer systems. For example, the modules 312, 314, 316 can include one or more processors and one or more computer-readable mediums encoding instructions that, when executed by the one or more processors, cause the one or more processors to perform operations such as portions of process 200 and / or other processes described herein. The modules 312, 314, 216 can be software modules executed by one or more processors based on instructions encoded in one or more computer-readable mediums. In some implementations, the modules 314, 314, 316 are modules of a mobile device, such that the processes described herein can advantageously be performed largely or entirely on a mobile device, in some cases reducing processing latency compared to processes that require execution on a remote system. The modules 312, 314, 316 need not be separate but, rather, can be at least partially integrated together as one or more combined modules or fully integrated together as a single program on the mobile device.
[0043] As shown in FIG. 2A, in the process 200, multiple images are captured using a camera of a mobile device (202), such as mobile device camera 302. For example, the multiple images can be images of a candidate mark, such as a barcode or other label, such as a printed label. The multiple images can be captured one after another (e.g., in a sequence, such as in a sequence of images that form a video or live image), and / or can be captured in spaced-out capturing operations, in some cases with intervening process(es) performed between capture of at least some of the multiple images. Capturing an image can include obtaining signals and / or data representative of the image (e.g., signals from photodetectors of the camera) and can include one or more processing steps performed onthe image, such as downscaling / upscaling, size and / or resolution adjustment, computational distortion compensation, brightness adjustment (e.g., to compensate for underexposure / overexposure), and / or one or more other image adjustment operations.
[0044] The process 200 further includes processing a first subset of the multiple images using two or more machine learning models (204). For example, the two or more machine learning models can include an orientation model 306, a dimensions model 308, and a corruption model 310. Some implementations according to this disclosure include the orientation model 306 and the corruption model 310 without including the dimensions model 308. In some implementations, operations discussed herein with respect to the dimensions model 308 can be performed with respect to a combined orientationdimension model configured as described with respect to both the orientation model 306 and the dimensions model 308. Moreover, in some implementations, the distinct models discussed herein can be wholly or partially combined into one or more combined models (e.g., as sub-models, stages, layers, or other portions of the combined model(s)) without departing from the scope of this disclosure. Nonetheless, using separate models can provide the advantage of flexibility in turning off or upgrading one of the models without impacting the other models, e.g., turning off the dimensions model 308 without affecting the orientation model 306 and the corruption model 310.
[0045] The orientation model 306 has been trained to detect an orientation of the candidate mark in the first subset of images. In some implementations, the orientation model 306 has been trained to detect a location of the candidate mark in the first subset of images and optionally correct any skew in the image of the candidate mark. Detecting the location of the candidate mark can include, for example, performing an object recognition process to identify the candidate mark in the first subset of images, and determining bounds of the candidate mark in the first subset of images. For example, the orientation model 306 can be trained to segment each of the first subset of images into a first portion that is the candidate mark, and a second portion that is background.
[0046] For detecting the orientation of the candidate mark, the orientation model 306 can be trained to determine whether the candidate mark is tilted with respect to a target orientation and, in some implementations, a degree and / or characteristic of the tilt. For example, the target orientation can be a right-side-up orientation in which a top side ofthe candidate mark faces a top of the image frame and a bottom side of the candidate mark faces a bottom of the image frame (e.g., a roll angle of zero), and / or in which the candidate mark is imaged head-on (e.g., pitch and / or yaw angles of zero). In some implementations, the orientation model 306 is trained to determine a quantity associated with the orientation, such as an angular degree of roll, pitch, and / or yaw. In some implementations, the orientation model 306 is trained to classify the orientation, for example, into “correct orientation,” “tilt,” or “inverted,” where “correct orientation” can be, for example, a tilt within a predetermined difference from the target orientation, and where “inverted” can be an up-side down image. “Tilt” can include a tilted orientation in one or more dimensions, e.g., pitch, yaw, and / or roll.
[0047] In some cases, image orientation is related to image fidelity. For example, if a candidate mark is captured at a tilted angle, portions of the candidate mark further from the camera may be captured with reduced resolution / fidelity based on the portions’ further distance. Moreover, in cases of tilted image capture, portions of the candidate mark may be out-of-focus. Accordingly, in some implementations, correction of image orientation results in capture of images more suitable for electronic signature generation.
[0048] FIG. 4 shows an example of an image 400 processed based on output from the orientation model 306. The orientation model 306 detects a candidate mark 402 (in this case, a printed barcode and associated printed white area) and determines a boundary 404 that contains the candidate mark 402. In some implementations, the boundary 404 is displayed as an overlay in a live image (e.g., as shown in FIGS. 6A and 7). In some implementations, the boundary 404 is used to define / generate a segmented version of the image 400 in which the area 406 outside the boundary 404 is cropped out. The segmented version, in some implementations, can be used for subsequent processing, such as by the dimensions model 308 and / or the corruption model 310.
[0049] As further shown in FIG. 4, the orientation model 306 determines and outputs a prediction 408 (sometimes referred to as an estimation, classification, or determination) indicative of an orientation of the candidate mark 402. In this example, the prediction 408 is a classification of “correct orientation” with a confidence level of 99%.
[0050] The dimensions model 308 has been trained to identify spatial dimensions of the candidate mark in the first subset of images or images derived therefrom. Forexample, the dimensions model 308 can determine a width and / or height (e.g., in inches, centimeters, etc.) of the candidate mark, and / or determine a scaling of the candidate mark with respect to a predetermined size. The spatial dimensions can be used to provide instructions that improve imaging of the candidate mark (e.g., improve a focus level of the imaging).
[0051] In some implementations, the dimensions model 308 is configured to process segmented images that are limited to the candidate mark itself, e.g., without background portions of the images. For example, the orientation model 306 can output, as input to the dimensions model 308, the first subset of images cropped to exclude at least some of the background in each image, and / or the orientation model 306 can output, as input to the dimensions model 308, information describing locations of the candidate mark in the first subset of images so that the dimensions model 308 crops the images to exclude at least some of the background. In some implementations, the dimensions model 308 is configured to implement an ensemble approach in which both the candidate mark itself and the background of the first subset of images are analyzed to determine the dimensions. For example, the dimensions model 308 can be trained to perform object recognition to identify the candidate mark in the first subset of images, and then to process either (i) versions of the first subset of images cropped to exclude at least some of the background, or (ii) the first subset of images including background portions and candidate mark portions, to determine the dimensions. In some implementations, analysis of background features can result in more-accurate dimensions determination, e.g., because the background features can be indicative of the dimensions of the candidate mark.
[0052] The corruption model 310 has been trained to identify an amount of corruption in the first subset of images or images derived therefrom. Corruption includes visual features that may inhibit analysis of the images to generate electronic signatures. For example, corruption in an image can include shadow, diffused glare, and / or spot glare, where shadowed portions of the candidate mark may be too dark for reliable detection of artifacts to generate electronic signature, and where portions of the candidate mark imaged with glare may be too bright for reliable detection of the artifacts. The corruption model 310 can identify portions of the images that exhibit corruption (arecorrupted), e.g., for generation of image masks, for display of overlays, and / or for limiting which portions of the images are used for signature generation, as discussed with respect to FIGS. 5 and 7 and with respect to process 208.
[0053] The amount of corruption in an image can include an area of the candidate mark in the image that is corrupted, e.g., where corruption is determined to be present. The area can include an absolute area (e.g., in cm2) and / or a proportion of the area of the candidate mark. For example, the corruption model 310 can be trained to determine that x% of a candidate mark in an image is obscured by corruption.
[0054] As discussed with respect to the dimensions model 308, in some implementations the corruption model 310 receives and / or generates, as input, cropped images based on outputs of the orientation model 306, where the cropped images are the first subset of images cropped to exclude background portions of the images. For example, the corruption model 310 can receive, as input, from the orientation model 306, data indicative of the boundary 404 shown in FIG. 4, and can crop the first subset of images using the data indicative of the boundary 404, to obtain cropped images that are largely or entirely limited to the candidate marks in the first subset of images. The corruption model 310 can then process the cropped images to determine an amount of corruption in the cropped images. This may reduce processing time and / or consumption of computational resources compared to analysis of uncropped images, e.g., because less image data needs to be analyzed for corruption detection.
[0055] As shown in FIG. 5, in some implementations, the corruption model 310 is trained to classify / label portions of images exhibiting different types of corruption, and to output the classification / labeling, e.g., as a mask, an overlay, and / or another form. In this example, the corruption model 310 processes an input image 500 that is cropped to exclude background portions of an image in the captured first subset of images, such that the input image 500 is substantially or entirely a candidate mark 502. The corruption model 310 is trained to identify portions of the input image 500 associated with different types and / or levels of corruption, such as an uncorrupted region 504 that exhibits little or no corruption; a diffused glare region 506 that exhibits diffused glare; and / or a spot glare portion 508 that exhibits spot glare (e.g., specular reflection). Although not shown in thisexample, in some implementations, the corruption model 310 can instead or additionally identify a shadow region that exhibits shadow.[0561 The corruption model 310 can be configured to output a mask representing identified regions such as 504, 506, 508. For example, the mask can indicate, on a pixel- by-pixel basis or other basis, whether each portion of the candidate mark in an image is corrupted and, in some implementation, if so, what type of corruption is exhibited. For example, a green overlay or coloring (508) can indicate spot glare, a red overlay or coloring (506) can indicate diffused glare, and a black overlay or coloring (504) can indicate a background portion. The mask can be used to provide an overlay on images of the candidate mark, e.g., as shown with segmentation overlay 510 in FIG. 5, in which different overlay colors correspond to different types of corruption. This image overlay can be displayed in real-time, thus giving the user of the mobile device immediate feedback on the usability of the image of a mark as the mark is being imaged. Further, in some implementations, the mask is used to limit a portion of an image that is used for generation of an electronic signature, e.g., as described with respect to process 208.
[0057] Referring again to FIG. 2A, the process 200 includes providing, at the mobile device, one or more instructions to modify the capturing of a second subset of the multiple images based at least on outputs of the orientation model and the corruption model (206). Providing the instructions can include, for example, displaying, on the mobile device, an instruction to a user of the mobile to adjust conditions of capture of the second subset of the multiple images (e.g., mobile device position, mobile device orientation, and / or lighting conditions), and / or can include sending, to the camera of the mobile device, an instruction that adjusts at least one camera setting to improve capture of the second subset of the multiple images. The instructions can be provided by the instruction module 312 based on outputs of one or more of the machine learning models.
[0058] In some implementations, instructions are displayed in the context of a live image of the candidate mark displayed on the user device. The live image is a real-time or near-real-time display of a series of images captured by the camera of the mobile device. For example, as a user moves the mobile device and changes a position / orientation of the camera, the live image changes based on the altered position / orientation. Instructions can be displayed on and / or in proximity to the liveimage to facilitate intuitive and user-friendly adjustment of the capture conditions. Note that the instructions to the user need not include written instructions and can be visual instructions, such as an animation showing how to move the mobile device and / or an image overlay showing usable portions of a live image in real time. Individual image frames of the live image can be processed as described herein, e.g., as described for processes 202, 804, 806, and / or 808. The live image can be enhanced with augmented reality (AR) overlays, e.g., as described in reference to FIGS. 6A-6B and 7.
[0059] For example, as shown in FIG. 6A, a display 600 of a mobile device (e g., the mobile device display 304, such as a screen of a smartphone) is controlled to present a live image 602 representing real-time captures of a camera of the mobile device. Overlaid on the live image 602 is a boundary 604 representing a boundary of a candidate mark 606, as estimated by the orientation model 306. In this example, a dimension 608 of the candidate mark 606 is also overlaid on the live image 602, the dimension 608 estimated by the dimensions model 308; in some implementations, the dimension is not overlaid. In some implementations, a display style (e.g., color, thickness, and / or line type) of the boundary 604 indicates whether the orientation of the candidate mark 606 is correct, e.g., as determined by the orientation model 306. For example, the boundary 604 can be orange to indicate an incorrect orientation, and green to indicate a correct orientation.
[0060] In the situation shown in FIG. 6A, the orientation model 306 determines that the as-captured candidate mark 606 has a tilted orientation. Accordingly, the instruction module 312 causes display of an instruction 610 to “Rotate the code [the candidate mark] by 180 degrees,” to cause the orientation to change to a “correct orientation” that is suitable for use of the imaged candidate mark 606 for electronic signature generation. An accompanying animation 614 shows a phone being rotated by 180 degrees, mimicking the motion that the user is being instructed to perform.
[0061] The instruction to rotate the mobile device, as shown in FIG. 6A, is an example of an instruction regarding how to move the mobile device in three-dimensional space to improve image suitable and / or reduce image corruption. Examples of other instructions (which can include textual, graphical, and / or audio instructions) for movement of the mobile device include instructions to tilt the mobile device, instructions to translate the mobile device laterally with respect to the candidate mark, andinstructions to move the mobile device closer to or further from the candidate mark. These instructions can be based on outputs of one or more of the machine learning models 306, 308, 310. For example, rotation and / or tilting can be instructed to cause the orientation to match a target orientation, based on the orientation model 306 indicating that the current orientation is not correct. As a further example, movement of the mobile device can be instructed to reduce glare and / or shadow in a captured image, and / or to move the glare and / or shadow to another portion of the image (e.g., from a location that obscures the candidate mark to a background location that does not obscure the candidate mark), based on an output of the corruption model 310.
[0062] In some implementations, an instruction regarding how to move the mobile device is based on an output of the dimensions model 308. The dimensions model 308 is configured to output physical dimensions of the candidate mark, and these dimensions are associated with a target imaging distance at which the candidate mark can be imaged with maximum focus and detail, without losing visibility of the candidate mark. For example, imaging distances closer than the target imaging distance may result in lost focus and / or an incompletely-imaged candidate mark, while imaging distances farther than the target imaging distance may result in reduced fidelity / detail of capture of the candidate mark. The target imaging distance can also be based, for example, on calibration data specific to the mobile device being used, e.g., an inferred dots-per-inch (DPI) of the camera of the mobile device. Based on the estimated spatial dimensions, the instruction module 312 can display an instruction to adjust a distance between the mobile device and the candidate mark (e.g., to move the mobile device closer to or further from the candidate mark) to achieve the target imaging distance.
[0063] As shown in FIG. 6B, in another example of an instruction provided based on model output, an instruction 612 advises a user to tilt the top of the mobile device down. This can provide a more head-on capture angle of the candidate mark 606 (e.g., based on the orientation model 306 indicating that capture angle is overly tilted) and / or can reduce a prominence of corruption, such as glare or shadow, on the candidate mark 606.
[0064] In some implementations, the instruction module 312 is configured to display an overlay on the live image, the overlay representing corruption in the live image. For example, as shown in FIG. 7, a display 700 presents a live image 702 of real-timecaptures by a camera of a mobile device. A boundary 704 (e.g., output by the orientation model 306 or another machine learning model) is displayed substantially defining the area of a candidate mark 706. In addition, a corruption overlay 708 is displayed on the candidate mark 706. The corruption overlay 708 includes a red portion 712 defining an area of diffused glare on the candidate mark 706 and a green portion 710 defining an area of spot glare on the candidate mark 706. Instructions 714 advise a user to reduce the glare. Based on the presentation of the overlay 708, a user can adjust the capture conditions (e.g., by moving the mobile device and / or the item on which the candidate mark 706 is printed, and / or by adjusting lighting conditions) to reduce the glare. Changes in the corruption are reflected in real-time by the corruption overlay 708 on the live image 702 to provide real-time feedback to the user, aiding in appropriate adjustment.
[0065] In some implementations, the instruction module 312 is configured to display a visual representation of a degree of progression in improving image capture. For example, a displayed representation of a barcode can be progressively filled-in as image suitability increases, e.g., as the orientation is improved, as an amount of corruption decreases, etc. In some implementations, the instruction module 312 is configured to display an overlay or other graphical indicator indicating an amount of area (e.g., area of the candidate mark) that must be captured for signature generation, e.g., with low or no corruption. For example, the graphical indicator can indicate (in absolute and / or fractional / percentage terms) how much more of the capture area or candidate mark must be “filled in” for accurate signature generation and a corresponding authentication response.
[0066] In some implementations, the instruction module 312 is configured to use haptic devices in the mobile device to provide tactile feedback. For example, vibrations and / or other tactile sensations can be provided to indicate proper alignment, successful capture of areas of the candidate mark, and / or a need for capture adjustment.
[0067] In some implementations, the instruction module 312 is configured to provide instructions by adjusting one or more settings of the mobile device camera 302, e.g., to improve focus, reduce corruption, and / or otherwise make captured images more suitable for signature generation. For example, one or more of camera focus (e.g., focus distance), zoom, exposure time, aperture setting (e.g., aperture width), or ISO (InternationalStandards Organization) value (for light sensitivity) can be adjusted to improve resolution, reduce corruption (e.g., by reducing an effect of glare), and / or otherwise improve image suitability based on outputs of one or more of the models 306, 308, 310.
[0068] In some implementations, the instruction module 312 is configured to provide instructions by adjusting one or more image processing steps. For example, the instruction module 312 can, based on the outputs of one or more of the models 306, 308, 310, adjust a resolution of the second subset of images, e.g., by adjusting a manner in which light representative of the second subset of images is sensed and / or by adjusting downscaling / upscaling of the second subset of images.
[0069] The instructions modify the capturing of a second subset of the multiple images. For example, in some implementations, the first subset of images (which are processed using the machine learning module 316) are captured, instructions are provided to adjust camera settings, mobile device position in space, etc., and the second subset of images are captured after provision of the instructions, using the adjusted camera settings, position in space, etc. In some implementations, the first subset of images and the second subset of images are at least partially captured prior to provision of the instructions, and the instructions dictate which of the multiple images are to be included in the second subset of images. In some implementations, the first subset of images and the second subset of images are at least partially captured prior to provision of the instructions, and the instructions dictate how the second subset of images are to be processed / used , e.g., by adjusting a resolution at which the second subset of images is captured and / or used for electronic signature generation.
[0070] Referring again to FIG. 2A, the process 200 includes generating, at the mobile device, an electronic signature of the candidate mark from two or more images of the second subset of images, in response to outputs of the machine learning models indicating that characteristics of the first subset of images are suitable for electronic signature generation (208). For example, the orientation model 306 can indicate that the orientation is suitable for signature generation (e.g., a “correct orientation” classification), and the corruption model 310 can indicate that the amount of corruption is satisfies a threshold condition, e.g., is less than a threshold value and / or otherwise acceptable. The electronic signature can be generated by the signature generation module 314. Asdiscussed with respect to FIG. 1, the electronic signature can be generated by encoding printing variations and / or other characteristics of the candidate mark, e.g., into a hash identifier. For example, dimensional characteristics (e.g., width, spacing, etc.) of bars or other features of the candidate mark in the second subset of images can be determined, and the electronic signature can be generated based on the dimensional characteristics. The electronic signature can be, for example, a string of values forming a hash encoding the characteristics. Details on examples of electronic signature generation can be found in U.S. Patent No. 9,519,942 and in PCT Application No. PCT / US2023 / 019484; however, other methods of electronic signature generation based on mark characteristics are also within the scope of this disclosure.
[0071] In some implementations, the first subset of images is processed at a first resolution, and the second subset of images is processed at a second resolution for generation of the electronic signature, where the second resolution is higher than the first resolution. For example, the subsets of images can be captured at different resolutions in that photodetectors of the camera of the mobile device are operated and / or read from differently between the subsets of images, resulting in differing resolutions. Instead, or additionally, the subsets of images can be processed and / or used differently to arrive at the different resolutions. For example, the first subset of images can be downscaled, or otherwise reduced in resolution, to the first resolution, and the second subset of images can be not downscaled, or can be downscaled less, to have the second resolution that is greater than the first resolution. In some implementations, this possible resolution reduction is performed by one or more of the models 306, 308, 310 themselves. In some implementations, another computing module performs the resolution reduction. In some implementations, the first subset of images and the second subset of images at least partially originate from common photodetection by the camera, and the subsets are differentiated based on different processing. For example, the first subset of images can include downscaled versions of images in the second subset of images.
[0072] The use of images having different resolutions for different portions of the process 200 can be advantageous for rapid processing and an improved user experience. For providing feedback / instructions on candidate mark imaging (e.g., to determine camera positioning), it can be desirable to capture and process images very quickly, sothat the user receives real-time feedback, e.g., displayed on a live, real-time-adjusted image of the candidate mark. For example, it may be desirable to perform image capture and processing (e.g., orientation detection and / or corruption detection) by the machine learning module 316 within 150 ms or 200 ms. The capture / use of lower-resolution images for this portion 204 of the process 200 (as the first subset of images) can facilitate processing under that time constraint, because (i) the camera can perform photodetection for imaging faster when capturing at a lower resolution, and / or (ii) the models 306, 308, 310 can process lower-resolution images faster than they can process higher-resolution images. Moreover, lower-resolution images can still provide sufficient fidelity for use by the models 306, 308, 310 for orientation determination, candidate mark detection, dimension determination, and corruption detection. In some implementations, the resolution of the first subset of images is less than 400x400 pixels, e.g., less than 200x200 pixels, such as 120x140 pixels. These resolution ranges have been found, for purposes of this disclosure, to provide rapid image processing.
[0073] For generation of the electronic signature (208), it may be desirable to use (as the second subset of images) high-resolution images, e g., images having resolution above 1000x1000 pixels, above 2000x2000 pixels, or above 3000x3000 pixels. These resolutions can provide sufficient fidelity for accurate and reproducible representation of small visual characteristics / variations of the candidate mark, so that electronic signatures for the candidate mark can be generated reliably, e.g., to match previously-generated electronic signature for the mark. In some cases, because relatively few images may be used for signature generation as the second subset of images (e.g., fewer than ten), it may be acceptable for capture and processing of the second subset of images to consume more time on a per-image basis than for the first subset of images.
[0074] In some implementations, capture of the second subset of images is performed in response to the orientation and corruption level being suitable for signature generation. As such, the capture of higher-resolution images (which may be slower than the capture of lower-resolution images) can be delayed until it is predicted that the to-be-captured images will be suitable for signature generation, and the possible increased time for capture of the higher-resolution images may have limited negative impact, as relatively few such images may be captured.
[0075] As noted above, electronic signature generation can be performed conditional on a detected amount of corruption, e.g., when the detected amount of corruption is less than a threshold. In some implementations, the corruption threshold is based on an area of the candidate mark that the corruption model 310 determines exhibits corruption. For example, the corruption model 310 can generate a mask of corruption on the candidate mark, and the amount of corruption can be based on an area of the mask at which corruption is present. The threshold can be, for example, a proportion of the area of the candidate mark that exhibits corruption, such as 5%, 10%, 15%, 20%, or another value. When the amount of corruption is greater than the threshold, in some cases, too little of the candidate mark is available for reliable electronic signature generation, e.g., because printing variations may not be visible in a shadowed portion of the candidate mark or a portion of the candidate mark exhibiting glare. The area of the candidate mark can be identified by appropriate mark detection and image segmentation, e g., as performed by the orientation model 306 or another model / module, as discussed above.
[0076] FIG. 3B shows an example of the system 300 in the context of electronic signature generation. In some implementations, at least the orientation model 306 and / or the corruption model 310 perform processing on the second subset of images (or images derived therefrom, e.g., reduced-resolution versions of the second subset of images) as discussed in reference to FIG. 3A and process 204, to obtain model outputs that are used for electronic signature generation.
[0077] For example, in some implementations, the orientation model 306 performs segmentation on the second subset of images or related images (e.g., lower-resolution versions of the second subset of images), to identify portions of the second subset of images that include the candidate mark. The segmentation can be provided to the signature generation module 314 as shown in FIG. 3B, such that the signature generation module 314 generates the signature based on cropped versions of the second subset of images, where the cropped version excludes at least some other portions of the second subset of images that do not include the candidate mark. This may reduce a time consumed in generating the electronic signature and / or improve the accuracy of generating the electronic signature by removing extraneous background features. In addition, or alternatively, the segmentation can be provided to the corruption model 310,so that the corruption model 310 limits corruption detection to an area of the candidate mark.[0781 As a further example, in some implementations, the corruption model 310 performs corruption detection on the second subset of images or related images (e.g., lower-resolution versions of the second subset of images), to generate a mask defining corrupted areas in the second subset of images. For example, referring to FIG. 5, the mask can have a value of “1” in regions 506 and 508 (indicating the presence of corruption) and a value of “0” in region 504 (indicating little or no corruption). The noncorrupted area need not be continuous but, rather, can include distinct portions separated by corrupted areas. As shown in FIG. 3B, the signature generation module 314 can obtain masks for the second subset of images and generate the electronic signature using masked versions of the second subset of images, the masked versions excluding corrupted areas indicated by the masks. This may improve the accuracy of signature generation by limiting the generation to use visible, unobscured features in the second subset of images. In some cases, this may result in faster signature generation, because image processing by the signature generation module 314 may be limited to effectively smaller images defined by the masks.
[0079] The electronic signature can be generated based on two or more images of the second subset of images. For example, in some implementations, respective electronic signatures are generated for each of the two or more images, and the resulting two or more electronic signatures are fused to form the electronic signature. In some implementations, the electronic signature is generated based on five, six, or seven images in the second subset of images. The use of multiple images for generating the electronic signature may improve the reliability of signature generation, e.g., because undesired imaging artifacts that may remain after removal of corrupted areas may be effectively canceled out / de-emphasized through the fusion of signatures corresponding to different images. Moreover, in some implementations, different images of the second subset of images have different corruption masks, such that the use of multiple images allows for a more complete analysis of the candidate mark. However, in some implementations, the electronic signature can be generated based on a single image of the two or more images.
[0080] Referring again to FIG. 2A, the process 200 includes sending the electronic signature of the candidate mark for authentication (210). For example, the mobile device can send the generated electronic signature to a remote computing system (e.g., a cloud computing system) over a network such as the Internet, and the remote computing system can compare the electronic signature to one or more stored electronic signatures, e.g., to determine an authenticity or identity of an item on which the candidate mark is printed. In some cases, transmission of the electronic signature, as opposed to transmission of the two or more images based on which the electronic signature was generated, can allow the process 200 to be performed more quickly, because the electronic signature may have a much smaller fde size than a fde size of the two or more images.
[0081] For example, as shown in FIG. 2B, the process 200 can include receiving, at the mobile device, an authentication response for the candidate mark (212). The authentication response can be a determination based on the electronic signature, e.g., can be indicative of whether the candidate mark is determined to be authentic, whether the candidate mark is determined to match a target mark, and / or indicative of an identity of an object on which the candidate mark is printed. The authentication response can be received, for example, from the remote computing system to which the electronic signature is sent.
[0082] Based on the authentication response, the mobile device determines (214) whether the candidate mark is authentic, e.g., as directly indicated by the authentication response, and / or by performing analysis using the authentication response. If the candidate mark is not authentic (e.g., does not have an electronic signature matching the electronic signature of a target mark), in some implementations, the mobile device displays an indication that the candidate mark is not authentic (216), e.g., using a display of the mobile device to present graphics and / or text. If the candidate mark is authentic, in some implementations, the mobile device displays an indication that the mobile device is authentic (218). Based on these indications, a user of the mobile device can take appropriate action, e.g., discarding or retaining the object on which the candidate mark is printed.
[0083] FIG. 8 shows an example of a process 800 that can be performed using the system 300. Some implementations of the process 800 correspond to someimplementations of the process 200, and some or all portions of the process 800 can be performed as described for portions of the process 200. The process 800 can be performed.
[0084] The process 800 includes capturing one or more images using a mobile device (802), e.g., as described with respect to process 202. An orientation model detects the orientation and position of a candidate mark in the images (804). For example, the orientation model can be orientation model 306 and can process the images as described with respect to orientation model 306 and process 204. For example, the orientation model can determine a segmentation of the images into portions that correspond to the candidate mark and portions that do not correspond to the candidate mark (e.g., using a suitable object recognition method) and can classify the orientation of the candidate mark as “correct orientation” or one or more other types, such as “tilted” and / or “inverted.”
[0085] In response to the orientation model indicating that the candidate mark has been imaged with a correct orientation, a dimensions model determines dimensions of the candidate mark (806). For example, the dimensions model can be dimensions model 308 and can process the images as described with respect to dimensions model 308. In some implementations (e.g., in some implementations of the processes 200, 800), the dimensions model executes independently of outputs of the orientation model, e.g., at any time during the process and / or without requiring that the candidate mark be correctly oriented.
[0086] In addition, a corruption model detects a level of corruption within cropped images (808), e.g., images cropped to exclude portions that do not correspond to the candidate mark, based on segmentation by the orientation model. For example, the corruption model can be corruption model 310 and can process the images as described with respect to corruption model 310 and process 204. In some implementations (e.g., in some implementations of the processes 200, 800), the corruption model executes in response to the orientation model indicating that the candidate mark has been imaged with a correct orientation.
[0087] Instructions can be provided to improve image suitability based on outputs of one or more of the orientation model, dimensions model, and / or corruption model, e.g., as described with respect to process 206 and operations of the instruction module 312.For example, based on the orientation model 306 detecting that the candidate mark is incorrectly oriented, instructions can be provided (812) for moving the mobile device in space to obtain a target orientation. Based on spatial dimensions estimated by the dimensions model, instructions can be provided (814) for moving the mobile device in space, and / or adjusting camera settings, to improve a focus of the candidate mark in the images. Based on corruption detected by the corruption model, instructions can be provided (816) for moving the mobile device in space, adjusting camera settings, and / or adjusting environmental conditions to reduce corruption.
[0088] In response to the level of corruption being less than a threshold, an electronic signature is generated and sent for authentication and / or other purposes (810). For example, generating the electronic signature can include capturing high-resolution images (e g., higher-resolution images than those evaluated in processes 804, 806, 808) and encoding characteristics of the candidate mark in two or more of the high-resolution images into a hash.
[0089] In some implementations, the corruption threshold based on which signature generation is performed is adjustable. For example, the acceptable level of corruption can depend on the ability of the signature generation module 314 to generate signatures with more or less available information (e.g., based on algorithms applied by the signature generation module 314). When less information is required from the two or more images based on which the signature is generated, a higher level of corruption may be acceptable.
[0090] For example, in some implementations, the corruption threshold can be determined in a testing procedure using the corruption model 310 and the signature generation module 314. The corruption model 310 is used to process images of marks to determine respective levels of corruption in the images, where the levels of corruption vary between the images. In addition, the signature generation module 314 processes the images to generate electronic signatures based on the marks. The generated electronic signatures are evaluated for accuracy in matching target electronic signatures for each mark. Because higher corruption is associated with less available data based on which to generate the signatures, the accuracy is expected to decrease with increasing corruption. As such, the corruption threshold can be determined based on the level of success in generating the signatures, e g., such that corruption levels at and below the thresholdprovide a satisfactory level of success. For example, if signatures generated with 10% corruption have a high rate of success in matching expected signatures, while signatures generated with 15% corruption tend to not match expected signatures, the threshold can be set at 10%. This process can be performed periodically to adjust the threshold based on improvements in the signature generation module 314, e.g., to increase the threshold as the signature generation module 314 becomes more capable. Advantageously, the threshold can be adjusted without having to retrain the corruption model 310, providing a rapid and easy means of tuning the processes 200, 800.
[0091] In some implementations, a feedback loop is employed for periodic or continuous adjustment of the corruption threshold. As electronic signatures are generated during repeated performance of the processes 200, 800 or related processes, the electronic signatures tend to be determined correctly at a high rate, the threshold can be decreased, to attempt to generate electronic signatures based on images with more corruption. If electronic signatures tend to be determined incorrectly at a high rate, the threshold can be increased, to enforce improved capture quality in subsequent processing.
[0092] In some implementations, the corruption threshold is adjusted based on a characteristic of the candidate mark and / or an item on which the candidate mark is printed. For example, different types of candidate mark, and / or different types of material on which the candidate mark is printed, can be associated with different thresholds. For example, glossier substrates can be associated with lower thresholds than more matte substrates. The machine learning module 316, the instruction module 312, the signature generation module 314, or another computing module can be configured to determine the corruption threshold based on the candidate mark and / or a characteristic of the item on which the candidate mark is printed.
[0093] As shown in FIG. 3B, the signature generation module 314 can receive, as an input, the corruption threshold or data indicative of the corruption threshold, such that the corruption threshold is remotely adjustable. For example, the corruption threshold can be received and / or remotely adjusted from / by the remote computing system to which the electronic signature is sent. Accordingly, as the signature authentication process improves at the remote computing system, it may not be necessary to recompile and reinstall software (e.g., modules 316, 312, and 304) that runs on the mobile device. Rather, theremote system may simply change an input to the signature generation module 314 that controls the level of corruption that is considered acceptable by the already-installed software on the mobile device. This arrangement may decrease computational resource usage and / or bandwidth consumption, and / or provide improved ease of use and maintenance, compared to requiring modification of software installed on the mobile device.
[0094] As shown in FIG. 9, the machine learning models 306, 308, 310 can be trained using labeled images of marks. In a process 900, training data 902 includes images of different types of marks (e.g., Universal Product Code (UPC), ID barcodes, QR codes, 2D barcodes, Data Matrix codes, etc.) captured under different capture conditions, for example, with different relative camera angles, camera-mark distances, and lighting conditions (e.g., bright sunlight, natural light, uneven shadows, spot glare, and / or diffuse glare, corresponding to varying types and levels of corruption in the images). In some implementations, the images in the training data 902 are low resolution images as discussed above, e.g., low resolution images obtained by image resizing, downscaling, and / or truncation. In some implementations, images in an initial dataset are assessed for similarity, and duplicative images (e.g., having a calculated similarity above a threshold value) are not used for training, e.g., are excluded from the training data 902. The similarity can be determined, for example, by a cosine similarity, structural similarity (SSIM), and / or mean-squared error method. For example, only a single image can be selected from each set of images in the initial dataset that are too-similar to one another. As a result, the training can be performed using images with reasonably high variance between one another, improving training robustness.
[0095] Labels 904 for the training can indicate, for the images of the training data, annotated locations of the marks in the images, annotated orientations of the marks, annotated types and / or locations of corruption in the images, and spatial dimensions of the marks. In a training process (906), based on the training data 902 and the labels 904 of the training data 902, one or more weights and / or other parameters of the models 306, 308, 310 are adjusted, to obtain trained models 908. Although FIG. 9 shows a combined process for training the models 306, 308, 310, in some implementations the models 306, 308, 310 can be trained in separate processes. For example, in each of the separateprocesses, the labels 904 can be appropriate for the specific model being trained, e.g., corruption annotations for training the corruption model 310.[0961 Because of the use of multiple types of marks in the training data 902, the models 306, 308, 310 can be trained to flexible detect and analyze the multiple types of marks. The models 306, 308, 310 can subsequently be retrained to learn to process additional types of marks, if desired.
[0097] Types of machine learning models within the scope of this disclosure include, for example, machine learning models that implement supervised, semi-supervised, unsupervised and / or reinforcement learning; neural networks, including deep neural networks, autoencoders, convolution neural networks, multi-layer perceptron networks, and recurrent neural networks; classification models; and regression models. The machine learning models described herein can be configured with one or more approaches, such as back-propagation, gradient boosted trees, decision trees, support vector machines, reinforcement learning, partially observable Markov decision processes (POMDP), and / or table-based approximation, to provide several non-limiting examples. Based on the type of machine learning model, the training 906 can include adjustment of one or more parameters. For example, in the case of a regression-based model, the training 906 can include adjusting one or more coefficients of the regression so as to minimize a loss function such as a least-squares loss function. In the case of a neural network, the training 906 can include adjusting weights, biases, number of epochs, batch size, number of layers, and / or number of nodes in each layer of the neural network, so as to minimize a loss function.
[0098] In some implementations, one or more of the machine learning models 306, 308, 310 is specially trained and / or configured for low-latency execution on mobile devices, which may have relatively low amounts of computational resources available. For example, in some implementations, the machine learning models are quantized so as to have reduced model sizes compared to un-quantized models, providing improved compatibility with the processing and / or memory limits of mobile devices. The machine learning module 316 can store metadata that maps outputs of the corruption model 310 to types of corruption, such as shadow, diffuse glare, and spot glare, a computational configuration that can be useful for execution on mobile devices. For example, thecorruption model 310, in some implementations, is initially trained to identify the different types of corruption; when the corruption model 310 is then quantized for execution on mobile devices, inference for these categories may be not maintained exactly. Metadata can be used to map outputs of the quantized model (e.g., numerical outputs) to the categories, obtaining corruption model outputs that can be used for overlays, for instructions, etc. in a manner that users can understand. In some implementations, one or more of the machine learning models 306, 308, 310 has an architecture that provides efficient performance on mobile devices, such as an EfficientNet architecture or a MobileNet architecture. The models 306, 308, 310 can be deployed using TENSORFLOW LITE™ to provide efficient inference on mobile devices, and the previously-described metadata can be metadata of the TENSORFLOW LITE ™ deployment.
[0099] Accordingly, based on the foregoing systems and processes, electronic signatures can be generated quickly and with high accuracy. These systems and processes can account for various complex factors associated with mark imaging, such as mark orientation, focus, and corruption, and (i) provide real-time feedback to improve imaging, and (ii) adapt to these factors for signature generation, such as by avoiding processing of corrupted areas using a corruption mask. Because, in some implementations, the systems and processes are compatible with performance by a mobile device, a dependency on cloud computing can be reduced, improving computational efficiency and reducing latency compared to processes that have to be performed using remote computing systems. For example, overlays displayed on live images according to this disclosure can be generated and updated in real-time using local processing rather than relying on remote processing, which may be associated with delayed and / or inaccurate overlays.
[0100] Some features described herein, such as the system 300 and elements thereof, may be implemented in digital and / or analog electronic circuitry or in computer hardware, firmware, software, or in combinations of them. Some features may be implemented in a computer program product tangibly embodied in an information carrier, e.g., in a machine-readable storage device, for execution by a programmable processor. Method steps may be performed by a programmable processor executing a program of instructions to perform functions of the described implementations by operating on inputdata and generating output, by discrete circuitry performing analog and / or digital circuit operations, or by a combination thereof.
[0101] Some described features may be implemented advantageously in one or more computer programs that are executable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that may be used, directly or indirectly, in a computer to perform a certain activity or bring about a certain result. A computer program may be written in any form of programming language (e.g., Objective- C, Java, Python, JavaScript, Swift), including compiled or interpreted languages, and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0102] Suitable processors for the execution of a program of instructions include, by way of example, both general and special purpose microprocessors, and the sole processor or one of multiple processors or cores, of any kind of computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer may communicate with mass storage devices for storing data fdes. These mass storage devices may include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including by way of example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, ASICs (application-specific integrated circuits). To provide for interaction with a user the features may be implemented on a computer having a display device such as a CRT (cathode ray tube), LED (light emitting diode) or LCD (liquid crystal display) display or monitor for displaying information to the author, a keyboard and a pointing device, such as a mouse or a trackball by which the author may provide input to the computer.
[0103] Examples within the scope of this disclosure include, but are not limited to, the following:
[0104] Example 1 : A method including: capturing, at a mobile device, multiple images using a camera of the mobile device; processing, at the mobile device, a first subset of the multiple images at a first resolution using two or more machine learning models, wherein a first of the two or more machine learning models has been trained to identify an orientation of a candidate mark in the first subset of the multiple images, and a second of the two or more machine learning models has been trained to identify an amount of corruption in an image of the candidate mark; providing, at the mobile device, one or more instructions to modify the capturing of a second subset of the multiple images based on a first output of the first model trained to identify the orientation of the candidate mark, and based on a second output of the second model trained to identify the amount of corruption in the image of the candidate mark; generating, at the mobile device, an electronic signature of the candidate mark from two or more images of the second subset of the multiple images, at a second resolution that is higher than the first resolution, in response to both the first output indicating the orientation is suitable for electronic signature generation and the second output indicating the amount of corruption is satisfies a threshold condition; and sending the electronic signature of the candidate mark for authentication.
[0105] Example 2: The method of Example 1, wherein the first subset of the multiple images is captured at the first resolution, and wherein the second subset of the multiple images is captured at the second resolution.
[0106] Example 3: The method of any one of Example 1-2, wherein the two or more machine learning models include three machine learning models including the first model trained to identify the orientation of the candidate mark, the second model trained to identify the amount of corruption in the image of the candidate mark, and a third model trained to identify spatial dimensions of the candidate mark.
[0107] Example 4: The method of Example 3, wherein the providing includes: displaying, on the mobile device, at least one instruction to a user of the mobile device to adjust a distance between the mobile device and the candidate mark, based on the identified spatial dimensions of the candidate mark.
[0108] Example 5: The method of any one of Examples 1-4, wherein the providing includes: displaying, on the mobile device, at least one instruction to a user of the mobile device regarding how to move the mobile device in three-dimensional space to improve image suitability, reduce image corruption, or both.
[0109] Example 6: The method of any one of Examples 1-5, wherein the providing includes: sending, to the camera of the mobile device, at least one instruction that adjusts at least one setting of the camera to improve image suitability, reduce image corruption, or both.
[0110] Example 7: The method of any one of Examples 1-6, wherein processing the first subset of the multiple images using the two or more machine learning models includes: identifying, using the first model, an area of a first image of the first subset of the multiple images, the area corresponding to the candidate mark; cropping the first image to obtain a cropped image that includes the identified area and excludes at least another area of the first image; and providing, as input to the second model, the cropped image.
[0111] Example 8: The method of any one of Examples 1-7, wherein generating the electronic signature includes: processing a first image of the two or more images using the second model to generate a corruption mask defining a corrupted area in the first image; and producing the electronic signature using a masked version of the first image based on the corruption mask, the masked version excluding the corrupted area.
[0112] Example 9: The method of any one of Examples 1-8, wherein the second of the two or more machine learning models has been trained to identify at least one of the following in the first subset of the multiple images: an area exhibiting glare, or an area that is shadowed.
[0113] Example 10: The method of any one of Examples 1-9, wherein the first of the two or more machine learning models has been trained to determine whether the candidate mark is tilted with respect to a target orientation.
[0114] Example 11 : The method of any one of Examples 1-10, wherein the second of the two or more machine learning models has been trained to identify a corrupted area in the image, and wherein providing the one or more instructions includes displaying an overlay on a live image of the candidate mark, the overlay covering a corrupted area in the live image.
[0115] Example 12: The method of Example 11, wherein a color of the overlay is indicative of a type of corruption in the corrupted area.
[0116] Example 13: The method of any one of Examples 1-12, wherein the candidate mark includes a barcode.
[0117] Example 14: The method of any one of Examples 1-13, wherein generating the electronic signature of the candidate mark is based on printing artifacts in the candidate mark.
[0118] Example 15: The method of any one of Examples 1-14, including training the two or more machine learning models using, as training data, a plurality of images of different printed marks, the plurality of images captured under different lighting conditions, from different perspectives, and with different levels of image corruption.
[0119] Example 16: The method of any one of Examples 1-15, wherein the threshold condition includes that the level of corruption is less than a threshold, and wherein the method includes: using the second model to determine levels of corruption in a plurality of images; generating electronic signatures of candidate marks in the plurality of images; and, based on a level of success in generating the electronic signatures, and based on the levels of corruption in the plurality of images, adjusting the threshold.
[0120] Example 17: The method of any one of Examples 1-16, wherein the amount of corruption includes a proportion of an area of the candidate mark that is corrupted.
[0121] Example 18: The method of any one of Examples 1-17, including: receiving, at the mobile device, an authentication response for the candidate mark; determining, by the mobile device, based on the authentication response, whether the candidate mark is authentic; and displaying, by the mobile device, an indication of the authenticity of the candidate mark.
[0122] Example 19. A non-transitory computer-readable medium tangibly encoding a computer program operable to cause data processing apparatus to perform operations of any of Examples 1-18.
[0123] Examples 20. A system including: one more computers programmed to authenticate electronic signatures of candidate marks; and a mobile device communicatively coupled with the one or more computers, the mobile device being programmed to perform operations of any of Examples 1-18.
[0124] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made. Elements of one or more implementations may be combined, deleted, modified, or supplemented to form further implementations. In yet another example, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. For example, processing by the models 306, 308, 310 can be sequential and / or parallel, in different implementations. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.
Claims
What is claimed is:
1. A method comprising: capturing, at a mobile device, multiple images using a camera of the mobile device; processing, at the mobile device, a first subset of the multiple images at a first resolution using two or more machine learning models, wherein a first of the two or more machine learning models has been trained to identify an orientation of a candidate mark in the first subset of the multiple images, and a second of the two or more machine learning models has been trained to identify an amount of corruption in an image of the candidate mark; providing, at the mobile device, one or more instructions to modify the capturing of a second subset of the multiple images based on a first output of the first model trained to identify the orientation of the candidate mark, and based on a second output of the second model trained to identify the amount of corruption in the image of the candidate mark; generating, at the mobile device, an electronic signature of the candidate mark from two or more images of the second subset of the multiple images, at a second resolution that is higher than the first resolution, in response to both the first output indicating the orientation is suitable for electronic signature generation and the second output indicating the amount of corruption is satisfies a threshold condition; and sending the electronic signature of the candidate mark for authentication.
2. The method of claim 1, wherein the first subset of the multiple images is captured at the first resolution, and wherein the second subset of the multiple images is captured at the second resolution.
3. The method of claim 1, wherein the two or more machine learning models comprise three machine learning models comprising the first model trained to identify the orientation of the candidate mark, the second model trained to identify the amount ofcorruption in the image of the candidate mark, and a third model trained to identify spatial dimensions of the candidate mark.
4. The method of claim 3, wherein the providing comprises: displaying, on the mobile device, at least one instruction to a user of the mobile device to adjust a distance between the mobile device and the candidate mark, based on the identified spatial dimensions of the candidate mark.
5. The method of claim 1, wherein the providing comprises: displaying, on the mobile device, at least one instruction to a user of the mobile device regarding how to move the mobile device in three-dimensional space to improve image suitability, reduce image corruption, or both.
6. The method of claim 1, wherein the providing comprises: sending, to the camera of the mobile device, at least one instruction that adjusts at least one setting of the camera to improve image suitability, reduce image corruption, or both.
7. The method of claim 1, wherein processing the first subset of the multiple images using the two or more machine learning models comprises: identifying, using the first model, an area of a first image of the first subset of the multiple images, the area corresponding to the candidate mark; cropping the first image to obtain a cropped image that includes the identified area and excludes at least another area of the first image; and providing, as input to the second model, the cropped image.
8. The method of claim 1, wherein generating the electronic signature comprises: processing a first image of the two or more images using the second model to generate a corruption mask defining a corrupted area in the first image; and producing the electronic signature using a masked version of the first image based on the corruption mask, the masked version excluding the corrupted area.
9. The method of claim 1, wherein the second of the two or more machine learning models has been trained to identify at least one of the following in the first subset of the multiple images: an area exhibiting glare, or an area that is shadowed.
10. The method of claim 1, wherein the first of the two or more machine learning models has been trained to determine whether the candidate mark is tilted with respect to a target orientation.
11. The method of claim 1, wherein the second of the two or more machine learning models has been trained to identify a corrupted area in the image, and wherein providing the one or more instructions comprises displaying an overlay on a live image of the candidate mark, the overlay covering a corrupted area in the live image.
12. The method of claim 11, wherein a color of the overlay is indicative of a type of corruption in the corrupted area.
13. The method of claim 1, wherein the candidate mark comprises a barcode.
14. The method of claim 1, wherein generating the electronic signature of the candidate mark is based on printing artifacts in the candidate mark.
15. The method of claim 1, comprising training the two or more machine learning models using, as training data, a plurality of images of different printed marks, the plurality of images captured under different lighting conditions, from different perspectives, and with different levels of image corruption.
16. The method of claim 1 , wherein the threshold condition comprises that the level of corruption is less than a threshold, wherein the method comprises: using the second model to determine levels of corruption in a plurality of images; generating electronic signatures of candidate marks in the plurality of images; and based on a level of success in generating the electronic signatures, and based on the levels of corruption in the plurality of images, adjusting the threshold.
17. The method of claim 1, wherein the amount of corruption comprises a proportion of an area of the candidate mark that is corrupted.
18. The method of claim 1, comprising: receiving, at the mobile device, an authentication response for the candidate mark; determining, by the mobile device, based on the authentication response, whether the candidate mark is authentic; and displaying, by the mobile device, an indication of the authenticity of the candidate mark.
19. A non-transitory computer-readable medium tangibly encoding a computer program operable to cause data processing apparatus to perform operations of any of claims 1-18.
20. A system comprising: one more computers programmed to authenticate electronic signatures of candidate marks; and a mobile device communicatively coupled with the one or more computers, the mobile device being programmed to perform operations of any of claims 1-18.