Ai-powered visual code detection and optimization engine

An AI-driven system optimizes camera parameters using neural networks to enhance visual code detection by calibrating zoom and sharpness settings and identifying regions of interest, addressing readability challenges in dynamic environments for improved code detection and data retrieval.

WO2025196756A1PCT designated stage Publication Date: 2025-09-25SODYO
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/IL2025/050254
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-19
Filing Date
2025-03-18
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing visual code scanning technologies struggle with readability challenges due to unfavorable lighting conditions, motion blur, and suboptimal camera focus, particularly in dynamic environments, leading to inefficiencies in code detection and data retrieval.

Method used

An AI-driven system utilizing machine learning to optimize camera parameters through neural networks for enhanced visual code detection, including calibrating zoom and sharpness settings, identifying regions of interest, and dynamically adjusting focus and exposure to improve code readability.

Benefits of technology

The system ensures high-speed, high-accuracy visual code detection in diverse conditions by minimizing digital zoom, optimizing optical zoom, and enhancing image clarity, thereby improving code readability and retrieval efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2025050254_25092025_PF_FP_ABST
    Figure IL2025050254_25092025_PF_FP_ABST
Patent Text Reader

Abstract

A computer implemented Al system for reading a visual code, including a camera, an imaging device having an optical zoom range extending from a minimal to a maximal optical magnification, and a processor which iteratively operates the device at a first zoom setting to capture a first image at a first magnification, and at a second zoom setting to capture a second image at a second magnification by applying both the maximal optical and a digital magnification to the second image. The processor calculates a ratio of the second to the first magnification, analyzes the second image to estimate the second image digital magnification, calculates a correction factor for selecting a target zoom setting, and uses the iterations to train the neural network to provide the target zoom setting. The processor operates the camera to capture a scene image containing the visual code, and reads the visual code.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AI-PQWERED VISUAL CODE DETECTION AND OPTIMIZATION ENGINE

[0002] CROSS-REFERENCE TO RELATED APPLICATION

[0003] This application claims the benefit of U.S. Provisional Patent Application 63 / 566,955, filed March 19, 2024, which is incorporated herein by reference.

[0004] FIELD OF THE INVENTION

[0005] This invention relates generally to Al-driven visual code scanning, and specifically to optimizing parameters of an imaging system through machine learning so as to enhance the readability of visual codes.

[0006] BACKGROUND OF THE INVENTION

[0007] Visual codes, such as quick response (QR) or color based codes, are increasingly being used in a variety of scenarios. The codes, which were initially developed for tracking automobile parts, are today widely used across multiple industries, including:

[0008] • Television broadcasting for interactive viewer engagement, second-screen experiences, and real-time call-to-action triggers.

[0009] • Retail and advertising for seamless product discovery, customer engagement, and conversion tracking.

[0010] • Logistics and supply chain management for tracking shipments, inventory management, and real-time scanning of goods across distribution networks.

[0011] While visual codes offer highly efficient, contactless data retrieval, readability challenges arise due to factors such as unfavorable lighting conditions, motion blur, poor contrast, or suboptimal camera focus. Consequently, existing device algorithms (autofocus, auto exposure, auto white balance) struggle to optimize visual code scanning in dynamic environments.

[0012] SUMMARY OF THE INVENTION

[0013] Embodiments of the present invention leverage artificial intelligence (Al) to enhance scanning performance across different use cases, ensuring high-speed, high- accuracy code detection in diverse real- world conditions. The embodiments employ Al-based image analysis to optimize camera parameters for enhanced visual code detection and readability. A disclosed embodiment consists of two primary Al-powered components:

[0014] 1. Al- Driven Device Optimization

[0015] A machine learning model trains a neural network to be used to calibrates a target device’s camera to determine the optimal zoom factor and sharpness settings for scanning visual codes.

[0016] The model minimizes reliance on digital zoom, which can degrade image quality, and prioritizes optical zoom to enhance clarity and ensure maximum readability.

[0017] The model assesses image sharpness to prevent excessive enhancement artifacts that may obscure the visual code.

[0018] 2. Al-Powered Scene Analysis and Code Detection

[0019] A neural network is trained to identify a Region of Interest (ROI) containing a visual code within a complex scene, such as a TV screen, a billboard, a warehouse rack, or a moving shipping container.

[0020] Using a result from the trained network allows a smartphone to dynamically adjust focus, exposure, and white balance for optimal code readability, rather than applying general adjustments to the entire image.

[0021] An embodiment of the present invention provides a computer implemented, artificial intelligence (Al), system for reading a visual code, including: a camera; an imaging device, having an optical zoom range extending from a minimal optical magnification to a maximal optical magnification; and a processor, configured to perform a plurality of iterations on the imaging device, so that at each iteration the processor: operates the device at a first zoom setting within the optical zoom range so as to capture a first image at a first magnification, and at a second zoom setting so as to capture a second image at a second magnification, greater than the maximal optical magnification, by applying both the maximal optical magnification and a digital magnification to the second image, calculates a ratio of the second magnification to the first magnification, analyzes the second image to estimate the digital magnification of the second zoom setting, in response to the estimated digital magnification, calculates a correction factor for selecting a target zoom setting, and inputs to a neural network the first image, the second image, the ratio, and the correction factor formed in each of the iterations so as to train the neural network to provide the target zoom setting; and wherein the processor is further configured: to operate the camera at the target zoom setting to capture a scene image containing the visual code; to identify the visual code in the scene image; and to read the identified visual code.

[0022] The processor may be configured to receive from the camera, prior to operating the camera at the target zoom setting, a first camera image having a first camera zoom setting and a second camera image having a second camera zoom setting greater than the first camera zoom setting, to provide the first camera image and the second camera image to the neural network, and in response to receive from the neural network a target correction factor, and apply the target correction factor to the camera to determine the target zoom setting.

[0023] In a disclosed embodiment the second zoom setting includes a second FOV setting, and the correction factor is a function of the second FOV setting.

[0024] In a further disclosed embodiment the imaging device has a digital zoom range, and the correction factor is a function of a digital component of the digital zoom range.

[0025] In a yet further disclosed embodiment analyzing the second image includes at least one of receiving a manual estimation of the digital magnification of the second image and measuring a modulation transfer function of the second image.

[0026] There is further provided, according to an embodiment of the present invention, a computer implemented system for setting a sharpness metric for an image of a visual code, including: a camera; an imaging device, having a sharpness setting range extending from a minimal sharpness setting to a maximal sharpness setting; and a processor, configured to perform a plurality of iterations on the imaging device, so that at each iteration the processor: operates the device at the maximal sharpness setting so as to capture an image at the maximal sharpness setting, formulates a sharpness correction factor, in response to a difference between the maximal sharpness setting and a preset optimal sharpness setting, for selecting a target sharpness setting, associates the sharpness correction factor with the image, inputs to a neural network the image and the correction factor formed in each of the iterations so as to train the neural network to provide the target sharpness setting; and wherein the processor is further configured: to operate the camera at the target sharpness setting to capture a scene image containing the visual code; to identify the visual code in the scene image; and to read the identified visual code.

[0027] The processor may be configured to receive from the camera, prior to operating the camera at the target sharpness setting, a camera image having a maximal camera sharpness setting, to provide the camera image to the neural network, and in response to receive from the neural network a target correction factor, and apply the target correction factor to the camera to determine the target sharpness setting.

[0028] There is further provided, according to an embodiment of the present invention, a computer implemented system for identifying a location of a region of interest (RO I) in an image, including: a camera; an imaging device; and a processor, configured to perform a plurality of iterations on the imaging device, so that at each iteration the processor: operates the device so as to capture an image of the ROI in a scene, the scene including a scene descriptor classifying the scene, identifies a location of the ROI within the scene, associates the location of the ROI and the scene descriptor with the image, inputs to a neural network the image and the associated ROI location and scene descriptor formed in each of the iterations so as to train the neural network to provide a target ROI location; and wherein the processor is further configured: to operate the camera to capture a camera scene image containing a target ROI; and to identify the target ROI location in the camera scene image.

[0029] The processor may be configured to receive from the camera the camera scene image, to provide the camera scene image to the neural network, and in response to receive from the neural network the target ROI location.

[0030] There is further provided, according to an embodiment of the present invention, a system for reading a visual code, including: a camera, having an optical zoom range extending from a minimal optical magnification to a maximal optical magnification; and a processor, configured: to operate the camera at a first zoom setting within the optical zoom range so as to capture a first image at a first magnification, and at a second zoom setting so as to capture a second image at a second magnification, greater than the maximal optical magnification, by applying both the maximal optical magnification and a digital magnification to the second image, to analyze a sharpness metric of the second image to estimate the digital magnification of the second zoom setting; in response to the estimated digital magnification, to select a target zoom setting within the optical zoom range; to operate the camera at the target zoom setting to capture a scene image containing the visual code; to identify the visual code in the scene image; and to read the identified visual code. In an alternative embodiment the first zoom setting includes a first field of view (FOV) setting and the second zoom setting includes a second FOV setting.

[0031] The first FOV setting may be an inverse of the first magnification, and the second FOV setting may be the inverse of the second magnification.

[0032] In a further alternative embodiment the camera has a digital zoom range, and the target zoom setting is at a boundary between the optical zoom range and the digital zoom range.

[0033] The sharpness metric may include a modulation transfer function.

[0034] There is further provided, according to an embodiment of the present invention, a system for determining a sharpness setting for imaging a visual code, including: an imaging device configured to image the visual code; and a processor, configured to: receive an optimal sharpness setting for reading the visual code; operate the imaging device at a maximal sharpness setting greater than the optimal sharpness setting; formulate a sharpness correction factor in response to the optimal sharpness setting and the maximal sharpness setting; and apply the sharpness correction factor to the imaging device to provide a target sharpness setting for the imaging device, and operate the imaging device at the target sharpness setting when imaging the visual code.

[0035] There is further provided, according to an embodiment of the present invention, a computer implemented, artificial intelligence (Al), method for reading a visual code, including: providing an imaging device with an optical zoom range extending from a minimal optical magnification to a maximal optical magnification; configuring a processor to perform a plurality of iterations on the imaging device, so that at each iteration the processor: operates the device at a first zoom setting within the optical zoom range so as to capture a first image at a first magnification, and at a second zoom setting so as to capture a second image at a second magnification, greater than the maximal optical magnification, by applying both the maximal optical magnification and a digital magnification to the second image, calculates a ratio of the second magnification to the first magnification, analyzes the second image to estimate the digital magnification of the second zoom setting, in response to the estimated digital magnification, calculates a correction factor for selecting a target zoom setting, and inputs to a neural network the first image, the second image, the ratio, and the correction factor formed in each of the iterations so as to train the neural network to provide the target zoom setting; and further configuring the processor: to operate a camera at the target zoom setting to capture a scene image containing the visual code; to identify the visual code in the scene image; and to read the identified visual code.

[0036] There is further provided, according to an embodiment of the present invention, a computer implemented method for setting a sharpness metric for an image of a visual code, including: providing an imaging device, having a sharpness setting range extending from a minimal sharpness setting to a maximal sharpness setting; and configuring a processor to perform a plurality of iterations on the imaging device, so that at each iteration the processor: operates the device at the maximal sharpness setting so as to capture an image at the maximal sharpness setting, formulates a sharpness correction factor, in response to a difference between the maximal sharpness setting and a preset optimal sharpness setting, for selecting a target sharpness setting, associates the sharpness correction factor with the image, inputs to a neural network the image and the correction factor formed in each of the iterations so as to train the neural network to provide the target sharpness setting; and further configuring the processor: to operate a camera at the target sharpness setting to capture a scene image containing the visual code; to identify the visual code in the scene image; and to read the identified visual code.

[0037] There is further provided, according to an embodiment of the present invention, a computer implemented method for identifying a location of a region of interest (RO I) in an image, including: providing an imaging device; and configuring a processor to perform a plurality of iterations on the imaging device, so that at each iteration the processor: operates the device so as to capture an image of the ROI in a scene, the scene including a scene descriptor classifying the scene, identifies a location of the ROI within the scene, associates the location of the ROI and the scene descriptor with the image, inputs to a neural network the image and the associated ROI location and scene descriptor formed in each of the iterations so as to train the neural network to provide a target ROI location; and further configuring the processor: to operate a camera to capture a camera scene image containing a target ROI; and to identify the target ROI location in the camera scene image.

[0038] There is further provided, according to an embodiment of the present invention, a method for reading a visual code, including: providing a camera with an optical zoom range extending from a minimal optical magnification to a maximal optical magnification; and operating the camera at a first zoom setting within the optical zoom range so as to capture a first image at a first magnification, and at a second zoom setting so as to capture a second image at a second magnification, greater than the maximal optical magnification, by applying both the maximal optical magnification and a digital magnification to the second image, analyzing a sharpness metric of the second image to estimate the digital magnification of the second zoom setting; in response to the estimated digital magnification, selecting a target zoom setting within the optical zoom range; operating the camera at the target zoom setting to capture a scene image containing the visual code; identifying the visual code in the scene image; and reading the identified visual code.

[0039] There is further provided, according to an embodiment of the present invention, a method for determining a sharpness setting for imaging a visual code, including: configuring an imaging device to image the visual code; and configuring a processor to: receive an optimal sharpness setting for reading the visual code, operate the imaging device at a maximal sharpness setting greater than the optimal sharpness setting, formulate a sharpness correction factor in response to the optimal sharpness setting and the maximal sharpness setting, and apply the sharpness correction factor to the imaging device to provide a target sharpness setting for the imaging device, and operate the imaging device at the target sharpness setting when imaging the visual code.

[0040] The present disclosure will be more fully understood from the following detailed description of the embodiments thereof, taken together with the drawings, in which:

[0041] BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Fig. 1 is a schematic illustration of a device being used to image and read a visual code, according to an embodiment of the present invention;

[0043] Fig. 2 illustrates the effect of zooming on the sharpness and readability of an image, according to an embodiment of the present invention; Fig. 3 illustrates the effect of different sharpness settings on an image, according to an embodiment of the present invention;

[0044] Fig. 4 is a flowchart of steps describing how an optimal zoom setting is provided to a smartphone, according to an embodiment of the present invention;

[0045] Fig. 5 is a graph illustrating some of the steps of the flowchart, according to an embodiment of the present invention;

[0046] Fig. 6 is a flowchart of steps for providing a sharpness setting to a smartphone, according to an embodiment of the present invention; and

[0047] Figs. 7A, 7B, and 7C are examples of images with different exposure times, according to an embodiment of the present invention; and

[0048] Fig. 8 is a flowchart describing steps for training and using a neural network to identify a region of interest for a smartphone, according to an embodiment of the present invention.

[0049] DETAILED DESCRIPTION OF EMBODIMENTS

[0050] Overview

[0051] A device used to capture and read a visual code typically has many device parameters used by a device processor to perform the capture and reading of an image including the visual code. Some of the parameters, e.g., the focal length of the lens of each device camera and the aperture (f-number) of each camera lens are fixed for a device. Other parameters, for example the fractions of optical and digital zoom applied, the values of the optical zoom and the digital zoom (within predetermined ranges), the focusing (lens - image sensor distance), the exposure time and the white balance applied to the image, are varied, but may not be directly set by a device user. These parameters may be set by algorithms of the device, and include, for example, autofocus, autoexposure, and auto white balance algorithms. Overall these parameters are herein termed non-user-adjustable device parameters, and the non-user-adjustable parameters are sub-divided into fixed parameters and algorithm controlled parameters.

[0052] Parameters of the device other than the non-user-adjustable parameters may be directly altered by the device user before an image is captured. Such parameters, herein termed user-adjustable parameters, include, for example, a zoom factor that the user can change, a sharpness of all or part of a scene, and whether flash is to be used. The devices used to acquire the visual code come as a multitude of platforms, for example, smartphones, computer laptops, tablets, logistics scanners, and smart cameras. For simplicity, except where otherwise stated, in the present description a device is assumed to comprise a smartphone.

[0053] The non-user-adjustable smartphone parameters are typically optimized for portraiture and / or landscape imaging. In many cases the non-user-adjustable parameters are sub-optimal for visual code imaging, which relies on good, minimally embellished, imaging to ensure correct reading of the code. The visual code imaging may be further impaired by the presence of extraneous factors, such as presenting the code on a TV screen surrounded by other images, or as a small entity in a large field of view being imaged.

[0054] Embodiments of the present invention address the problems of visual code imaging described above using two components to read the code. Both components use machine learning applied to corpuses of images to enhance the efficiency of the component.

[0055] For a first component, also herein termed the device component, a target smartphone (the phone intended to be used to read a target visual code) is calibrated, semi-automatically, to find an optimal zoom factor and an optimal sharpness setting.

[0056] The optimal zoom factor comprises an optimal optical image magnification which inherently also achieves a good sharpness of the image. The optimal sharpness setting enhances the sharpness provided by the optimal optical image magnification. The target smartphone then uses the optimal zoom factor and the optimal sharpness setting when imaging the target visual code.

[0057] To calibrate the target smartphone for the optimal zoom factor, a corpus of images from different devices is provided to a neural network. The corpus comprises pairs of images - each pair from a given device comprising a minimally zoomed image having a wide field of view (FOV) and a highly zoomed image having a narrow FOV. Embodiments of the invention assume that there is a linear relationship between the image zoom value and the reciprocal of the image FOV. While different devices typically use different product- specific descriptors for the zoomed images, embodiments of the present invention standardize the pairs across different devices by using the ratio of the FOVs of the pair (which corresponds to the reciprocal of the zoom ratios).

[0058] The corpus of images is annotated before being provided to the neural network. Each pair of images comprises a minimally zoomed image annotated with a known maximal FOV, a maximally zoomed image annotated with a known minimal FOV, and the pair is also annotated with the ratio of the FOVs. After training, the network is able to receive, from the target smartphone, a pair of images - one minimally zoomed, the other maximally zoomed, and estimate the ratio of the FOVs for the target smartphone.

[0059] Each zoomed image is formed as a combination of an optically zoomed component and a digitally zoomed component. From the point of view of reading a visual code, any digitally zoomed component present in an image being read may detract from the ability of a processor to read the code. In addition, the optimal image presented to the processor should have as large a magnification as possible. As explained below, the neural network is further trained to find, for each pair of images in the corpus, a zoom setting, between the minimum and maximum zoom settings of the pair, having a maximum magnification and a minimal digitally zoomed component.

[0060] For the further training, for each pair in the corpus the maximally zoomed image is analyzed to find the digital component fraction of the image. Using the fraction and the linear relationship between the reciprocal of the FOV and the zoom setting referred to above, an estimate is made for the FOV value of the optimal image (the image having maximum magnification and minimal digitally zoomed component. The FOV value of the optimal image lies “between” the two FOVs of the corpus pair and has a digital component, herein termed the minimal zoom fraction.

[0061] The ratio of the FOV value of the maximally zoomed image to the FOV value of the optimal image is an FOV factor which identifies the optimal image. Thus, in the further training, each maximally zoomed image is provided with an FOV factor identifying the FOV of the optimal image. After the further training, the network is able to formulate the factor for the maximally zoomed image received from the target smartphone and provide the factor to the target smartphone, so that the phone can use the factor to set the optimal zoom setting. As stated above, for the device component the target smartphone is also calibrated to find an optimal sharpness setting. A neural network is trained, using a corpus of images formed with maximal sharpness settings, and formed on different devices. Each image is manually annotated to find a sharpness factor to be applied to the maximal sharpness setting to achieve a pre-defined image sharpness standard. In an embodiment the standard is assumed to be a preset value of the modulation transfer function at 50% contrast (MTF50).

[0062] To find the optimal sharpness setting for the target smartphone, the phone provides an image with a maximal sharpness setting to the trained neural network. In response the network provides the sharpness factor to the target smartphone, which then applies the factor to set the optimal sharpness setting for the phone.

[0063] The second component for reading the visual code, also herein termed the scene component, is used to analyze the scene being imaged, to identify a region of interest (ROI) within the scene, and to provide the processor of the target smartphone with the ROI information. The processor then applies existing target smartphone algorithms, such as those referred to above, to set algorithm controlled parameters for the ROI when the optimally zoomed image is acquired.

[0064] On acquisition of an image using the optimal zoom value and the optimal sharpness setting, the processor of the target smartphone uses existing target smartphone algorithms to set algorithm controlled parameters of the smartphone, as is described above with reference to the scene component. As is also described above, within the optimally zoomed image the processor identifies the ROI, which comprises the visual code being imaged, and reads the code.

[0065] Detailed Description

[0066] In the following description, like elements in the drawings are identified by like numerals.

[0067] Reference is now made to Fig. 1, which is a schematic illustration of a device 20 being used to image and read a visual code 24 in a scene 26, according to an embodiment of the present invention. Device 20 comprises one or more image capturing devices 28, also termed cameras 28, which are controlled by a local processor 32 of the device. Device 20 is also configured to communicate remotely with a processor 36. In some embodiments local processor 32 and remote processor 36 are configured as a single processor, which may be local to or remote from device 20.

[0068] Device 20 may be formed on different types of platforms, such as a smartphone, a tablet, or a laptop computer; visual code 24 may be presented as different types of code, for example a quick response (QR) code or a color based code. For simplicity, except where otherwise stated in the disclosure, device 20 is assumed to comprise a smartphone, and is also referred to as smartphone 20 or as target smartphone 20 and those having ordinary skill in the art will be able to adapt the disclosure to encompass other types of platform. In addition, while smartphone 20 may have more than one camera 28, for simplicity the disclosure refers to the one or more image capturing devices of smartphone 20 in the singular, as camera 28. Camera 28 has a range of optical zoom settings and a range of digital zoom settings, each of which may be set by local processor 32.

[0069] As described in more detail below, a user 40 calibrates smartphone 20, and after calibration operates the smartphone to capture and read an image of visual code 24. The calibration provides smartphone 20 with a factor which, inter alia, enables processor 32 to operate camera 28 with an optimal zoom value when acquiring an image of scene 26 that includes visual code 24. The calibration also provides smartphone 20 with sharpness settings which enhance the quality of the acquired image.

[0070] In order to efficiently read a visual code, the image of the code should be as large, and as unembellished, as possible. The increased size required may be accomplished in smartphone 20 by zooming the image, and the zooming applied typically comprises a combination of an optical zoom setting and a digital zoom setting. Any optical zoom setting typically improves the readability of the code. However, any digital zoom setting applied may decrease the readability, even though the application of the digital zoom may enhance the appearance of the zoomed image to user 40, for example by adding smoothing pixels to imaged edges. Consequently, embodiments of the invention identify an optimal zoom setting for smartphone 20 so that the image acquired at the setting has a maximal optical zoom component and a minimal digital zoom component.

[0071] While the presence of a digital zoom component typically reduces the readability, of the image, a moderate sharpness setting may increase the readability of the image. Embodiments of the invention consequently also identify an optimal sharpness setting for smartphone 20.

[0072] Fig. 2 illustrates the effect of zooming on the sharpness and readability of an image, according to an embodiment of the present invention. The image on the left has an increased zoom factor, and thus reduced readability, compared to the image on the right.

[0073] Fig. 3 illustrates the effect of different sharpness settings on an image, according to an embodiment of the present invention. The image on the left has a minimal sharpness setting. The image on the right has a high sharpness setting, but this introduces artifacts into edges of the image.

[0074] Embodiments of the invention use trained neural networks for the optimal sharpness setting and for the optimal zoom setting identification. The flowchart of Fig. 4 and the graph of Fig. 5 describe how the optimal zoom setting is identified.

[0075] Zoom Setting

[0076] Fig. 4 is a flowchart 100 of steps describing how an optimal zoom setting is provided to a smartphone, and Fig. 5 is a graph 150 illustrating some of the steps, according to an embodiment of the present invention. In Fig. 4 the first set of steps describe how a neural network is trained, to be used by any smartphone, and the remaining steps describe how the trained network is used to provide the optimal zoom setting to target smartphone 20. In Fig. 5 graph 150 plots zoom settings vs. field of view (FOV) settings.

[0077] Except where otherwise stated, the steps of flowchart 100 are assumed to be performed by processor 36.

[0078] As is described below, graph 150 illustrates FOV being used as a metric for the magnification set by a given smartphone, and this enables embodiments of the invention to standardize magnifications for different smartphones, which in practice typically use different product specific device descriptors, such as 1 - 99 or x0.5 - xlO for the magnifications.

[0079] In an initial step 104 of the flowchart a pair of images is acquired from a device generally similar to smartphone 20. The pair of images comprises a first, minimally zoomed, image and a second maximally zoomed image, i.e., having a zoom setting larger than that of the first image. In some cases the pair of images may be acquired consecutively of the same scene. In alternative cases the second image is generated artificially by cropping the first image, then resizing the cropped image to generate a magnified image.

[0080] In embodiments of the invention, the zoom setting of a device is assumed to be inversely proportional to the FOV of the device, i.e., where z is the zoom setting of the device, and

[0081] FOV is the linear angular field of view of the device.

[0082] In the disclosure a zoom setting may also be referred to as a zoom magnification, or just as a magnification.

[0083] A curve 154 corresponds to equation (1) and as illustrated, a minimally zoomed image having a zoom setting corresponds to an image with a maximal FOV setting maximally zoomed image having a zoom setting corresponds to an image with a minimal FOV setting

[0084] In an annotation step 108, each image is manually annotated with its respective linear FOV, typically measured in degrees. For example, the minimally zoomed image may have an FOV value 80°, and the higher zoomed image may have an FOV value 10°. For the pair the ratio of the zoom settings, corresponding to the inverse of the ratio of the FOV settings, is calculated: where R is the ratio of the zoom settings.

[0085] As is apparent from equation (2), ratio R is device independent. As is described further below, embodiments of the invention use a corpus of data, including images derived from diverse devices 20, to train a neural network. The independence property of R standardizes the data from the diverse devices, enabling the data to be used for the training.

[0086] As is known in the art, the zoom value set by a camera typically comprises a digital zoom component and an optical zoom component. A desired optical zoom component, herein termed corresponds to a desired FOV value, herein termed The desired values correspond to the optimal zoom and FOV settings to be provided to a smartphone.

[0087] A factor r may be defined in terms (from equation (2)).

[0088] Assuming a linear relation between r and a zoom value z an equation relating z and r is:

[0089] It is noted from equation (4) that z is the sum of two components: a fraction p = p + q = 1 (5)

[0090] In a digital zoom step 112 in addition to recording the FOV (in step 108) of the second image, i.e., the more highly magnified image, the second image is annotated with a factor AF, according to equation (6) below. AF is the ratio of the FOV of the second, more highly zoomed image, to the FOV of a desired image, i.e., the image having a maximal optical zoom, and a method of annotation of the image is described further below.

[0091] In an embodiment of the invention steps 104 - 112 are reiterated, as indicated by an arrow 120, to construct a corpus of data for training a neural network operated by processor 36. For the training the annotations of step 112, to determine the digital zoom component, may be provided manually, by a user of the network, and / or by applying a known digital zoom value, for example by cutting and / or resizing a selected image so as to generate the second image.

[0092] In an alternative embodiment one iteration of steps 104 - 112 is used by local processor 32 on target smartphone 20 to find correction factor AF for the target smartphone. In this alternative embodiment, processor 32 may estimate the FOV of the two images by any convenient method known in the art, such as by enumerating pixels, and may find AF of the second image by analyzing a sharpness metric, such as a modulation transfer function, of the image.

[0093] Continuing with flowchart 100, the iterations of steps 104 - 112 are typically acquired on devices formed on different platforms, e.g. smartphones and tablets. The reiterations produce a corpus of a set of pairs of images, together with their associated FOV ratio R and correction factor AF, corresponding to expression (7): where I1, I2, are the first and second images, and n is an index of the set.

[0094] In a training step 124, assumed herein to be implemented by remote processor

[0095] 36, a convolutional neural network (CNN) is trained using the corpus of data corresponding to expression (7). The CNN may comprise a known pre-trained network, such as one of the VGGNet or ResNet models, that has been adapted. The adaptation typically comprises applying transfer learning using an existing image dataset, such as ImageNet. The training may use gradient descent, and in one embodiment a loss function, for the training to achieve minimal cost, may comprise the sum, taken over the corpus of data, of the absolute difference of the predicted value of ratio R with the correct R.

[0096] The trained network is stored with processor 36, and provides a value of correction factor AF given a minimally zoomed image and a maximally zoomed image.

[0097] In a calibration step 128, user 40 operates target smartphone 20 to acquire a minimally zoomed image and a maximally zoomed image, and provides the images to processor 36. Processor 36 calculates a target phone zoom ratio RJARG 'or l'lcPa'r°f images and provides the trained network with the pair of images and the ratio RjARG- In response the network outputs a value, AFJARG- °f the correction factor for the smartphone.

[0098] From equation (3):

[0099] Substituting values from equation (8) in equation (3) gives: Processor 32 of the phone applies the received zoom value when acquiring subsequent images of visual codes such as visual code 24.

[0100] Sharpness Setting

[0101] Target smartphone 20 is assumed to have multiple sharpness settings for image capture, so that a captured image may vary from an insufficiently sharpened image that is blurry, to an excessively sharpened image that typically introduces artifacts into the image. Both types of sharpening reduce the readability of a scanned visual code. However, moderate image sharpening may enhance the readability of a visual code such as code 24, and the flowchart of Fig. 6 describes how the sharpness may be set for a smartphone.

[0102] Fig. 6 is a flowchart 180 of steps for providing a sharpness setting to target smartphone 20, according to an embodiment of the present invention. The steps are assumed to be performed using remote processor 36.

[0103] Prior to performing the flowchart processor 36 is provided with the value of a sharpness metric, herein by way of example assumed to comprise a modulation transfer function at 50% contrast (MTF50) for maximal distance reading of a visual code. The sharpness metric is herein termed

[0104] In an imaging step 188 an image capturing device, such as a smartphone 20, with its sharpness setting maximized, captures an image of a scene. It is noted that a visual code need not be present in the imaged scene. The captured image is analyzed to determine its MTF50 value, and the image is tagged with this value.

[0105] In a correction step 192 a sharpness correction, comprising a difference, AMTF50, between the MTF50 value of the image of step 188 and is determined. The image of step 188 is also tagged with difference AMTF50.

[0106] An arrow 196 indicates that steps 188 and 192 are reiterated to generate a corpus of tagged sharpened images. The reiteration is typically implemented using multiple image capturing devices on multiple platforms.

[0107] In a training step 200 processor 36 trains a CNN, generally similar to the CNN described above with regard to the zoom setting, with the corpus of tagged images. The trained network is stored with processor 36 and provides a value of a sharpness correction given an image generated with maximum sharpness.

[0108] In a calibration step 204, user 40 operates target smartphone 20 to acquire a maximally sharpened image, and provides the image for the CNN to processor 36. In response the network outputs a value, AMTF50, of the sharpness correction for the smartphone.

[0109] In a final step 208, processor 36 transmits sharpness correction factor AMTF50 to the smartphone, and processor 32 of the phone applies the correction factor to set the sharpness of the phone when acquiring images of visual codes such as visual code 24.

[0110] Scene Parameters

[0111] The algorithms installed by the manufacturer of target smartphone 20 enable the phone to automatically adjust the focus, the exposure time, and the white balance of the scene being imaged. However, the automatic adjustment may not be optimized for imaging a visual code such as visual code 24, which may be present in one of many extremely diverse scenes. The automatic adjustment may be optimized for the overall scene, rather than for a region of interest (RO I) comprising the targeted visual code.

[0112] Figs. 7A, 7B, and 7C are examples of images with different exposure times, according to an embodiment of the present invention. Fig. 7A is underexposed for a displayed ROI and Fig. 7B is overexposed for the ROI. Fig. 7C is correctly exposed once the ROI is used to set the algorithms of the phone.

[0113] Embodiments of the invention train a neural network to identify such regions of interest, and deploy the trained network for use on target smartphone 20, as is described below.

[0114] Fig. 8 is a flowchart 250 describing steps for training and using a neural network to identify an ROI for target smartphone 20, according to an embodiment of the present invention. The training described in the flowchart is assumed to be implemented by processor 36.

[0115] In an initial step 254, a diverse set of images, including a visual code, are assembled. The images have different scenarios, such as a television screen, a restaurant menu (in tablet or printed form), an advertising billboard, a patterned surface, or a plain un-patterned surface. The lighting of such scenes may also vary tremendously, from bright sunlight or its equivalent to dim night lighting. The orientation of the visual code may also be varied.

[0116] In an annotation step 258 each image of the set of images is manually annotated with the location of the visual code, herein also termed the region of interest (ROI).

[0117] In a training step 262, processor 36 trains a CNN, generally similar to the CNN described above with regard to the zoom setting, with the corpus of annotated images. The trained network is able to identify the location of an ROI, similar to a visual code, in an image presented to the network.

[0118] In a deployment step 266, the trained network is deployed in target smartphone 20. The deployment may typically be implemented before or during the zoom and sharpness calibrations described above.

[0119] In an imaging step 270 and in a final step 274, when user 40 is using target smartphone 20 to image an ROI where a visual code such as visual code 24 may be present, processor 32 of the target phone accesses the network to identify the ROI. The processor is then able to adjust the focus, the exposure time, and the white balance of the scene being imaged to optimally image the ROI.

[0120] It will be appreciated that the embodiments described above are cited by way of example, and that the present invention is not limited to what has been particularly shown and described hereinabove. Rather, the scope of the present invention includes both combinations and subcombinations of the various features described hereinabove, as well as variations and modifications thereof which would occur to persons skilled in the art upon reading the foregoing description and which are not disclosed in the prior art.

Claims

CLAIMS1. A computer implemented, artificial intelligence (Al), system for reading a visual code, comprising: a camera; an imaging device, having an optical zoom range extending from a minimal optical magnification to a maximal optical magnification; and a processor, configured to perform a plurality of iterations on the imaging device, so that at each iteration the processor: operates the device at a first zoom setting within the optical zoom range so as to capture a first image at a first magnification, and at a second zoom setting so as to capture a second image at a second magnification, greater than the maximal optical magnification, by applying both the maximal optical magnification and a digital magnification to the second image, calculates a ratio of the second magnification to the first magnification, analyzes the second image to estimate the digital magnification of the second zoom setting, in response to the estimated digital magnification, calculates a correction factor for selecting a target zoom setting, and inputs to a neural network the first image, the second image, the ratio, and the correction factor formed in each of the iterations so as to train the neural network to provide the target zoom setting; and wherein the processor is further configured: to operate the camera at the target zoom setting to capture a scene image containing the visual code; to identify the visual code in the scene image; and to read the identified visual code.

2. The system according to claim 1, wherein the processor is configured to receive from the camera, prior to operating the camera at the target zoom setting, a first camera image having a first camera zoom setting and a second camera image having a second camera zoom setting greater than the first camera zoom setting, to provide the first camera image and the second camera image to the neural network, and in response toreceive from the neural network a target correction factor, and apply the target correction factor to the camera to determine the target zoom setting.

3. The system according to claim 1, wherein the second zoom setting comprises a second FOV setting, and wherein the correction factor is a function of the second FOV setting.

4. The system according to claim 1, wherein the imaging device has a digital zoom range, and wherein the correction factor is a function of a digital component of the digital zoom range.

5. The system according to claim 1, wherein analyzing the second image comprises at least one of receiving a manual estimation of the digital magnification of the second image and measuring a modulation transfer function of the second image.

6. A computer implemented system for setting a sharpness metric for an image of a visual code, comprising: a camera; an imaging device, having a sharpness setting range extending from a minimal sharpness setting to a maximal sharpness setting; and a processor, configured to perform a plurality of iterations on the imaging device, so that at each iteration the processor: operates the device at the maximal sharpness setting so as to capture an image at the maximal sharpness setting, formulates a sharpness correction factor, in response to a difference between the maximal sharpness setting and a preset optimal sharpness setting, for selecting a target sharpness setting, associates the sharpness correction factor with the image, inputs to a neural network the image and the correction factor formed in each of the iterations so as to train the neural network to provide the target sharpness setting; and wherein the processor is further configured: to operate the camera at the target sharpness setting to capture a scene image containing the visual code; to identify the visual code in the scene image; andto read the identified visual code.

7. The system according to claim 6, wherein the processor is configured to receive from the camera, prior to operating the camera at the target sharpness setting, a camera image having a maximal camera sharpness setting, to provide the camera image to the neural network, and in response to receive from the neural network a target correction factor, and apply the target correction factor to the camera to determine the target sharpness setting.

8. A computer implemented system for identifying a location of a region of interest (ROI) in an image, comprising: a camera; an imaging device; and a processor, configured to perform a plurality of iterations on the imaging device, so that at each iteration the processor: operates the device so as to capture an image of the ROI in a scene, the scene comprising a scene descriptor classifying the scene, identifies a location of the ROI within the scene, associates the location of the ROI and the scene descriptor with the image, inputs to a neural network the image and the associated ROI location and scene descriptor formed in each of the iterations so as to train the neural network to provide a target ROI location; and wherein the processor is further configured: to operate the camera to capture a camera scene image containing a target ROI; and to identify the target ROI location in the camera scene image.

9. The system according to claim 8, wherein the processor is configured to receive from the camera the camera scene image, to provide the camera scene image to the neural network, and in response to receive from the neural network the target ROI location.

10. A system for reading a visual code, comprising:a camera, having an optical zoom range extending from a minimal optical magnification to a maximal optical magnification; and a processor, configured: to operate the camera at a first zoom setting within the optical zoom range so as to capture a first image at a first magnification, and at a second zoom setting so as to capture a second image at a second magnification, greater than the maximal optical magnification, by applying both the maximal optical magnification and a digital magnification to the second image, to analyze a sharpness metric of the second image to estimate the digital magnification of the second zoom setting; in response to the estimated digital magnification, to select a target zoom setting within the optical zoom range; to operate the camera at the target zoom setting to capture a scene image containing the visual code; to identify the visual code in the scene image; and to read the identified visual code.

11. The system according to claim 10, wherein the first zoom setting comprises a first field of view (FOV) setting and wherein the second zoom setting comprises a second FOV setting.

12. The system according to claim 10, wherein the first FOV setting is an inverse of the first magnification, and wherein the second FOV setting is the inverse of the second magnification.

13. The system according to claim 10, wherein the camera has a digital zoom range, and wherein the target zoom setting is at a boundary between the optical zoom range and the digital zoom range.

14. The system according to claim 10, wherein the sharpness metric comprises a modulation transfer function.

15. A system for determining a sharpness setting for imaging a visual code, comprising: an imaging device configured to image the visual code; anda processor, configured to: receive an optimal sharpness setting for reading the visual code; operate the imaging device at a maximal sharpness setting greater than the optimal sharpness setting; formulate a sharpness correction factor in response to the optimal sharpness setting and the maximal sharpness setting; and apply the sharpness correction factor to the imaging device to provide a target sharpness setting for the imaging device, and operate the imaging device at the target sharpness setting when imaging the visual code.

16. A computer implemented, artificial intelligence (Al), method for reading a visual code, comprising: providing an imaging device with an optical zoom range extending from a minimal optical magnification to a maximal optical magnification; configuring a processor to perform a plurality of iterations on the imaging device, so that at each iteration the processor: operates the device at a first zoom setting within the optical zoom range so as to capture a first image at a first magnification, and at a second zoom setting so as to capture a second image at a second magnification, greater than the maximal optical magnification, by applying both the maximal optical magnification and a digital magnification to the second image, calculates a ratio of the second magnification to the first magnification, analyzes the second image to estimate the digital magnification of the second zoom setting, in response to the estimated digital magnification, calculates a correction factor for selecting a target zoom setting, and inputs to a neural network the first image, the second image, the ratio, and the correction factor formed in each of the iterations so as to train the neural network to provide the target zoom setting; and further configuring the processor: to operate a camera at the target zoom setting to capture a scene image containing the visual code;to identify the visual code in the scene image; and to read the identified visual code.

17. The method according to claim 16, wherein the processor is configured to receive from the camera, prior to operating the camera at the target zoom setting, a first camera image having a first camera zoom setting and a second camera image having a second camera zoom setting greater than the first camera zoom setting, to provide the first camera image and the second camera image to the neural network, and in response to receive from the neural network a target correction factor, and apply the target correction factor to the camera to determine the target zoom setting.

18. The method according to claim 16, wherein the second zoom setting comprises a second FOV setting, and wherein the correction factor is a function of the second FOV setting.

19. The method according to claim 16, wherein the imaging device has a digital zoom range, and wherein the correction factor is a function of a digital component of the digital zoom range.

20. The method according to claim 16, wherein analyzing the second image comprises at least one of receiving a manual estimation of the digital magnification of the second image and measuring a modulation transfer function of the second image.

21. A computer implemented method for setting a sharpness metric for an image of a visual code, comprising: providing an imaging device, having a sharpness setting range extending from a minimal sharpness setting to a maximal sharpness setting; and configuring a processor to perform a plurality of iterations on the imaging device, so that at each iteration the processor: operates the device at the maximal sharpness setting so as to capture an image at the maximal sharpness setting, formulates a sharpness correction factor, in response to a difference between the maximal sharpness setting and a preset optimal sharpness setting, for selecting a target sharpness setting, associates the sharpness correction factor with the image,inputs to a neural network the image and the correction factor formed in each of the iterations so as to train the neural network to provide the target sharpness setting; and further configuring the processor: to operate a camera at the target sharpness setting to capture a scene image containing the visual code; to identify the visual code in the scene image; and to read the identified visual code.

22. The method according to claim 21, wherein the processor is configured to receive from the camera, prior to operating the camera at the target sharpness setting, a camera image having a maximal camera sharpness setting, to provide the camera image to the neural network, and in response to receive from the neural network a target correction factor, and apply the target correction factor to the camera to determine the target sharpness setting.

23. A computer implemented method for identifying a location of a region of interest (RO I) in an image, comprising: providing an imaging device; and configuring a processor to perform a plurality of iterations on the imaging device, so that at each iteration the processor: operates the device so as to capture an image of the ROI in a scene, the scene comprising a scene descriptor classifying the scene, identifies a location of the ROI within the scene, associates the location of the ROI and the scene descriptor with the image, inputs to a neural network the image and the associated ROI location and scene descriptor formed in each of the iterations so as to train the neural network to provide a target ROI location; and further configuring the processor: to operate a camera to capture a camera scene image containing a target ROI; and to identify the target ROI location in the camera scene image.

24. The method according to claim 23, wherein the processor is configured to receive from the camera the camera scene image, to provide the camera scene image to the neural network, and in response to receive from the neural network the target ROI location.

25. A method for reading a visual code, comprising: providing a camera with an optical zoom range extending from a minimal optical magnification to a maximal optical magnification; and operating the camera at a first zoom setting within the optical zoom range so as to capture a first image at a first magnification, and at a second zoom setting so as to capture a second image at a second magnification, greater than the maximal optical magnification, by applying both the maximal optical magnification and a digital magnification to the second image, analyzing a sharpness metric of the second image to estimate the digital magnification of the second zoom setting; in response to the estimated digital magnification, selecting a target zoom setting within the optical zoom range; operating the camera at the target zoom setting to capture a scene image containing the visual code; identifying the visual code in the scene image; and reading the identified visual code.

26. The method according to claim 25, wherein the first zoom setting comprises a first field of view (FOV) setting and wherein the second zoom setting comprises a second FOV setting.

27. The method according to claim 25, wherein the first FOV setting is an inverse of the first magnification, and wherein the second FOV setting is the inverse of the second magnification.

28. The method according to claim 25, wherein the camera has a digital zoom range, and wherein the target zoom setting is at a boundary between the optical zoom range and the digital zoom range.

29. The method according to claim 25, wherein the sharpness metric comprises a modulation transfer function.

30. A method for determining a sharpness setting for imaging a visual code, comprising: configuring an imaging device to image the visual code; and configuring a processor to: receive an optimal sharpness setting for reading the visual code, operate the imaging device at a maximal sharpness setting greater than the optimal sharpness setting, formulate a sharpness correction factor in response to the optimal sharpness setting and the maximal sharpness setting, and apply the sharpness correction factor to the imaging device to provide a target sharpness setting for the imaging device, and operate the imaging device at the target sharpness setting when imaging the visual code.

Citation Information

Patent Citations

  • Barcode Recognition Using Data-Driven Classifier

    US20130240628A1

  • Context-based priors for object detection in images

    US20170011281A1

  • Digital image capturing method and electronic device for digital image capturing

    US20180157885A1

  • Artificial intelligence-based machine readable symbol reader

    US20190303636A1

  • Method for Identifying Objects in an Image of a Camera

    US20190354783A1