A method for performing rapid and accurate autofocusing using deep learning techniques

IN595240BActive Publication Date: 2026-07-13SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
IN · IN
Patent Type
Patents
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2022-03-07
Publication Date
2026-07-13

AI Technical Summary

Technical Problem

Existing autofocusing methods in smartphones are either slow and inaccurate in bright conditions or costly and require additional hardware in low light conditions, failing to provide rapid and accurate focusing without hardware integration.

Method used

A method using deep learning techniques, specifically Generative Adversarial Networks (GANs), to detect focus regions, correct pixel intensity, and generate missing pixel information for rapid and accurate autofocusing, eliminating the need for additional hardware.

Benefits of technology

Enables rapid and accurate autofocusing in smartphones without additional hardware, allowing users to capture high-quality images quickly and cost-effectively, even in challenging lighting conditions.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A method for performing rapid and accurate autofocusing using deep learning techniques [0033] The present invention discloses a method for performing rapid and accurate autofocusing using deep learning techniques. One or more focus regions in an image frame are detected and divided into plurality of segments, wherein segmentation is performed based on the Red Green Blue (RGB) colour intensity of the focus regions. The segments are prioritized based on at least one RGB colour intensity, wherein the pixel intensity of the prioritized segments are subsequently corrected using a pixel intensity correcting GAN. The missing pixels information for the prioritized and intensity corrected segments are generated through a pixel information generating GAN using Artificial Intelligence (AI) and deep learning techniques. Further, an autofocus driver in the image capturing device adjusts the position of the lens based on the phase differences of the pixels information obtained from the pixel information generating GAN. (FIG 1)
Need to check novelty before this filing date? Find Prior Art

Description

Technical field of the invention

[0002] The present invention discloses a method for performing rapid andaccurate autofocusing using deep learning techniques. The invention particularlyrelates to a method for correcting the colour intensity of an image frame using apixel intensity correcting Generative Adversarial Networks (GAN) andsubsequently generating the missing pixel information for the intensity correctedimage frame using a pixel information generating GAN.Background of the invention

[0003] The image capturing feature in smartphones has evolved by many foldssince the past decade, wherein the quality of images captured from the cameraprovided in the smartphone has improved drastically. However, with theimprovement in the technology to capture highly focused, near reality images, aconsequential time delay is observed due to the amount of real-time imageprocessing which takes place at the background. The smartphone users expect tocapture highly focused images rapidly and at the precise moment. However, thechallenge faced by smartphone manufacturers is to achieve the required autofocusin a very short duration of time.

[0004] Various attempts have been made for improving the ability ofautofocusing within a short duration of time. One such method includes passiveautofocusing using Contrast Detect Autofocus (CDAF), wherein focus iscalculated while moving the lens and searching for the focus position with thehighest contrast. Another mechanism in passive autofocusing is Phase DetectionAuto-Focus (PDAF), wherein it determines whether the light incident through thelens is focused by dividing it into pairs. Further, Depth From Defocus (DFD) isInternal Ref: OR22C004another passive-autofocusing technique which calculates on focus position bymeasuring defocus change in image. Though the afore-mentioned passiveautofocusing techniques are cost effective and do not require additional hardware,they do not work effectively in low light regions, and they are usually slow,inaccurate and involve complex computation due to which they are not preferredfor the purpose of improving the ability of autofocusing within a short duration oftime.

[0005] Yet another method of auto-focusing is known as active-autofocusingwhich involves a plurality of techniques such as measurement of laser reflectiontime and calculation of distance using laser focus. In another technique involvingthe measurement using an infrared system, the distance is calculated by measuring(triangulation) the angle at which the strength of the reflected signal ismaximized. Furthermore, in the active auto-focusing technique using supersonicwaves, the time until the ultrasonic wave is reflected is measured and the distanceusing the measured time is calculated. Another mechanism of active-autofocusingis the Time Of Flight (TOF) autofocusing technology which reads the amount oflight received to a complement receptor in which the ON-OFF light source phaseis synchronized to determine the distance. Though the afore-mentioned passiveautofocusing techniques are accurate, fast and work well in low light regions, theyare often expensive and require an additional hardware integration.

[0006] For instance, US Patent Application No. US20210066371A1 titled "Animage sensor comprising an array of pixels for reducing the shadowing effect, andcorresponding designing method and computer program product" discloses animage sensor comprising an array of pixels, wherein a set of pixels of the arraycomprises pixels with different height levels arranged according to their heightlevel and relative position to an optical axis of the image sensor, wherein eachpixel of the set can take one height level i among N different height levels, N≥2,where i=1 is the smallest height level and i=N is the highest height level.According to the disclosure, for a pixel of the set having a first height level equalto n, 2≥n≥N, and an adjacent pixel of the set having a second height level equal tom, lower than the first height level, in at least one of horizontal, vertical, orInternal Ref: OR22C004diagonal scanning direction, from a point at which the optical axis intersects theimage sensor to at least one rim of the image sensor, the second height level m isequal to n-1. However, "US20210066371A1" only discloses the method forminimizing shadow effect and does not disclose details pertaining to mechanismsto improve autofocusing of image frames.

[0007] For instance, US Patent Application No. US20210120198A1 titled "Imagesensors including phase detection pixel" discloses an image sensor which includesa pixel array including a plurality of image sensing pixels in a substrate, a phasedetection shared pixel in the substrate, the phase detection shared pixel includingtwo phase detection subpixels arranged next to each other, a color filter fencedisposed on the plurality of image sensing pixels, and the phase detection sharedpixel, the color filter fence defining a plurality of color filter spaces, a plurality ofcolor filter layers respectively disposed in the plurality of color filter spaces on theplurality of image sensing pixels, and the phase detection shared pixel, a firstmicro-lens disposed on each of the plurality of image sensing pixels to have a firstheight, and a second micro-lens disposed to vertically overlap the two phasedetection subpixels of the phase detection shared pixel and to have a secondheight greater than the first height. However, "US20210120198A1" only disclosesthe method for phase detection using two phase detecting pixels and does notdisclose details pertaining to the mechanisms to improve autofocusing of imageframes.

[0008] Hence, there exists a need for a cost-effective, rapid and accurateautofocusing mechanism without the requirement of an additional hardwareintegration.Summary of the invention:

[0009] The present invention overcomes the drawbacks of the prior art bydisclosing a method for performing rapid and accurate autofocusing using deeplearning techniques. The method comprises the steps of detecting one or morefocus regions in an image frame while a user captures an image using an imagecapturing device and dividing the detected focus regions into a plurality ofInternal Ref: OR22C004segments, wherein the segmentation is performed based on the Red Green Blue(RGB) colour intensity of the focus regions. Further, the segments are prioritizedbased on at least one RGB colour intensity of the focus regions, wherein the pixelintensity of the prioritized segments are subsequently corrected using a pixelintensity correcting Generative Adversarial Networks (GAN) which comprises agenerator-discriminator network.

[0010] Further, the missing pixels information for the prioritized and intensitycorrected segments of the image are generated using a pixel informationgenerating GAN, wherein the pixel information generating GAN generatesmissing pixels information using Artificial Intelligence (AI) and deep learningtechniques. Further, the missing pixels information generated by the pixelinformation generating GAN is transmitted to an autofocus driver provided in theimage capturing device, wherein the autofocus driver adjusts the position of thelens based on the phase differences of the pixels information obtained from thepixel information generating GAN.

[0011] Thus, the present invention provides a mechanism to improve the ability toauto-focus in an image capturing device such as a smartphone without therequirement of any additional hardware integration. Additionally, the user of theimage capturing device is allowed to capture highly focused images within a shortduration of time thereby ensuring that the user is able to capture special andimportant moments with high precision. Further, the implementation of themethod disclosed in the present invention is cost-effective in comparison to themechanisms disclosed in the prior art.Brief description of the drawings:

[0012] The foregoing and other features of embodiments will become moreapparent from the following detailed description of embodiments when read inconjunction with the accompanying drawings. In the drawings, like referencenumerals refer to like elements.Internal Ref: OR22C004

[0013] FIG 1 illustrates a method for performing rapid and accurate autofocusingusing deep learning techniques.

[0014] FIG 2 illustrates a method of detecting one or more focus regions in animage frame.

[0015] FIG 3 illustrates a method for dividing the detected focus regions into aplurality of segments based on RGB colour intensity.

[0016] FIG 4 illustrates a block diagram of a pixel intensity correctingGenerative Adversarial Networks (GAN) which comprises a generatordiscriminator network.

[0017] FIG 5 illustrates a method of generating the missing pixel information forthe prioritized and intensity corrected segments using the pixel informationgenerating GAN.Detailed description of the invention:

[0018] Reference will now be made in detail to the description of the presentsubject matter, one or more examples of which are shown in figures. Eachexample is provided to explain the subject matter and not a limitation. Variouschanges and modifications obvious to one skilled in the art to which the inventionpertains are deemed to be within the spirit, scope and contemplation of theinvention.

[0019] FIG 1 illustrates a method for performing rapid and accurate autofocusingusing deep learning techniques, wherein the method (100) comprises the steps ofdetecting one or more focus regions in an image frame while a user captures animage using an image capturing device in step (101), wherein in one embodiment,the image capturing device may be a smartphone device. In step (102), thedetected focus regions are divided into a plurality of segments, wherein thesegmentation is performed based on the Red Green Blue (RGB) colour intensityof the focus regions. Subsequently, in step (103), the segments are prioritizedbased on at least one RGB colour intensity of the focus regions and the pixelInternal Ref: OR22C004intensity of the prioritized segments are corrected using a pixel intensitycorrecting Generative Adversarial Networks (GAN), wherein the pixel intensitycorrecting GAN comprises a generator-discriminator network.

[0020] In one embodiment, the generator network comprises an encoder, bottleneck layer and a decoder, wherein the encoder encodes the segments of the inputimage frame by performing 3x3 convolutions at three levels and doubles thenumber of feature channels in each level. Further, the bottle neck layer performs1x1 convolutions and doubles the number of feature channels to 1024.Furthermore, the decoder decodes the encoded image frame by performingtransposed convolutions at three levels with 3x3 convolutions at each level,thereby reducing the number of feature channels as a result of which the imageframe of the same dimension as the input image frame is generated at the outputof the generator network. Hence, the generator network in the pixel intensitycorrecting GAN provides an output which is identical to the input image framethereby deceiving the discriminator network to interpret the output of thegenerator network as the original input image.

[0021] In one embodiment, the discriminator network in the pixel intensitycorrecting GAN accepts the output of the generator network as an input andperforms 3x3 convolutions with a gradual increase in the number of feature maps,wherein the discriminator network provides an output image frame with thecorrected pixel intensity such that the cross-entropy loss between the input imageframe and output image frame of the discriminator network is minimum.

[0022] In step (104), the missing pixels information for the prioritized andintensity corrected segments of the image are generated using a pixel informationgenerating GAN, wherein the pixel information generating GAN generatesmissing pixels information using Artificial Intelligence (AI) and deep learningtechniques. Further, the missing pixels information generated by the pixelinformation generating GAN is provided to an autofocus driver provided in theimage capturing device, wherein the autofocus driver adjusts the position of theInternal Ref: OR22C004lens based on the phase differences of the pixels information obtained from thepixel information generating GAN in step (105).

[0023] FIG 2 illustrates a method of detecting one or more focus regions in animage frame, wherein the method (101) comprises the steps of creating scalespace and performing log approximation for a given input image frame in step(101a). Further, in step (101b), one or more key points in the given image frameare detected using a feature enhancement technique known as Difference ofGaussians (DoG), wherein DoG is estimated by the blurring of the given imageframe at two different scales as a result of which noise present in the image frameis eliminated and key points are highlighted. In one embodiment, the given imageframe is checked for local extremas by comparing each point of the image framewith 8 neighboring pixels of the current image frame and with 9 neighboringpixels in the scales above and below the current image frame. Each point isselected only if it is larger or smaller than all the pixels with which is it compared,wherein such selected points are the required key points which determine thepoints whose intensity changes sharply irrespective of the scale of the image.After obtaining the required key points in the given image frame, all the poorlylocalized points along the edges (i.e., noise) are removed and the important cornerpoints and edge points of the image frame are retained. After obtaining therequired key points in the given image frame, low contrast key points and poorlylocalized points along the edges (i.e., noise) of the image frame are eliminated andthe high contrast corner points and edge points of the image frame are retained instep (101c).

[0024] FIG 3 illustrates a method for dividing the detected focus regions into aplurality of segments based on RGB colour intensity, wherein the method (102)comprises the steps of choosing one or more clusters based on the identified focusregions in step (102a) and checking if the given image frame is in Green (Y), Blue(Cb), Red (Cr) (YCbCr) colour format in step (102b). If the given input frame isnot as per the YCbCr colour format, the image frame is converted from the Red,Green, Blue (RGB) colour format into the YCbCr colour format in step (102c).Conversely, if the given input frame is as per the YCbCr colour format, theInternal Ref: OR22C004colours in the CbCr space of the image frame are classified using the technique ofK Nearest Neighbor (KNN) clustering in step (102d). KNN is one of the simplestMachine Learning algorithms based on supervised learning technique whichstores all the available data and classifies a new data point based on the similarity.KNN clustering is a well-known technique that has been extensively covered bythe prior art and hence not described in detail.

[0025] In step (102e), new centroids are updated post classification of the imageframe using KNN clustering and new centroids are compared to the old centroidsin step (102f). If the new centroids are not equal to the old centroids, the coloursin the CbCr space of the image frame are re-classified using KNN clustering instep (102g). However, if the new centroids are equal to the old centroids, theidentified focus regions are segmented based on colour intensity in step (102h).

[0026] FIG 4 illustrates a block diagram of a pixel intensity correcting GenerativeAdversarial Networks (GAN) which comprises a generator-discriminator network.The generator network in the pixel intensity correcting GAN provides an outputwhich is identical to the input image frame thereby deceiving the discriminatornetwork to interpret the output of the generator network as the original inputimage. The discriminator network in the pixel intensity correcting GAN acceptsthe output of the generator network and provides an output image frame with thecorrected pixel intensity such that the cross-entropy loss between the input imageframe and output image frame of the discriminator network is minimum.

[0027] FIG 5 illustrates a method of generating the missing pixel information forthe prioritized and intensity corrected segments using the pixel informationgenerating GAN, wherein the method (104) comprises the steps of convertinghigh-resolution images into low-resolution images through the process ofdownsampling, wherein the high-resolution images are output images and thedownsampled low resolution images are input images in step (104a). The inputimages and output images are converted to vector format and provided as inputsto a deep neural network for the purpose of generating one or more trainingmodels in step (104b). In step (104c), the features from both, the input images andInternal Ref: OR22C004the output images are extracted, wherein the colours of the input images aremapped with the colours of the output images thereby resulting in a pixelinformation generating GAN output. In step (104d), a real-time image frame iscompared with one or more trained models by the pixel information generatingGAN and the missing pixel information is generated in the real-time image framefor the purpose of improving the focusing ability of the real-time image frame.The resultant loss, i.e., pixel loss, content loss, adversarial loss and the PeakSignal to Noise Ratio (PSNR) observed at the output of the pixel informationgenerating GAN is minimum.

[0028] At least one of the plurality of modules may be implemented through anAI model. A function associated with AI may be performed through the nonvolatile memory, the volatile memory, and the processor. The processor mayinclude one or a plurality of processors. At this time, one or a plurality ofprocessors may be a general-purpose processor, such as a central processing unit(CPU), an application processor (AP), or the like, a graphics-only processing unitsuch as a graphics processing unit (GPU), a visual processing unit (VPU), and / oran AI-dedicated processor such as a neural processing unit (NPU).

[0029] The one or a plurality of processors control the processing of the inputdata in accordance with a predefined operating rule or artificial intelligence (AI)model stored in the non-volatile memory and the volatile memory. The predefinedoperating rule or artificial intelligence model is provided through training orlearning. Here, being provided through learning means that, by applying alearning algorithm to a plurality of learning data, a predefined operating rule or AImodel of a desired characteristic is made. The learning may be performed in adevice itself in which AI according to an embodiment is performed, and / o may beimplemented through a separate server / system.

[0030] The AI model may consist of a plurality of neural network layers. Eachlayer has a plurality of weight values and performs a layer operation throughcalculation of a previous layer and an operation of a plurality of weights.Examples of neural networks include, but are not limited to, Convolutional NeuralInternal Ref: OR22C004Network (CNN), Deep Neural Network (DNN), Recurrent Neural Network(RNN), Restricted Boltzmann Machine (RBM), Deep Belief Network (DBN),Bidirectional Recurrent Deep Neural Network (BRDNN), Generative AdversarialNetworks (GAN), and deep Q-networks. The learning algorithm is a method fortraining a predetermined target device (for example, a robot) using a plurality oflearning data to cause, allow, or control the target device to make a determinationor prediction. Examples of learning algorithms include, but are not limited to,supervised learning, unsupervised learning, semi-supervised learning, orreinforcement learning.

[0031] The present invention provides a mechanism to improve the ability toauto-focus in an image capturing device such as a smartphone without therequirement of any additional hardware integration. Additionally, the user of theimage capturing device is allowed to capture highly focused images within a shortduration of time thereby ensuring that the user is able to capture special andimportant moments with high precision. Further, the implementation of themethod (100) disclosed in the present invention is cost-effective in comparison tothe mechanisms disclosed in the prior art.

[0032] While at least one exemplary embodiment has been presented in theforegoing detailed description, it should be appreciated that a vast number ofvariations exist.

Claims

1. A method for performing rapid and accurate autofocusing using deep learning techniques, the method (100) comprising the steps of: a. detecting one or more focus regions in an image frame while a user captures an image using an image capturing device; b. dividing the detected focus regions into a plurality of segments, wherein the segmentation is performed based on the Red Green Blue (RGB) colour intensity of the focus regions; c. prioritizing the segments based on at least one RGB colour intensity of the focus regions and correcting the pixel intensity of the prioritized segments using a pixel intensity correcting Generative Adversarial Networks (GAN), wherein the pixel intensity correcting GAN comprises a generator-discriminator network; d. generating the missing pixels information for the prioritized and intensity corrected segments of the image using a pixel information generating GAN, wherein the pixel information generating GAN generates missing pixels information using deep learning techniques; e. providing the missing pixels information generated by the pixel information generating GAN to an autofocus driver provided in the image capturing device, wherein the autofocus driver adjusts the position of the lens based on the phase differences of the pixels information obtained from the pixel information generating GAN.

2. The method (100) as claimed in claim 1, wherein the method of detecting one or more focus regions in an image frame comprises the steps of: a. creating scale space and performing log approximation for a given input image frame; b. detecting one or more key points in the given image frame using a feature enhancement technique known as Difference of Gaussians (DoG), wherein DoG is estimated by the blurring of the given image frame at two different scales as a result of which noise present in the image frame is eliminated and key points are highlighted; c. eliminating low contrast key points and poorly localized key points along the edges of the image frame and retaining the high contrast corner points and edge points of the image frame.

3. The method (100) as claimed in claim 1, wherein the method for dividing the detected focus regions into a plurality of segments based on RGB colour intensity comprises the steps of: a. choosing one or more clusters based on the identified focus regions; b. checking if the given image frame is in Green (Y), Blue (Cb), Red (Cr) (YCbCr) colour format, wherein if the given input frame is: i. not as per the YCbCr colour format, the image frame is converted from the Red, Green, Blue (RGB) colour format into the YCbCr colour format; ii. as per the YCbCr colour format, the colours in the CbCr space of the image frame are classified using the technique of K Nearest Neighbor (KNN) clustering; c. updating new centroids post classification of the image frame using KNN clustering; d. checking if the new centroids are equal to the old centroids, wherein if the new centroids are: i. not equal to the old centroids, the colours in the CbCr space of the image frame are re-classified using KNN clustering; ii. equal to the old centroids, the identified focus regions are segmented based on colour intensity.

4. The method (100) as claimed in claim 1, wherein the pixel intensity correcting GAN for correcting the intensity of the prioritized segments comprises a generator-discriminator network whereby the generator network comprises an encoder, bottle neck layer and a decoder.

5. The method (100) as claimed in claim 1, wherein the generator network in the pixel intensity correcting GAN used for pixel intensity correction comprises: a. an encoder for encoding the segments of the input image frame by performing 3x3 convolutions at three levels and doubles the number of feature channels in each level; b. a bottle neck layer for performing 1x1 convolutions and doubling the number of feature channels to 1024; c. a decoder for decoding the encoded image frame by performing transposed convolutions at three levels with 3x3 convolutions at each level, thereby reducing the number of feature channels as a result of which the image frame of the same dimension as the input image frame is generated at the output of the generator network.

6. The method (100) as claimed in claim 1, wherein the discriminator network in the pixel intensity correcting GAN accepts the output of the generator network as an input and performs 3x3 convolutions with a gradual increase in the number of feature maps.

7. The method (100) as claimed in claim 1, wherein the generator network in the pixel intensity correcting GAN provides an output which is identical to the input image frame thereby deceiving the discriminator network to interpret the output of the generator network as the original input image.

8. The method (100) as claimed in claim 1, wherein the discriminator network in the pixel intensity correcting GAN provides an output image frame with the corrected pixel intensity such that the cross-entropy loss between the input image frame and output image frame of the discriminator network is minimum.

9. The method (100) as claimed in claim 1, wherein the method of generating the missing pixel information for the prioritized and intensity corrected segments of the image using the pixel information generating GAN comprises the steps of: a. converting high-resolution images into low-resolution images through the process of downsampling, wherein the high-resolution images are output images and the downsampled low resolution images are input images; b. providing the input images and output images in vector format as inputs to a deep neural network for the purpose of generating one or more training models; c. extracting the features from both, the input images and the output images, wherein the colours of the input images are mapped with the colours of the output images thereby resulting in a pixel information generating GAN output; d. comparing a real-time image frame with one or more trained models by the pixel information generating GAN and generating the missing pixel information in the real-time image frame for the purpose of improving the focusing ability of the real-time image frame.