Methods and electronic devices for managing image artifacts
By using Generative Adversarial Networks (GANs) to classify and optimize image artifacts, the problems of excessive user intervention and incomplete artifact removal in existing technologies are solved, achieving intelligent image artifact management and quality improvement.
Patent Information
- Application Number
- CN202180075818.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-12
- Filing Date
- 2021-11-23
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2041-11-23
AI Technical Summary
Existing technologies require multiple user interventions when managing image artifacts, and cannot effectively remove artifacts such as glare by changing image contrast.
Multiple generative adversarial networks (GANs) are used to classify artifacts in the image, generate binary masks to preserve or remove artifact parts, and optimize image quality by adjusting the gamma value.
It enables intelligent management of image artifacts with minimal user intervention, improving image quality and visibility, and is applicable to various artifact types.
Smart Images

Figure CN116569207B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to electronic devices, and more specifically to methods and electronic devices for managing image artifacts. Background Technology
[0002] Typically, cameras in electronic devices offer various options for editing images. However, during image capture, various artifacts (such as shadows, flare, exposure inconsistencies, and color discrepancies) can be added to the image. These artifacts degrade the user experience and image quality. For example, as... Figure 1 As shown, artifacts (such as shadows) can obscure the texture details captured in photo 102, create unwanted patterns on the face of the human subject captured in photo 104, and may make it difficult to recognize the characters in optical character recognition (OCR) when artifacts appear on the characters shown in photo 106. If the artifact is glare, it can add unwanted patterns to photo 108 and may hinder the visibility of the object captured in photo 108. Furthermore, there is no single method to effectively manage various artifacts.
[0003] Existing solutions for managing artifacts require multiple user intervention instances. For example, users might need to use methods such as... Figure 2 The shadow slider shown allows you to adjust the image contrast to select the degree of artifact removal needed; however, the shadow slider's function is not effective at eliminating shadows. Artifacts (such as glare) in an image cannot be simply removed by changing its contrast. Summary of the Invention
[0004] Technical issues
[0005] Existing solutions for managing artifacts require multiple user intervention instances. Artifacts (such as glare) in an image cannot be simply removed by changing its contrast.
[0006] Technical solution
[0007] The primary objective of these embodiments is to provide a method and electronic device for managing image artifacts. Aside from the user's selection of an artifact management icon on the electronic device, the method according to these embodiments requires minimal or no manual intervention from the user. Furthermore, the method is not limited to any particular artifact type and can be used to manage all forms of artifacts (e.g., shadows, glare, etc.).
[0008] Another objective of this embodiment is to use multiple Generative Adversarial Networks (GANs) and classify artifacts or portions of artifacts into desired and unwanted artifacts. Therefore, portions classified as desired artifacts are retained in the output image, and portions classified as unwanted artifacts are removed from the output image. Thus, the method and electronic device according to the embodiment may not completely remove artifacts, but rather intelligently determines whether a portion of an artifact enhances image details and may retain that portion.
[0009] Beneficial effects
[0010] The embodiments described herein provide a method and electronic device for managing image artifacts. Attached Figure Description
[0011] The above and / or other aspects will become more apparent from the description of certain exemplary embodiments with reference to the accompanying drawings, in which:
[0012] Figure 1 Example images including artifacts according to existing technology are shown;
[0013] Figure 2 This illustrates an example scenario of removing shadows from an image using existing techniques;
[0014] Figure 3 This is a block diagram of an electronic device for managing image artifacts according to embodiments disclosed herein;
[0015] Figure 4a This is a flowchart illustrating a method for managing image artifacts by an electronic device 100 according to embodiments disclosed herein;
[0016] Figure 4b This is a flowchart illustrating a method for managing image artifacts by an electronic device according to embodiments disclosed herein;
[0017] Figure 5a This is an example illustrating the extraction of multiple features from an input image by an electronic device according to embodiments disclosed herein;
[0018] Figure 5b This is an example illustrating the use of Gaussian blur to extract the colors of an input image by an electronic device according to embodiments disclosed herein;
[0019] Figure 5c This is an example illustrating the extraction of texture from an input image using Otsu thresholding by an electronic device according to embodiments disclosed herein;
[0020] Figure 6aThe following is illustrated: an architecture comprising multiple GANs for obtaining an intermediate output image without artifacts by an electronic device, according to embodiments disclosed herein.
[0021] Figure 6b The architectures of a first GAN and a third GAN according to embodiments of an image artifact controller disclosed herein are shown;
[0022] Figure 6c The architecture of attention blocks and removal blocks for each of the first and third GANs according to embodiments disclosed herein is shown;
[0023] Figure 6d The architecture of a long short-term memory (LSTM) for an attention block of a first GAN according to an embodiment disclosed herein is shown.
[0024] Figure 6e A schematic diagram of a second GAN and a thinning generator according to an embodiment of the image artifact controller disclosed herein is shown;
[0025] Figure 6f The detailed internal architecture of the second GAN and the thinning generator according to the embodiments of the image artifact controller disclosed herein is shown;
[0026] Figure 6g The transformations associated with multiple layers of the second GAN and the refinement generator according to embodiments disclosed herein are illustrated.
[0027] Figure 6h The architecture of the negative discriminator, the segmentation discriminator, and the refinement discriminator of an image artifact controller according to embodiments disclosed herein is shown;
[0028] Figure 7a This is a flowchart illustrating a method for determining whether an artifact is desired or undesirable by a mask classifier according to embodiments disclosed herein;
[0029] Figure 7b This is an example illustrating the processing of a mask classifier according to embodiments disclosed herein;
[0030] Figure 7c Examples illustrating desired and unwanted artifacts in images according to embodiments disclosed herein;
[0031] Figure 8a The architecture of an image quality controller is shown, which determines the image with the best quality after a binary mask is superimposed on an input image according to embodiments disclosed herein.
[0032] Figure 8b This is a flowchart illustrating a method for determining an image of optimal quality after superimposing a binary mask onto an input image, according to embodiments disclosed herein.
[0033] Figure 8c This is a table illustrating NIQE relative to gamma for obtaining various images by changing gamma values according to embodiments disclosed herein;
[0034] Figure 8d This is a graph showing NIQE versus gamma for obtaining various images by changing gamma values according to embodiments disclosed herein.
[0035] Figure 8e This is a flowchart illustrating a method for determining the ideal IQE for obtaining optimal quality from the NIQE relative to the gamma plot, according to embodiments disclosed herein;
[0036] Figure 8f This illustrates a scenario where gamma-corrected color compensation is used in an input image including artifacts, according to embodiments disclosed herein.
[0037] Figure 9a This is an example illustrating a scene depicting optical characters according to existing technology;
[0038] Figure 9b This is an example illustrating a scenario of OCR according to embodiments disclosed herein;
[0039] Figure 10a This is an example illustrating a scene with shadows along a black background, according to existing technology;
[0040] Figure 10b This is an example illustrating a scene with shadows along a black background according to embodiments disclosed herein;
[0041] Figure 10c This is an example illustrating a scene of intelligent overlay of artifacts by an electronic device according to embodiments disclosed herein;
[0042] Figure 10d This is an example illustrating a scenario of artifact removal in a live preview mode of an electronic device according to embodiments disclosed herein;
[0043] Figure 10e This is an example illustrating an artifact management active element in a live preview mode of an electronic device according to embodiments disclosed herein;
[0044] Figure 10f This is an example illustrating an example mode of managing artifacts by an electronic device according to embodiments disclosed herein;
[0045] Figure 11a This is a flowchart illustrating a method for managing image artifacts by an electronic device according to embodiments disclosed herein; and
[0046] Figure 11b This is an example illustrating the management of image artifacts by an electronic device according to embodiments disclosed herein. Detailed Implementation
[0047] Best mode
[0048] According to one aspect of this disclosure, a method for processing image data may include: receiving an input image; extracting a plurality of features from the input image, wherein the plurality of features may include texture of the input image, color composition of the input image, and edges in the input image; determining at least one region of interest (RoI) in the input image including at least one artifact based on the plurality of features; generating at least one intermediate output image by removing the at least one artifact from the input image using a plurality of generative adversarial networks (GANs); generating a binary mask using the at least one intermediate output image, the input image, edges in the input image, and edges in the at least one intermediate output image; classifying the at least one artifact into a first artifact category or a second artifact category based on the binary mask of the input image; and obtaining a final output image by processing the input image based on the category of the at least one artifact corresponding to the first artifact category or the second artifact category.
[0049] The first artifact category corresponds to the desired artifact, and the second artifact category corresponds to the unwanted artifact.
[0050] The method may further include: generating multiple versions of at least one intermediate output image by changing the gamma value associated with the final output image; determining a Natural Image Quality Evaluator (NIQE) value for each of the multiple versions of the at least one intermediate output image; and displaying one of the multiple versions of the at least one intermediate output image as the final output image, the one version having the minimum NIQE value among the multiple versions of the at least one intermediate output image.
[0051] The extraction of the multiple features from the input image may include: extracting texture, color composition, and edges from the input image based on Otsu thresholding, Gaussian blur, and edge detection, respectively.
[0052] The multiple GANs may include a negative generator, an artifact generator, a partition generator, a negative discriminator, a partition discriminator, a refinement generator, and a refinement discriminator.
[0053] Generating the at least one intermediate output image may include: determining a set of loss values for each of the plurality of GANs; generating a first GAN image using the negative generator by removing the darkest regions of the at least one artifact in the input image; generating a second GAN image using the artifact generator by removing at least one of the color-continuous regions and texture-continuous regions of the at least one artifact in the input image; generating a third GAN image using the partitioning generator by removing the brightest regions of the at least one artifact and adding white patch regions to the at least one artifact in the input image; and generating the at least one intermediate output image without the at least one artifact using at least one of the first GAN image, the second GAN image, and the third GAN image.
[0054] The input image may be a previous input image, and the method may further include: receiving a new input image; and overlaying the binary mask obtained for the previous input image onto the new input image.
[0055] The at least one artifact can be a shadow or a glare.
[0056] According to another aspect of the present invention, an electronic device for processing image data may include: a memory, storing instructions; and a processor configured to execute instructions to perform the following operations: receiving an input image; extracting a plurality of features from the input image, wherein the plurality of features may include texture of the input image, color composition of the input image, and edges in the input image; determining a region of interest (RoI) in the input image including at least one artifact based on the plurality of features; generating at least one intermediate output image by removing the at least one artifact from the input image using a plurality of generative adversarial networks (GANs); generating a binary mask using the at least one intermediate output image, the input image, edges in the input image, and edges in the at least one intermediate output image; classifying the at least one artifact as a desired artifact or an unwanted artifact based on the binary mask; and obtaining a final output image from the input image based on the category of the at least one artifact corresponding to the desired artifact or the unwanted artifact.
[0057] The processor may also be configured to: generate multiple versions of the at least one intermediate output image by changing the gamma value associated with the final output image; determine a Natural Image Quality Evaluator (NIQE) value for each of the multiple versions of the at least one intermediate output image; and display a version of the final output image having the minimum NIQE value among the NIQE values for the multiple versions.
[0058] The processor can also be configured to extract texture, color composition, and edges based on Otsu thresholding, Gaussian blur, and edge detection, respectively.
[0059] The multiple GANs may include a negative generator, an artifact generator, a partition generator, a negative discriminator, a partition discriminator, a refinement generator, and a refinement discriminator.
[0060] The processor may also be configured to: determine a set of loss values for each of the plurality of GANs; generate a first GAN image using the negative generator by removing the darkest region of the at least one artifact in the input image; generate a second GAN image using the artifact generator by removing at least one of the color-continuous regions and texture-continuous regions of the at least one artifact in the input image; generate a third GAN image using the partitioning generator by removing the brightest region of the at least one artifact and adding white patch regions to the at least one artifact in the input image; and generate at least one intermediate output image without the at least one artifact using at least one of the first GAN image, the second GAN image, and the third GAN image.
[0061] The input image may be a previous input image, and the processor may also be configured to: receive a new input image; and overlay the binary mask obtained for the previous input image onto the new input image.
[0062] The at least one artifact can be a shadow or a glare.
[0063] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided, the program being executable by one or more processors to perform a method for processing image data. The method may include: obtaining an input image; identifying types of artifacts included in the input image; generating a plurality of intermediate output images, wherein the artifacts are processed differently according to their types in the plurality of intermediate output images; generating a binary mask based on the plurality of intermediate output images and the input image; and applying the binary mask to the input image to selectively retain or remove the artifacts and obtain a final output image, wherein at least one of the artifacts is retained and at least one other artifact is removed in the final output image.
[0064] The types of artifacts may include the darkest areas of shadows in the input image, continuous color and texture areas of shadows, and the brightest areas of shadows. Generating the plurality of intermediate output images may include: removing the darkest areas of shadows from the input image to generate a first intermediate output image; removing continuous color and texture areas of shadows from the input image to generate a second intermediate output image; and removing the brightest areas of shadows from the input image to generate a third intermediate output image.
[0065] Removing the brightest area of a shadow to generate a third intermediate output image may include adding a white patch to the brightest area of the shadow after removing the brightest area of the shadow.
[0066] The method may further include combining a first intermediate output image, a second intermediate output image, and a third intermediate output image into a combined intermediate image of multiple intermediate output images. Generating the binary mask may include generating the binary mask based on the combined intermediate image.
[0067] Invention Model
[0068] The exemplary embodiments are described in more detail below with reference to the accompanying drawings.
[0069] In the following description, the same reference numerals are used for the same elements, even in different figures. Matters defined in the description (such as detailed constructions and elements) are provided to aid in a comprehensive understanding of the exemplary embodiments. However, it will be apparent that the exemplary embodiments can be practiced without those specifically defined matters. Furthermore, well-known functions or structures are not described in detail because they would be described in an unnecessarily obscure manner. The terms and words used in the following description and claims are not limited to their literal meaning but are used solely by the inventors to enable a clear and consistent understanding of this disclosure. Therefore, it will be apparent to those skilled in the art that the following description providing various embodiments of this disclosure is for illustrative purposes only and is not intended to limit the purpose of this disclosure as defined by the appended claims and their equivalents.
[0070] Unless otherwise indicated, the term "or" as used herein means non-exclusive or. Expressions such as "at least one of..." modify the entire list of elements when preceding it, rather than individual elements within the list. For example, the expression "at least one of a, b, and c" should be understood to include only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or any variation of the foregoing examples. The term "(one or more) images" as used herein refers to one or more images. For example, "(one or more) images" may indicate one image or multiple images.
[0071] Embodiments in this disclosure may be described and illustrated in blocks that perform one or more of the described functions. These blocks (which may be referred to herein as units or modules, etc.) are physically implemented by analog or digital circuitry (such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuitry, etc.) and may optionally be driven by firmware. The circuitry may, for example, be implemented in one or more semiconductor chips or on a substrate support such as a printed circuit board. The circuitry constituting a block may be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware performing some functions of the block and a processor performing other functions of the block. Without departing from the scope of this disclosure, each block of an embodiment may be physically divided into two or more interactive and discrete blocks. Similarly, without departing from the scope of this disclosure, the blocks of an embodiment may be physically combined into more complex blocks.
[0072] Although the terms first, second, etc., may be used in this document to describe various elements, these elements should not be limited to these terms. These terms are generally used only to distinguish one element from another.
[0073] Therefore, embodiments of this document disclose a method and electronic apparatus for managing image artifacts. The method includes receiving an input image and extracting multiple features from the input image. The features include the texture of the input image, the color composition of the input image, and edges in the input image. Furthermore, the method includes determining a region of interest (RoI) in the input image, including artifacts, based on the features, and generating an intermediate output image by removing artifacts using multiple GANs. Artifacts can be shadows or glare. Additionally, the method includes generating a binary mask using the intermediate output image, the input image, an image showing edges in the input image, and an image showing edges in the intermediate output image, and obtaining a final output image by applying the generated binary mask to the input image.
[0074] In related technologies, electronic devices allow users to edit input images to remove shadows. However, shadow removal is not very effective, thus reducing image quality. Furthermore, there are no methods or systems to address all forms of artifacts (such as shadows, glare, etc.). Related methods and systems focus more on allowing users to correct exposure and edit images by changing contrast or color composition to remove or mitigate artifacts.
[0075] Unlike related methods and systems, the electronic device according to the embodiment uses multiple GANs to generate a binary mask that defines portions of artifacts as desired and unwanted. Furthermore, the binary mask is overlaid on the input image to remove or retain artifacts or portions of artifacts. Therefore, the electronic device according to the embodiment intelligently manages artifacts in the input image.
[0076] Figure 3 This is a block diagram of an electronic device 100 for managing image artifacts as described in the embodiment.
[0077] Reference Figure 3 The electronic device 100 may be, but is not limited to, a laptop computer, a handheld computer, a desktop computer, a mobile phone, a smartphone, a personal digital assistant (PDA), a tablet computer, a wearable device, an Internet of Things (IoT) device, a virtual reality device, and an immersive system. In an embodiment, the electronic device 100 includes a memory 120, a processor 140, an image artifact controller 160, and a display 180.
[0078] Memory 120 is configured to store one or more binary masks generated using one or more intermediate output images, one or more input images, an image showing the edges in one or more input images, and an image showing the edges in one or more intermediate output images. Memory 120 may include non-volatile storage elements. Examples of such non-volatile storage elements may include magnetic hard disks, optical disks, floppy disks, flash memory, or electrically programmable memory (EPROM) or electrically erasable programmable memory (EEPROM). Additionally, in some examples, memory 120 may be considered a non-transitory storage medium. The term "non-transitory" may indicate that the storage medium is not implemented in a carrier wave or propagating signal. However, the term "non-transitory" should not be construed as meaning that memory 120 is non-removable. In some instances, a non-transitory storage medium may store data that may change over time (e.g., in random access memory (RAM) or cache memory).
[0079] Processor 140 may include one or more processors. The one or more processors may be general-purpose processors (such as central processing unit (CPU), application processor (AP), etc.), graphics-only processing units (such as graphics processing unit (GPU), vision processing unit (VPU)), and / or artificial intelligence (AI) dedicated processors (such as neural processing unit (NPU), AI accelerator, or machine learning accelerator). Processor 140 may include multiple cores and is configured to execute instructions stored in memory 120.
[0080] In an embodiment, the image artifact controller 160 includes an image feature extraction controller 162, an artifact management generative adversarial network (GAN) 164a-164g, a mask classifier 166, an image quality controller 168, and an image effect controller 170. Although Figure 3 The image artifact controller 160 is shown as a separate or different element from the processor 140, but the image artifact controller 160 may be incorporated into the processor 140 or may be implemented as another processor.
[0081] Image feature extraction controller 162 is configured to extract features from one or more input images. Features may include, but are not limited to, the texture of one or more input images, the color composition of one or more input images, and edges in one or more input images. The texture of one or more input images may be extracted, for example, by Otsu thresholding, which involves returning a single intensity threshold (e.g., ...) that separates pixels in one or more input images into foreground and background. Figure 5c Image feature extraction controller 162 (described herein). The color composition of one or more input images is extracted, for example, by Gaussian blurring, which highlights different regions of the input images including regions of interest (ROIs) with artifacts. Edge detection of the input images is performed using any edge detection technique, such as Sobel edge detection, Canny edge detection, Prewitt edge detection, Roberts edge detection, and fuzzy logic edge detection. The extracted features are sent by image feature extraction controller 162 to the corresponding artifact management GAN164a-164g. Image feature extraction controller 162 is implemented by processing circuitry (such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuitry, etc.) and may optionally be driven by firmware. The circuitry may be implemented, for example, in one or more semiconductor chips or on a substrate support such as a printed circuit board.
[0082] The artifact management GANs 164a-164g include a first GAN 164a, a second GAN 164b, a third GAN 164c, a negative discriminator 164d, a division discriminator 164e, a thinning generator 164f, and a thinning discriminator 164g. The artifact management GANs 164a-164g are implemented by processing circuitry (such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits, etc.) and may optionally be driven by firmware. The circuitry may be implemented, for example, in one or more semiconductor chips or on a substrate support such as a printed circuit board.
[0083] Artifact management GANs 164a-164g are configured to determine the Regions of Interest (ROIs) in one or more input images, including artifacts, based on extracted features. Artifacts can be, for example, shadows and glare. Furthermore, artifact management GANs 164a-164g are configured to determine a set of loss values for each of the first GAN 164a, the third GAN 164c, and the thinning generator 164f. Additionally, the first GAN 164a is configured to generate a first GAN image by removing the darkest regions of artifacts from one or more input images and using the texture of one or more input images extracted by the image feature extraction controller 162. The first GAN 164a is, for example, a negative generator. The second GAN 164b is configured to generate a second GAN image by removing color-continuous regions and texture-continuous regions of artifacts from one or more input images and using the color composition of one or more input images extracted by the image feature extraction controller 162. The second GAN 164b is, for example, an artifact generator. The third GAN 164c is configured to generate a third GAN image by removing the brightest regions of artifacts and adding white patch regions to the artifacts in one or more input images, and using edges in one or more input images extracted by the image feature extraction controller 162. The third GAN 164c is, for example, a segmentation generator. Furthermore, the negative discriminator 164d is configured to receive the first GAN image and one or more input images and send its output to the thinning GAN 164f. Similarly, the segmentation discriminator 164e is configured to receive the third GAN image and one or more input images and send its output to the thinning GAN 164f. The thinning GAN 164f receives the outputs from the negative discriminator 164d, the segmentation discriminator 164e, and the second GAN 164b. Furthermore, the thinning GAN 164f is connected to the thinning discriminator 164g to generate an intermediate output image (in the image obtained by completely eliminating artifacts from one or more input images). Figures 6a-6h (See detailed description in the text). The artifact management GAN164a-164g is trained using multiple images containing various forms of artifacts to generate the corresponding output images.
[0084] Mask classifier 166 is configured to receive an intermediate output image, an image showing edges in the intermediate output image, an input image, and an image showing edges in at least one input image, and to generate one or more binary masks. The binary mask is used to classify artifacts or portions of artifacts as desired or unwanted artifacts. Desired artifacts may be those that do not need to be removed from the image, and unwanted artifacts may be those that are suitable for removal from the image. The white portion of the binary mask indicates the portion of the artifact that will be removed from the input image, i.e., the unwanted artifact. For example, a shadow on text in an image of a document is classified as an unwanted artifact because the shadow makes the text difficult to read. In another example, sand dunes in a desert landscape are considered desirable because the shadow provides depth to the landscape. Furthermore, mask classifier 166 is configured to apply the binary mask to each pixel of the input image to preserve or remove appropriate portions of artifacts. The mask classifier 166 is implemented by processing circuitry (such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuitry, etc.) and may optionally be driven by firmware. The circuitry may be implemented, for example, in one or more semiconductor chips, or on a substrate support such as a printed circuit board.
[0085] Image quality controller 168 is configured to generate multiple versions of the output image by varying the gamma value associated with the final output image. Furthermore, a graph with a Natural Image Quality Evaluator (NIQE) value is plotted for each version of the output image to determine the optimal quality output image. The optimal quality output image is the version of the final output image that has the minimum NIQE value in the stable region of the gamma graph (within...). Figures 8a-8f (See detailed description below). The image quality controller 168 is implemented by processing circuitry (such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuitry, etc.) and may optionally be driven by firmware. The circuitry may, for example, be implemented in one or more semiconductor chips or on a substrate support such as a printed circuit board.
[0086] Image effects controller 170 is configured to retrieve a binary mask from memory 120 and apply the binary mask directly to any other image to add effects associated with artifacts. For example, image effects controller 170 may add glare or shadows to an image. Image effects controller 170 is implemented by processing circuitry (such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuitry, passive electronic components, active electronic components, optical components, hardwired circuitry, etc.) and may optionally be driven by firmware. The circuitry may, for example, be implemented in one or more semiconductor chips or on a substrate support such as a printed circuit board.
[0087] In this embodiment, the display 180 is configured to display a final output image containing portions of artifacts classified as desired. In other words, artifacts classified as desired are not removed from the image, but artifacts classified as unwanted are removed. The display 180 is implemented using touch-sensitive technology and includes one of a liquid crystal display (LCD), a light-emitting diode (LED), or the like.
[0088] although Figure 3 The hardware components of electronic device 100 are shown, but it should be understood that other embodiments are not limited thereto. In other embodiments, electronic device 100 may include fewer or more components. Furthermore, the designations or names of components are for illustrative purposes only and do not limit the scope of the invention. One or more components may be combined together to perform the same or substantially the same function.
[0089] Figure 4a This is a flowchart illustrating a method for managing image artifacts by an electronic device 100 according to embodiments disclosed herein.
[0090] Reference Figure 4a In operation 421, electronic device 100 receives an input image. For example, in... Figure 3 In the electronic device 100 shown, the image artifact controller 160 is configured to receive an input image.
[0091] In operation 422, electronic device 100 extracts features from the input image. For example, in... Figure 3 In the illustrated electronic device 100, the image artifact controller 160 is configured to extract features from an input image. These features may include the texture of the input image, the color composition of the input image, and edges within the input image.
[0092] In operation 423, electronic device 100 determines the Region of Interest (ROI) in the input image, including artifacts, based on extracted features. For example, in... Figure 3 In the electronic device 100 shown, the image artifact controller 160 is configured to determine the ROI in the input image, including artifacts, based on extracted features.
[0093] In operation 424, electronic device 100 generates an intermediate output image by using multiple GANs to remove artifacts. For example, in... Figure 3 In the electronic device 100 shown, the image artifact controller 160 is configured to generate an intermediate output image by removing artifacts using multiple GANs 164a-164g.
[0094] In operation 425, electronic device 100 uses an intermediate output image, an input image, an image showing edges in the input image, and an image showing edges in the intermediate output image to generate a binary mask and determine whether artifacts are desired or unwanted. For example, in... Figure 3 In the illustrated electronic device 100, the image artifact controller 160 is configured to generate a binary mask using an intermediate output image, an input image, an image showing the edges in the input image, and an image showing the edges in the intermediate output image, and to determine whether the artifacts are desired or not.
[0095] In operation 426, electronic device 100 classifies at least one artifact into a first artifact category or a second artifact category based on a binary mask of the input image. For example, as mentioned in operation 425, the artifact can be classified as a desired artifact or an unwanted artifact. In this case, the first artifact category can be a desired artifact, and the second artifact category can be an unwanted artifact. The binary mask can indicate whether an artifact or a portion of an artifact is a desired artifact or an unwanted artifact. The white portion of the binary mask can indicate the portion of the artifact that will be removed from the input image, i.e., the unwanted artifact, and the black portion of the binary mask can indicate the portion of the artifact that will be retained. The white portion of the binary mask can have a value of 1, and the black portion of the binary mask can have a value of 0.
[0096] In operation 427, the electronic device 100 obtains a final output image by processing the input image based on the category of at least one artifact corresponding to a first artifact category and / or a second artifact category.
[0097] Figure 4b This is a flowchart illustrating a method for managing image artifacts by an electronic device 100 according to embodiments disclosed herein.
[0098] Reference Figure 4b In operation 402, electronic device 100 receives one or more input images. For example, in... Figure 3 In the electronic device 100 shown, the image artifact controller 160 is configured to receive one or more input images.
[0099] In operation 404, electronic device 100 extracts features from one or more input images. For example, in... Figure 3 In the electronic device 100 shown, the image artifact controller 160 is configured to extract features from one or more input images.
[0100] In operation 406, electronic device 100 determines the Region of Interest (ROI) in one or more input images, including artifacts, based on extracted features. For example, in... Figure 3In the electronic device 100 shown, the image artifact controller 160 is configured to determine the ROI in one or more input images containing artifacts based on extracted features.
[0101] In operation 408, electronic device 100 generates an intermediate output image by using multiple GANs to remove artifacts. For example, in... Figure 3 In the electronic device 100 shown, the image artifact controller 160 is configured to generate an intermediate output image by removing artifacts using multiple GANs 164a-164g.
[0102] In operation 410, electronic device 100 uses an intermediate output image, an input image, an image showing edges in the input image, and an image showing edges in the intermediate output image to generate a binary mask and determine whether artifacts are desired or unwanted. For example, in... Figure 3 In the illustrated electronic device 100, the image artifact controller 160 is configured to generate a binary mask using an intermediate output image, an input image, an image showing the edges in the input image, and an image showing the edges in the intermediate output image, and to determine whether the artifacts are desired or not.
[0103] In operation 412, electronic device 100 obtains a final output image by applying the generated binary mask to one or more input images. For example, in... Figure 3 In the electronic device 100 shown, the image artifact controller 160 is configured to obtain a final output image by applying a generated binary mask to one or more input images.
[0104] The various actions, behaviors, blocks, steps, etc., in the method can be executed in the provided order, in different orders, or simultaneously. Furthermore, in some embodiments, some of the actions, behaviors, blocks, steps, etc., may be omitted, added, modified, skipped, etc., without departing from the scope of the invention.
[0105] Figure 5a This is an example illustrating the extraction of multiple features from one or more input images by an electronic device 100 according to embodiments disclosed herein.
[0106] Reference Figure 5a In operation 502, electronic device 100 receives an input image. In operation 504, image feature extraction controller 162 is used to manipulate the input image by performing standard operations on the image matrix of the input image to highlight artifacts (such as shadows, glare areas) in the input image and preserve texture in the input image.
[0107] Multiple features can be extracted, allowing the texture features 530, edge features 550, and color composition 540 of the input image to be extracted and preserved during the training and inference phases. For example, texture feature 530 could be the result of Otsu thresholding of the input image, color composition 540 could be the result of Gaussian blurring of the input image, and edge feature 550 could be the result of edge detection.
[0108] Removing artifacts from an input image can be considered as removing information from the input image. Therefore, during artifact removal, regions containing artifacts may lose some information (such as texture, color composition, hue, saturation and value (HSV) maps, edges, etc.). Image feature extraction controller 504 (with...) Figure 3 The image feature extraction controller 162 provides more details about the removal information, allowing the removed information to be incorporated into the final output image. Furthermore, based on the structural similarity and functionality of the extracted features, the extracted features are mapped as input to corresponding GANs 164a-164g. For example, texture is mapped to the first GAN 164a, color composition to the second GAN 164b, and the edges of artifacts extracted from the input image are mapped to the third GAN 164c. The extracted features can then be used in the GAN generator.
[0109] Figure 5b This is an example illustrating the use of Gaussian blur to extract the colors of an input image by electronic device 100 according to embodiments disclosed herein.
[0110] Reference Figure 5b The input image is filtered by a Gaussian filter in the image feature extraction controller 504 to obtain a Gaussian blur for extracting the color composition of the input image. Furthermore, the Gaussian blur can be used to extract the texture of the input image. To determine the Gaussian blur, a Gaussian filter with a 3×3 kernel size and convolution are performed on the input image to obtain different smoothed images, thereby highlighting different regions of artifacts in the input image. The Gaussian filter helps preserve the color composition and texture of the input image. The Gaussian blur reduces image noise and detail.
[0111] As shown by reference numeral 542, edges, shadows, and glare in the input image are identified, and in reference numeral 544, the edges of the ROI are labeled 544a. Furthermore, in reference numerals 546 and 548, various curves and patterns depicting the texture of the input image are identified. As described above, the Gaussian filter helps preserve the color composition and texture of the input image. Gaussian blur reduces image noise and detail. Therefore, when the input image passes through the Gaussian filter of the image feature extraction controller 504, the color composition or texture of the input image is extracted from the Gaussian-blurred input image.
[0112] Figure 5c This is an example illustrating the extraction of texture from an input image by electronic device 100 using Otsu thresholding according to embodiments disclosed herein.
[0113] Reference Figure 5c The Otsu thresholding algorithm is applied to the input image to extract its texture. Image thresholding is used to binarize the image based on pixel intensity. The input to this thresholding algorithm is typically a grayscale image and a threshold. The output is a binary image. If the intensity of a pixel in the input image is greater than the threshold, the corresponding output pixel is marked as white (foreground), and if the intensity of the input pixel is less than or equal to the threshold, the output pixel location is marked as black (background). In Otsu thresholding, the weights, mean, and variance {W} used for the foreground (greater than Vth) are determined. f Weight, μ f Mean and : variance for the foreground and weights, mean, and variance for the background (below Vth) {W} b Weight, μ b Mean and : Variance background}. Then determine the variance within each category. . The highest optimal Vth is selected as the threshold, and the background pixels are set to zero, while the foreground is set to the highest pixel value.
[0114] ...(1)
[0115] In operation 532, an input image is received, and in operation 534, a histogram of the input image is determined, wherein N is the sum of the values of all pixels in the input image. bins =20. In operation 536, based on the histogram, electronic device 100 determines the internal variance for each pixel in the input image. Furthermore, as indicated by reference numeral 537, it is assumed that the internal variance is highest at pixel value 165, therefore pixel value 165 is considered the threshold. Therefore, pixel values less than 165 are assigned as pixel values of 0, and pixel values greater than 165 are assigned as pixel values of 255. In other words, pixels with values less than 165 are classified as foreground, and pixels with values greater than 165 are classified as background. Thus, the background and foreground in the input image are separated by processing the input image according to Otsu thresholding. In operation 538, the result of Otsu thresholding for the input image is provided as an output image. The result of Otsu thresholding may include texture extracted from the input image.
[0116] Figure 6a The diagram illustrates the architecture of multiple GANs 164a-164g according to embodiments disclosed herein, including those for obtaining artifact-free intermediate output images by electronic device 100.
[0117] Reference Figure 6a The multiple GANs 164a-164g include the first GAN 164a, the second GAN 164b, the third GAN 164c, the negative discriminator 164d, the split discriminator 164e, the refinement generator 164f, and the refinement discriminator 164g.
[0118] At point 1, the first GAN 164a can be a negative generator that receives an input image and the texture of the input image extracted by the image feature extraction controller 162. The result of Otsu thresholding of the input image can be used as the texture feature of the input image and input into the first GAN 164a. The first GAN 164a can generate negative samples for the input. For example, the first GAN 164a can remove the darkest regions of artifacts in the input image and generate a first GAN image as a negative image. At point 2, the second GAN 164b can be an artifact generator that receives an input image and a Gaussian blurred image, the Gaussian blurred image including features composed of colors of the input image extracted by the image feature extraction controller 162.
[0119] The second GAN 164b removes color and texture continuity regions of artifacts from the input image and generates a second GAN image. At step 3, the third GAN 164c is a segmentation generator that receives the input image and the edges of the input image detected by the image feature extraction controller 162. The results of edge detection can be used as the edges of the input image. The third GAN 164c removes the brightest regions of artifacts but adds white patches and generates a third GAN image. In operation 4, the negative discriminator 164d receives the first GAN image and compares it with the input image to determine if the first GAN image is genuine. Similarly, in operation 6, the segmentation discriminator 164e receives the third GAN image and compares it with the input image to determine if the third GAN image is genuine.
[0120] In operation 5, the first GAN image and the input image are fed into an adder. Similarly, in operation 7, the third GAN image and the input image are fed into a multiplier. Then, in operation 8, the outputs from the multiplier and adder are fed into the adder to generate the combined final image. The combined final image is then passed through a thinning generator 164f to thin the image (operation 9), and in operation 10, the image is passed through a thinning discriminator 164g to generate an intermediate output image by removing artifacts. The intermediate output image is then passed through a mask classifier 166 to determine whether the artifacts / artifact portions in the input image are desired or unwanted.
[0121] Furthermore, based on whether the artifact is a desired or unwanted artifact, the electronic device 100 assigns different weights and preferences to the corresponding GAN. The weights of each GAN are determined based on variable loss calculations.
[0122] Adversarial Loss: This is the joint adversarial loss for multiple GANs 164a-164g:
[0123] Loss Adv =[{log(D(I in ))+log(1-D(I G4 ))+log(D(I out -I in ))+log(D(I G1 ))+log(D(I out / I in ))+log(D(I G3 ))}] ..... (2)
[0124] Segmentation loss: Calculate the L1 norm (absolute loss) between the segmented map and the image generated by the third GAN 164c segmentation generator.
[0125] Loss Div =abs((I out / I in )-I G3 (3)
[0126] Negative Loss: A measure of the absolute loss between the negative image and the image generated by the first GAN 164c negative generator. A measure of brightness in the shadow regions is given.
[0127] Loss Neg =abs((I out / I in )-I G1 (4)
[0128] Feature loss: The absolute loss calculated between the features of the original image and the image generated by the thinning generator 164f.
[0129] Loss feat =abs(I infeat / I (G4(feat) (5)
[0130] Therefore, the total loss is calculated by summing all losses according to the weights provided based on the type of artifact:
[0131] Loss Total = Loss Adv. + λ1 Loss Div. + λ2 Loss Neg. + λ3 Loss feat.. .... (6)
[0132] Among them I out = Output pairs of images provided during training, I in =Input pairs of images provided during training, D(I) Gi ) = the discriminator loss of the i-th generator with respect to the output image, I Gi = Output from the i-th generator λ 1. λ 2 and λ 3 is a hyperparameter and is appropriately used to train the model and obtain the desired output. Loss Adv It uses the standard loss that is generally used by each GAN network as suggested in the original GAN paper, but the results are improved by using other losses.
[0133] Because multiple losses need to be computed for each of the multiple GAN 164a-164g pairs, the training time is increased by a factor of three in terms of performance and computation time when using multiple GAN164a-164g pairs. However, when inferring the output, the computation remains the same because inference is obtained only using a refined discriminator 164g that gives intermediate output images by completely removing artifacts.
[0134] Each individual GAN in the multiple GAN 164a-164g performs the task separately, but the combination of multiple GAN 164a-164g provides the best possible results. Due to the use of multiple GAN 164a-164g, the electronic device 100 can easily handle both complex and simple scenes. Furthermore, the multiple GAN 164a-164g performing multiple functions allows the electronic device 100 to extract the output of a generator and use that output for other purposes if needed. For example, in the case of a negative generator, the user may be able to directly use the negative of an image.
[0135] Figure 6b The architectures of a first GAN 164a and a third GAN 164c of an image artifact controller 160 according to embodiments disclosed herein are shown.
[0136] Reference Figure 6bThe architectures of GAN 164a and GAN 164c are the same. However, the inputs to GAN 164a and GAN 164c are different, and therefore their outputs are also different. GAN 164a is a negative generator that takes the input image and Otsu thresholded texture extraction as input. GAN 164c is a segmentation generator that takes the input image and detected edges as input. GAN 164a and GAN 164c provide additional support to the master generator (i.e., GAN 164b and thinning generator 164f).
[0137] Both the first GAN 164a and the third GAN 164c include networks with attention blocks and removal blocks. Attention blocks selectively choose what the network wants to observe, locate artifacts in the input image, and focus the attention of the removal block on the detected regions. Each block is refined in a coarse-to-fine manner to handle complex scenes captured in the images. The subsequent three blocks within the attention and removal blocks provide a good trade-off between performance and complexity.
[0138] Figure 6c The architecture of the attention block and removal block of each of the first GAN 164a and the third GAN 164c according to embodiments disclosed herein is shown.
[0139] Reference Figure 6c Attention blocks provide attention (i.e., observation and localization of ROIs including artifacts) to make processing more focused. Attention blocks consist of ten (10) convolutional layers (Conv) with batch normalization (BN) and leaked corrected linear unit (ReLU) activation functions (Conv+BN+Leaked ReLU). Attention blocks also include long short-term memory (LSTM) layers (in... Figure 6d (Further explanation follows) to utilize information from previous outputs in the preceding steps, and learn from the previous and current states required in the generator, and generate attention maps for subsequent output levels.
[0140] The removal block generates an artifact-free image. The removal block consists of eight (8) (Conv+BN+Leaked ReLU) layers to extract multiple features from the image. Additionally, the removal block includes eight deconvolutional layers (Deconv+BN+Leaked ReLU) with batch normalization and leaked ReLU activation functions to map to a specific distribution. Skip connections are used for multiple channels and to preserve the contextual information of the layer. The final convolutional layer, two (Conv+BN+Leaked ReLU) layers, extracts the feature map after deconvolution. The final convolutional and sigmoid layers transform the feature map into a 3-channel space of the same size as the input.
[0141] Figure 6dThe architecture of a long short-term memory (LSTM) for an attention block of a first GAN 164a according to an embodiment disclosed herein is shown.
[0142] Reference Figure 6d This paper presents an LSTM for the attention block of the first GAN 164a. LSTM is an artificial recurrent neural network (RNN) architecture used for deep learning. LSTM has feedback connections and can process single data points (such as images). LSTM has the ability to add or remove cell states, carefully tuned by gates. The LSTM for the attention block includes 4 gates and 3 sigmoid and 1 hyperbolic tangent (tanh) classifiers for protection and control. The sig gates provide values between 0 and 1. In the proposed method, the sig gates provide values close to 0 for the shadow and glare portions of the image; that is, the shadow and glare portions of the image are discarded and not used in subsequent layers of the first GAN 164a.
[0143] Figure 6e A schematic diagram of a second GAN 164b and a thinning generator 164f of an image artifact controller 160 according to embodiments disclosed herein is shown.
[0144] Reference Figure 6e The schematic diagrams of the second GAN 164b and the thinning generator 164f are the same. However, the inputs to the second GAN 164b and the thinning generator 164f are different, and therefore the corresponding outputs are different. The second GAN 164b removes artifacts in the first pass, and the thinning generator 164f of the image artifact controller 160 combines the outputs from the previous stage generator and modifies the output from a coarse output to a fine output.
[0145] Figure 6f The detailed internal architecture of the second GAN 164b and the thinning generator 164f of the image artifact controller 160 according to embodiments disclosed herein is shown.
[0146] Reference Figure 6f The detailed internal architecture of the second GAN 164b and the refinement generator 164f is the same. However, the inputs of the second GAN 164b and the refinement generator 164f are different, and therefore the corresponding outputs are also different.
[0147] Each of the second GAN 164b and the refinement generator 164f includes six (6) (ReLU + 2 Conv + average pooling) blocks (650) in the encoder layer, which performs down-transformation after each block and skips the decoder layer connected to perform up-transformation after each block. The last bottom layer has 15 (ReLU + 2 Conv + average pooling) blocks (650). At the output stage, a tanh classifier is present to transform the image into a three (3) channel output image.
[0148] The second GAN 164b receives the input image and Gaussian blur features associated with the color composition of artifacts in the input image. The output of the second GAN 164b is generated by removing color and texture continuity regions of the artifacts. The thinning generator 164f receives the first GAN image, the second GAN image, and the third GAN image as input and thins the output of the previously combined GANs.
[0149] Figure 6g The diagram illustrates the transformations associated with multiple layers of the second GAN 164b and the refinement generator 164f according to embodiments disclosed herein.
[0150] Reference Figure 6g The individual layers of the second GAN 164b and the refinement generator 164f are provided. An encoder layer performing down-conversion after each block is shown. A decoder layer performing up-conversion after each block is shown. The encoder and decoder layers are connected using skip connections for multiple channels and retain the layer's context information.
[0151] Figure 6h The architecture of the negative discriminator 164d, the segmentation discriminator 164e, and the refinement discriminator 164g of the image artifact controller 160 according to embodiments disclosed herein is shown.
[0152] Reference Figure 6h The negative discriminator 164d, the dividing discriminator 164e, and the refining discriminator 164g have the same architecture. However, the inputs to each discriminator are different, and therefore the corresponding outputs are also different.
[0153] The negative discriminator 164d receives the input image and the first GAN image as input. The first GAN image is generated by the first GAN 164a by removing the darkest regions of artifacts from the input image. The negative discriminator 164d concatenates the input image and the first GAN image. Furthermore, the concatenated image is passed through several levels of convolution to determine whether the first GAN image is real or fake. Real images are forwarded to the next level, while fake images are discarded.
[0154] Similarly, the discriminator 164e receives the input image and the third GAN image as input. The third GAN image is generated by removing the brightest regions of artifacts from the input image and adding white patches.
[0155] The refined discriminator 164g provides the final discriminator for the intermediate output image by removing all artifacts from the input image.
[0156] Figure 7a This is a flowchart illustrating a method for determining whether an artifact is desired or undesirable by a mask classifier 166, according to embodiments disclosed herein.
[0157] Mask classifier 166 was used to preserve artifacts that were necessary for the image but were removed by multiple GANs 164a-164g.
[0158] Reference Figure 7a In operation 702, the electronic device 100 determines whether artifacts in the intermediate output image should be removed.
[0159] If the electronic device 100 determines that artifacts in the intermediate output image should be removed, in operation 704, the electronic device 100 obtains the output from the thinning generator 164f, and in operation 706, the electronic device obtains the input image. In operation 708, the electronic device 100 performs Canny edge detection on the input image and the output from the thinning generator 164f, respectively. In operation 712, the electronic device 100 obtains the Canny input image, and in operation 714, the electronic device 100 obtains the Canny output from the thinning generator 164f.
[0160] Furthermore, in operation 710, a grayscale subtraction output is obtained using the output of the thinning generator 164f and the input image. Therefore, in operation 716, the input binary mask is obtained.
[0161] In operation 718, the mask classifier 166 uses the Canny input image, the Canny output from the thinning generator 164f, and the input binary mask, and in operation 720, obtains an output binary mask that intelligently determines the portions of artifacts that need to be retained and the portions of artifacts that need to be removed. If the electronic device 100 determines that artifacts in the intermediate output image should not be removed, in operation 722, an output binary mask containing artifacts or portions of artifacts is generated. In this case, the white portion of the binary mask indicates the portion of the artifact that will ultimately be removed. Furthermore, the white portion of the binary mask is the area that will be copied from the output of the thinning generator 164f onto the input image. In other words, unwanted artifacts indicated by the white portion of the binary mask are removed from the input image, and the unwanted artifacts are replaced by the corresponding portions of the output from the thinning generator 164f. Desired artifacts indicated by the black portion of the binary mask are not removed from the input image, and the desired artifacts are retained without being replaced by the output from the thinning generator 164f.
[0162] Furthermore, in operation 724, the electronic device 100 performs pixel-by-pixel overlay of the output binary mask on the input image to obtain the final enhanced output image. The method performed by the mask classifier 166 can be extended to user-guided selection of artifact regions, where only the white portion of the binary mask is copied from the output of the thinning generator 164f onto the input image to generate the output image.
[0163] Figure 7b This is an example illustrating the processing of a mask classifier 166 according to an embodiment disclosed herein.
[0164] Masking allows control over the transparency level of artifacts in the input image without affecting the actual background. Combined with... Figure 7a Reference Figure 7bIn operation 11, mask classifier 166 receives an intermediate output image from thinning generator 164f; in operation 12, it receives a canny output from thinning generator 164f; in operation 13, it receives an input image with shadows; and in operation 14, it receives a canny input image. These inputs provide key information to mask classifier 166 for making pixel-by-pixel classification decisions from previous artifact masks to intelligent artifact masks. Furthermore, in operation 15, mask classifier 166 generates an output binary mask based on the received multiple inputs. The output binary mask indicates the desired and unwanted portions of artifacts (such as shadows) in the input image. In operations 16 and 17, the output binary mask is superimposed on the input image to obtain the final output image. In other words, unwanted artifacts indicated by the white portions of the binary mask are removed from the input image, and the corresponding portions from the output of thinning generator 164f are used to replace the unwanted artifacts. The desired artifacts, indicated by the black portions of the binary mask, are not removed from the input image, and are preserved without being replaced by the output from the thinning generator 164f. In the final output image, the shadow portions obscuring the sharp appearance of the object's face are removed, while the shadow portions in the object's neck that provide depth to the image are preserved. The mask classifier 166 is a neural network model trained on multiple images to determine whether artifacts are desired or not.
[0165] Figure 7c These are examples illustrating desired and unwanted artifacts in images according to embodiments disclosed herein.
[0166] Shadows can help draw attention to specific points in a composition. They can reveal form or conceal features that are best left unseen. They can also be used to add a touch of drama, emotion, interest, or mystery to an image. Furthermore, shadows can emphasize light and draw attention to highlights in an image. Therefore, depending on the image, it may be necessary to remove shadows from the image, and sometimes it may be necessary to retain shadows in the image.
[0167] Reference Figure 7c Multiple images with artifacts are used to train a mask classifier 166 to determine desired and unwanted artifacts. Figure 7c The artifact in image (a) is the glare portion (shown in the lower part of the image of the reflected flower), which makes the image complete and also enhances its aesthetic value. Therefore, the glare in the image is considered the desired artifact. Figure 7cThe artifacts in image (b) are the shadows associated with each dune, which provide depth information and aesthetic value to the image. Therefore, shadows in the image are considered desirable artifacts. Generally, removing shadows that provide depth meaning in landscape images results in a loss of the image's 3D properties; shadows that capture the main object in the image; shadows that enhance emotion and bring a true-to-life tone to the image; and monochrome image shadows that provide contrast and thus enhance the image's meaning and value are all considered desirable shadows.
[0168] Figure 7c The artifacts in image (c) are shadows cast by the electronic device 100 on the book due to an incorrect angle at which the image was captured. These shadows on the book or any text hinder OCR reading and the recovery and authentication of any document. Figure 7c The artifacts in image (d) are the shadows of multiple trees on the road. Distorted, unstructured shadows in the image reduce the visibility and attractiveness of a good object image. Similarly, shadows that occlude the main object, facial expressions, and image features are considered unwanted shadows.
[0169] Furthermore, the proposed method can be used to suggest various directions and other camera effect values (such as white balance, exposure, ISO, etc.). Shadows are formed when there is occlusion in the path of light. Therefore, the direction of the shadow indicates the direction from which the light is coming, and suggestions can be shown to the user regarding which direction to move in to prevent shadows from ruining the image. The higher the light intensity, the darker the shadow, and vice versa. Therefore, by evaluating the shadow, the electronic device 100 can automatically suggest various camera controls (such as white balance, ISO, exposure, etc.) to deliver a naturally enhanced and meaningful photograph.
[0170] Figure 8a The architecture of an image quality controller 168 for determining an image with optimal quality after a binary mask is overlaid on an input image is shown according to embodiments disclosed herein.
[0171] Reference Figure 8a This provides the architecture for an image quality controller 168. The image quality controller 168 is a natural image quality evaluator (NIQE) that measures the quality of images with arbitrary distortion. The image quality controller 168 uses a trained model to calculate a quality score. The model is trained using the same predictable statistical features known as natural scene statistics (NSS).
[0172] NSS is based on normalized brightness coefficients in the spatial domain and is modeled as a multidimensional Gaussian distribution. Distortions due to perturbations in the Gaussian distribution may exist.
[0173] The quality score for each image is provided using the following equation:
[0174] ...(7)
[0175] Where v1, v2 and ∑1, ∑2 are the mean vector and covariance matrix of the natural Gaussian model and the Gaussian model of the distorted image, respectively. Therefore, unlike traditional methods and systems, in the proposed method, the electronic device 100 not only intelligently applies binary masks based on the determination of desired and unwanted artifacts, but also selects the optimal image based on gamma correction.
[0176] Figure 8b This is a flowchart illustrating a method for determining an image of optimal quality after a binary mask is superimposed on an input image, according to embodiments disclosed herein.
[0177] Reference Figure 8b In operation 810, electronic device 100 receives one or more input images with applied binary masks and determines the step size for performing gamma correction. Furthermore, in operation 820, electronic device 100 plots a graph of gamma values varying between 0 and 10 with a determined step size. In operation 830, electronic device 100 extracts data from the graph (in...) Figure 8d (As explained in the text) Find the ideal Image Quality Evaluator (IQE). In operation 840, the electronic device 100 determines and sets the delta deviation from the ideal IQE. In operation 850, the electronic device 100 determines an IQE close to the ideal from the graph.
[0178] Figure 8c This is a table illustrating NIQE relative to gamma for obtaining various images by changing gamma values according to embodiments disclosed herein.
[0179] Figure 8d This is a graph showing the NIQE versus gamma for obtaining various images by changing the gamma value according to embodiments disclosed herein.
[0180] Reference Figure 8c and Figure 8d The electronic device 100 determines the graph by generating images with variable gamma values. For each image with a specific gamma value, a corresponding Natural Image Quality Evaluator (NIQE) value is determined, such as... Figure 8c As shown. Furthermore, the lowest NIQE value in the stable region of the graph was determined, and the image including that gamma value was selected as the best quality image.
[0181] from Figure 8d In this study, a gamma value of 1.5 was determined to be the most stable; therefore, after applying a binary mask, the image associated with a gamma value of 1.5 was selected as the best quality image. Figure 8e The process for determining the gamma value at the stability point is explained in the document, and... Figure 8f This provides the best quality images you can get.
[0182] Figure 8e This is a flowchart illustrating a method for determining the ideal IQE for obtaining optimal quality from the NIQE relative to a gamma plot, according to embodiments disclosed herein.
[0183] Reference Figure 8e , combined Figure 8c and Figure 8d In operation 832, electronic device 100 acquires three (3) image IQEs, namely input image IQE ( Figure 8c and Figure 8d The IQE of point A in the image), and the first image quality IQE with a gamma value of 10 ( Figure 8c and Figure 8d The IQE of point B in the image and the second image quality IQE with a gamma value of 5. Figure 8c and Figure 8d (IQE of point C in the middle).
[0184] In operation 834, electronic device 100 determines whether a first image quality IQE is greater than a second image quality IQE. In operation 836, in response to determining that the first image quality IQE (IQE(B)) is greater than the second image quality IQE (IQE(C)), electronic device 100 calculates a threshold, which is calculated by finding a point with a lower NIQE between the maximum gamma (gamma=10) and half of the maximum gamma (i.e., gamma=5):
[0185] Threshold = (IQE(C) - IQE(A)) / 2
[0186] Furthermore, the ideal NIQE is calculated by subtracting half the difference between the above-mentioned comparison result and the initial NIQE from the initial NIQE, as follows:
[0187] Ideal IQE = IQE(A) + threshold (8)
[0188] In operation 838, in response to determining that the first image quality IQE is not greater than the second image quality IQE, the electronic device 100 calculates the threshold as follows:
[0189] Threshold = (IQE(B) - IQE(A)) / 2 (9)
[0190] Ideal IQE = IQE(A) + threshold (10)
[0191] Therefore, in Figure 8dIn this process, electronic device 100 determines a point with a lower NIQE between the maximum gamma and half of the maximum gamma. The ideal NIQE is 17.87 - (17.87 - 12.65) / 2 = 15.26. The ideal NIQE is searched where the absolute difference in delta is |15.54 - 15.26| < 0.5, and the stable point D is obtained with gamma = 1.5.
[0192] Figure 8f This illustrates a scenario where gamma-corrected color compensation is used in an input image including artifacts, according to embodiments disclosed herein.
[0193] Reference Figure 8f , combined Figure 8e NIQE provides 1.5 as the optimal gamma correction value for the input image. Multiple output images are provided by varying the gamma correction value, and the optimal image is found at gamma correction value = 1.5 in both shadow and glare conditions.
[0194] Figure 9a This is an example illustrating a scenario of optical character recognition in related technologies.
[0195] OCR is commonly used to assist blind people in reading, to save books and scripts in digital format, and for electronic signatures and processing digital documents. Typically, with scanned images, users face various lighting conditions and the resulting artifacts in the scanned images. (See reference...) Figure 9a The presence of shadows in the image obscures characters, causing OCR to provide unacceptably incomplete and erroneous results. Therefore, users are forced to perform tedious image editing tasks to meet digital document requirements.
[0196] Figure 9b This is an example illustrating a scenario in which OCR is performed by an electronic device 100 according to an embodiment disclosed herein.
[0197] Reference Figure 9b The electronic device 100 captures clear and well-defined images that provide complete and accurate results that can be used in digital documents.
[0198] Therefore, using the proposed method, the electronic device 100 removes shadows from the input image and then applies OCR to obtain better results with very little editing required.
[0199] Figure 10a This is an example of a scene where shadows are captured along a black background in a related technique.
[0200] Reference Figure 10aThe image of the object was captured and placed along a black background. In images 1, 2, and 3, the shadow along the black background remains visible and has not been removed from the image. Furthermore, the shadow along the black background reduces the focus on the main object in the image, thus distorting the user experience.
[0201] Figure 10b This is an example illustrating a scene captured by electronic device 100 along a black background according to an embodiment disclosed herein.
[0202] Reference Figure 10a , combined Figure 10b Unlike related methods and systems, the electronic device 100 according to the embodiment detects shadows along a black background and intelligently determines whether the shadows are desired or unwanted. In the example case, since shadows along a black background do not add any value to images 1, 2, and 3, the shadows are classified as unwanted and removed. Therefore, even artifacts along a black / dark background can be distinguished and managed by the electronic device 100 according to the embodiment, where the identification and differentiation of artifacts in such cases may be impossible or inefficient in conventional methods.
[0203] Figure 10c This is an example illustrating a scene where an electronic device 100 intelligently overlays artifacts according to an embodiment disclosed herein.
[0204] Reference Figure 10c Consider a user attempting to capture an image in the live preview mode of electronic device 100. In live preview mode, electronic device 100 offers a "smart overlay" option. Smart overlay refers to a binary mask stored in electronic device 100, which the user can choose to apply to any image. Furthermore, in step 1302c, the user manually selects the "smart overlay" option via a drop-down menu, and the binary mask is automatically applied to the image while it is being captured.
[0205] Figure 10d This is an example illustrating a scene of artifact removal in a live preview mode of an electronic device 100 according to embodiments disclosed herein.
[0206] Reference Figure 10d The electronic device 100 allows the user to choose whether to capture an image in live preview mode with or without shadows. In this example, during operation 1302d, the electronic device 100 provides a live preview image of a book, with its shadow covering a portion of the book cover. The user then selects the "artifact removal" option from a drop-down menu, and the electronic device 100 captures an image of the book by removing the shadows.
[0207] Figure 10eThis is an example illustrating an artifact management active element in the real-time preview mode of an electronic device 100 according to an embodiment disclosed herein.
[0208] Reference Figure 10e The electronic device 100 provides an artifact management active element, which may be, for example, an icon, a button, etc. The user needs to manually activate the artifact management active element to enable artifact management as provided by the proposed method. In operation 1302e, the user has not yet activated the artifact management active element, therefore the image preview includes shadows on the user's face. In operation 1304e, the user has enabled the artifact management active element, therefore the electronic device 100 identifies the shadows as unwanted artifacts and removes them while capturing the image.
[0209] Figure 10f This is an example of an example mode in which an electronic device 100 manages artifacts according to embodiments disclosed herein.
[0210] Reference Figure 10f Consider an input image including shadows on a user's face in regions 1 and 2, as shown in operation 1302f. Consider a user mode where electronic device 100 allows the user to select desired and unwanted shadow areas. In operation 1304f, electronic device 100 allows the user to retain the shadowed region 1 and remove the shadowed region 2. In operation 1306f, electronic device 100 allows the user to change the opacity of the shadow in region 1, thereby allowing the user to further enhance the image.
[0211] Consider the automatic mode, where the electronic device 100 automatically determines the desired and unwanted shadow areas and automatically applies a binary mask generated for them. In operation 1308f, the electronic device 100 determines that both shadow area 1 and shadow area 2 are unwanted, and therefore both are removed.
[0212] Figure 11a This is a flowchart illustrating a method for managing image artifacts by an electronic device 100 according to embodiments disclosed herein.
[0213] Reference Figure 11a The electronic device 100 can be an augmented reality (AR) device. In operation 1102, consider a user wearing the electronic device 100 and selecting "artifact mode" to be enabled. In operation 1104, the electronic device 100 captures an image frame, and in operation 1106, the electronic device 100 sends the captured image frame to a connected device. Furthermore, in operation 1108, the connected device identifies and removes image artifacts, and in operation 1110, the image frame is resent to the electronic device 100. In operation 1112, the electronic device 100 displays the processed frame. Therefore, the proposed method can be performed by a third device that is not responsible for capturing image frames.
[0214] Figure 11b This is an example illustrating the management of image artifacts by electronic device 100 according to embodiments disclosed herein.
[0215] Reference Figure 11b , combined Figure 11a In operation 1120, the user has not yet activated the artifact mode available in the electronic device 100, so the captured image frame, along with the shadows along the face of the captured object in the image frame, is displayed as is. In operation 1130, the user activates the artifact mode in the electronic device 100, and thus the shadows along the face of the object are automatically and intelligently removed, resulting in a clearer image. Therefore, the proposed method is executed in real time to obtain an image without artifacts.
[0216] The foregoing exemplary embodiments are merely illustrative and should not be construed as limiting. This teaching can be readily applied to other types of devices. Furthermore, the description of the exemplary embodiments is intended to be illustrative and not to limit the scope of the claims, and many alternatives, modifications, and variations will be apparent to those skilled in the art. While this disclosure has been shown and described with reference to various embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of this disclosure as defined by the appended claims and their equivalents.
Claims
1. A method for processing image data, the method comprising: Receive input image; Multiple features are extracted from the input image, including the texture of the input image, the color composition of the input image, and the edges in the input image; Based on the aforementioned features, at least one region of interest (RoI) is determined in the input image, including at least one artifact. At least one intermediate output image is generated by removing at least one artifact from the input image using multiple generative adversarial networks (GANs); A binary mask is generated using the at least one intermediate output image, the input image, the edges in the input image, and the edges in the at least one intermediate output image; Based on the binary mask of the input image, the at least one artifact is classified into a first artifact category or a second artifact category; and The final output image is obtained by processing the input image based on the category of at least one artifact corresponding to the first artifact category and / or the second artifact category.
2. The method according to claim 1, further comprising: Multiple versions of at least one final intermediate output image are generated by changing the gamma value associated with the final output image; Determine the Natural Image Quality Evaluator (NIQE) value for each of the plurality of versions of the at least one final intermediate output image; as well as Display one of the plurality of versions of the at least one final intermediate output image as the final output image, the version having the minimum NIQE value among the plurality of NIQE values for the at least one final intermediate output image.
3. The method according to claim 1, wherein, The extraction of the multiple features from the input image includes: Texture, color composition, and edges are extracted from the input image based on Otsu thresholding, Gaussian blur, and edge detection, respectively.
4. The method according to claim 1, wherein, The multiple GANs include a negative generator, an artifact generator, a partition generator, a negative discriminator, a partition discriminator, a refinement generator, and a refinement discriminator.
5. The method according to claim 4, wherein, Generating the at least one intermediate output image includes: Determine a set of loss values for each of the plurality of GANs; The negative generator is used to generate a first GAN image by removing the darkest region of at least one artifact in the input image; The artifact generator is used to generate a second GAN image by removing at least one of the color-continuous regions and texture-continuous regions of at least one artifact from the input image; The partitioning generator is used to generate a third GAN image by removing the brightest region of the at least one artifact and adding white patch regions to the at least one artifact in the input image; and At least one of the first GAN image, the second GAN image, and the third GAN image is used to generate the at least one intermediate output image without the at least one artifact.
6. The method according to claim 1, wherein, The input image is a previously input image, and the method further includes: Receive new input images; and The binary mask obtained for the previous input image is superimposed on the new input image.
7. The method according to claim 1, wherein, The at least one artifact is a shadow or glare.
8. The method according to claim 1, wherein, The first artifact category corresponds to the desired artifact, and the second artifact category corresponds to the unwanted artifact.
9. An electronic device for processing image data, the electronic device comprising: Memory, storing instructions; as well as The processor is configured to execute the instructions to perform the following operations: Receive input image; Multiple features are extracted from the input image, including the texture of the input image, the color composition of the input image, and the edges in the input image; Based on the aforementioned features, a region of interest (RoI) is determined in the input image, including at least one artifact. At least one intermediate output image is generated by removing at least one artifact from the input image using multiple generative adversarial networks (GANs); A binary mask is generated using the at least one intermediate output image, the input image, the edges in the input image, and the edges in the at least one intermediate output image; Based on the binary mask, classify the at least one artifact into a first artifact category or a second artifact category; and The final output image is obtained from the input image based on the category of at least one artifact corresponding to the first artifact category or the second artifact category.
10. The electronic device according to claim 9, wherein, The processor is also configured to: Multiple versions of at least one final intermediate output image are generated by changing the gamma value associated with the final output image; Determine the Natural Image Quality Evaluator (NIQE) value for each of the plurality of versions of the at least one final intermediate output image; as well as The version of the final output image that has the lowest NIQE value among the multiple versions.
11. The electronic device according to claim 9, wherein, The processor is also configured to: Texture, color composition, and edges are extracted based on Otsu thresholding, Gaussian blur, and edge detection, respectively.
12. The electronic device according to claim 9, wherein, The multiple GANs include a negative generator, an artifact generator, a partition generator, a negative discriminator, a partition discriminator, a refinement generator, and a refinement discriminator.
13. The electronic device according to claim 12, wherein, The processor is also configured to: Determine a set of loss values for each of the plurality of GANs; The negative generator is used to generate a first GAN image by removing the darkest region of at least one artifact in the input image; The artifact generator is used to generate a second GAN image by removing at least one of the color-continuous regions and texture-continuous regions of at least one artifact from the input image; The partitioning generator is used to generate a third GAN image by removing the brightest region of the at least one artifact and adding white patch regions to the at least one artifact in the input image; as well as At least one of the first GAN image, the second GAN image, and the third GAN image is used to generate the at least one intermediate output image without the at least one artifact.
14. The electronic device according to claim 9, wherein, The input image is a previously input image, and the processor is further configured to: Receive new input images; and The binary mask obtained for the previous input image is superimposed on the new input image.
15. The electronic device according to claim 9, wherein, The at least one artifact is a shadow or glare.
Citation Information
Patent Citations
Motion shadow removal method based on multi-characteristic fusion
CN107038690A
Shadow area detection method based on full convolutional network and mean shift
CN109754440A