Linear transformation model trained on non-pairwise data using diffusion model

By learning linear transformation models in the semantic latent space, using diffusion automatic encoder and mask transformation, the problem of removing opaque artifacts in the image is solved, and the glare and reflections are efficiently removed without changing other attributes of the image. It is suitable for client devices such as smart phone cameras and webcams.

CN120283255APending Publication Date: 2025-07-08GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380078137.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-11
Filing Date
2023-11-11
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively remove opaque artifacts such as glare and reflections in images, especially in the absence of paired input images, and prior methods are difficult to create pixel-aligned paired data.

Method used

The diffusion automatic encoder is used to learn a linear transformation model in the semantic latent space. Through the diffusion model, the image data is gradually converted and linear loss and mask transformation are applied, and the editing is restricted in the opaque artifact area and the accumulation properties of the opaque artifact are removed.

Benefits of technology

It realizes efficient removal of opaque artifacts without changing other properties of the image, and provides a more realistic output image, suitable for various client devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283255A_ABST
    Figure CN120283255A_ABST
Patent Text Reader

Abstract

A method may include receiving an image including identifying a tag including at least one opaque artifact, generating a transformed semantic potential space based on the image using a linear transformation model, generating a noisy image based on the image, generating a first estimated image based on the transformed semantic potential space using a diffusion model, and generating a second estimated image based on the noisy image using the diffusion model. A second estimated image is generated based on the transformed semantic potential space and the noisy image using a diffusion model, and the linear transformation model is trained based on the first estimated image, the second estimated image, and a loss that implements a linear variation in the linear transformation model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 383,416, filed on November 11, 2022, the disclosure of which is hereby incorporated by reference in its entirety. Technical Field

[0003] The present disclosure relates to image manipulation, and more particularly, to a robust method for authentically removing glasses glare from an input image. Background Art

[0004] Glare and reflections (opaque artifacts) on glasses are common in input images such as portrait photos, video conferencing streams, or other settings where the face of an object is captured in an image. Unfortunately, these artifacts (glare and reflections) are often inevitable when an image is captured in the presence of strong sunlight, bright lights, a nearby screen, etc. Opaque artifacts obscure the eyes of the object, thereby affecting the aesthetic quality of the portrait and interfering with the perception of the object's expression. Removing such artifacts from an image computationally has significant value as it enhances the quality of the image and broadens the environments in which good portrait photos and good object - centered videos can be taken. Summary of the Invention

[0005] In some aspects, the techniques described herein relate to a method for removing opaque artifacts (e.g., glare, reflections) from lenses in an image. Specifically, a glare removal model is trained that learns to remove reflections given only a binary class label - i.e., a set of images with and without reflections. In particular, a diffusion auto - encoder is used to learn the latent embedding of the input image and then edit the embedding to remove the opaque artifacts. Since opaque artifacts are additive in the image space, the implementation can include a novel linear loss that uses the additive nature of the opaque artifacts to find the editing direction. To further constrain the editing to remove the opaque artifacts without changing other attributes or while minimizing the change to other attributes, the implementation can include a mask transformation in the feature space of the denoising network to limit the editing to the eye region. The implementation can create pixel - aligned paired data that provides a more realistic resulting image compared to existing methods that rely on paired data.

[0006] In general aspects, an apparatus, a system, a non-transitory computer-readable medium (having computer-executable program code stored thereon that can be executed on a computer system), and / or a method can perform a process using the following method, which includes: receiving an image including at least one opaque artifact, and generating an enhanced image by minimizing the at least one opaque artifact using a trained linear transformation model. The trained linear transformation model is trained using: a first estimated image generated based on a semantic latent space using a diffusion model, a second estimated image generated based on the semantic latent space and a noisy image using a diffusion model, a loss implementing a linear transformation in the trained linear transformation model, and using the same training images to generate the semantic latent space and the noisy image.

[0007] In another general aspect, an apparatus, a system, a non-transitory computer-readable medium (having computer-executable program code stored thereon that can be executed on a computer system), and / or a method can perform a process using the following method, which includes: receiving an image including a label identifying at least one opaque artifact; generating a transformed semantic latent space based on the image using a linear transformation model; generating a noisy image based on the image; generating a first estimated image based on the transformed semantic latent space using a diffusion model; generating a second estimated image based on the transformed semantic latent space and the noisy image using a diffusion model; and training the linear transformation model based on the first estimated image, the second estimated image, and a loss implementing a linear transformation in the linear transformation model.

[0008] The foregoing exemplary summary and other exemplary objects and / or advantages of the present disclosure and their implementation are further explained in the following detailed implementation and its accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 A computing device including a glare removal model is shown in accordance with a possible implementation of the present disclosure.

[0010] Figure 2 A data flow diagram of a method for determining a transformation of a glare removal model is shown in accordance with a possible implementation of the present disclosure.

[0011] Figure 3 A difference between a global transformation and a regional transformation is shown in accordance with a possible implementation of the present disclosure.

[0012] Figure 4 A global transformation applied by a network to a semantic latent space is shown.

[0013] Figure 5Illustrates a regional transformation applied by a network to a semantic latent space according to a possible implementation of the present disclosure.

[0014] Figure 6 Is a block diagram of a method for generating an enhanced image according to an example implementation.

[0015] Figure 7 Illustrates a block diagram of a method for training a diffusion model according to at least one implementation of the present disclosure.

[0016] Figure 8A and Figure 8B Compare the output of the currently described techniques with glare removal techniques that rely on paired input.

[0017] The components in the figures are not necessarily drawn to scale relative to each other. Throughout several views, like reference numerals designate corresponding parts. Detailed Description

[0018] Implementations relate to a system and method for removing opaque artifacts from an input image. Specifically, implementations relate to training a machine learning model for removing opaque artifacts from an image that does not rely on paired input images. In other words, one image can be used in each training iteration (e.g., without using a ground truth image). For example, opaque artifacts may be associated with glare. Thus, some implementations relate to removing glare from an image. For example, some implementations may involve training a machine learning model for removing glare from glasses worn by an object in an image, where the training technique does not rely on paired input images. Opaque artifacts can include, for example, glare, shadows, and / or image discontinuities. In some implementations, the opaque artifact can be a human skin disorder such as, for example, rash, hives, vitiligo, eczema, etc. In some implementations, the opaque artifact can be an environmental discontinuity such as, for example, a tree missing some leaves, a discolored wall, a missing patch of grass, etc. Other opaque artifacts are also within the scope of the present disclosure. Some implementations not only fill in missing information but can also change parts of the image without deforming that part. For example, the described techniques can be used to change the color of the leaves of a tree without deforming the leaves. (e.g., changing the style from summer to autumn).

[0019] Previous machine learning methods for opaque artifact reduction / removal relied on paired images for pixel-by-pixel supervised learning. In such supervised learning, one input image represents the ground truth output image or the desired output image (without opaque artifacts), and the other image represents the same image with opaque artifacts. The model is then trained to produce the ground truth image given the input image. However, there are technical problems where the number of paired images that affect the quality of the model used in training is large, and the cost and difficulty of obtaining a sufficient number of real-world (e.g., non-synthetic) image pairs with and without opaque artifacts are high. Therefore, the ability to build a robust model based on such real-world data is limited.

[0020] To address the lack of manually curated image pairs for supervised training, other existing methods have generated synthetic image pairs, e.g., using physically based rendering, or taking images with and without a glass plane. However, these methods are not suitable for creating pixel-aligned paired data for opaque artifacts. For example, it is difficult to model glasses reflections due to the wide variety of lens geometries, tints, and coatings that can introduce effects that cause distortions such as those due to refraction, color shift, etc. Further, it is difficult to capture a pair of pixel-aligned images with and without glasses reflections because a human subject may move between the capture and the removal of the reflection source - e.g., a bright screen - which will change the lighting of the entire scene.

[0021] In contrast, the disclosed implementations include a technical solution having a model that learns to remove opaque artifacts (e.g., glare and reflections) from an image without paired input-output examples. Instead of such supervised learning using image pairs, the technical solution can include some implementations that learn a linear transformation in the semantic latent space used by generative methods in synthesis and restoration. In some implementations, the model (in inference mode) encodes the input image into the semantic and stochastic latent spaces (sometimes referred to as the latent space or semantic latent space), applies the learned linear transformation to the semantic latent space (the output is sometimes referred to as the latent space or the transformed semantic latent space), and decodes the image using the original stochastic latent. The resulting linear loss and latent mask semantic transformation help preserve the appearance of the regions of the image without opaque artifacts while only removing the opaque artifacts, resulting in a more realistic output image. Once trained, the model can be pushed to various client devices / included in various client devices for various purposes to remove opaque artifacts from images, photos, and / or videos. For example, the model can be pushed to a smartphone camera / included in a smartphone camera to remove opaque artifacts (e.g., glare, shadows, etc.) from photos, used in a webcam to remove opaque artifacts (e.g., glare, shadows, etc.) from a video conference feed, etc.

[0022] A diffusion autoencoder may include a diffusion model. The diffusion model may be configured to gradually transform data (e.g., image data) into noise and then train a neural network to learn to invert the noisy data back into the original data type. The increment may include reducing the noise of the noisy data by replacing some of the information in the information masked by the noise. In some implementations, starting from pure noise and incrementing through the diffusion model may generate new data.

[0023] In some implementations, a diffusion autoencoder (such as DiffAE) may be modified to utilize semantic and stochastic latent spaces. In some implementations, the diffusion model may be modified to utilize semantic and stochastic latent spaces. Given an unpaired set of images from two domains, the diffusion autoencoder may learn the latent directions and transform an image from one domain to another by editing the latent code in this direction. However, since the latent editing is global and the two domains often contain some unwanted biases, such editing often changes the image more than desired. Examples of such distortions are changing identity, head pose, and deforming 3D shapes. Due to the cumulative nature of reflections, the implementations include a novel linear loss to ensure that any semantic editing along the latent editing direction can only produce images with varying glare intensities. In other words, the output image may be an image that is a weighted blend of images with and without opaque artifacts.

[0024] This may lead to a constrained optimization that penalizes changes in the image space that are non-linear such as pose changes, 3D shape changes, etc. To spatially confine the editing to regions that include opaque artifacts, some implementations may include a feature transformation in the diffusion model. While some diffusion autoencoder methods may apply per-channel weighting to features in the diffusion model, the implementations extend this aspect to per-pixel transformations. This can ensure that an opaque artifact removal transformation is applied to regions that may contain opaque artifacts and thus avoid spurious changes in regions that do not contain opaque artifacts. Thus, the implementations may include a diffusion-based opaque artifact (e.g., reflection and glare) removal method that can learn from an unpaired set of images with and without opaque artifacts.

[0025] Some implementations may include a linear loss that confines the search in the latent space to directions that do not distort the image. In other words, the linear loss may minimize or eliminate changes to the input image other than opaque artifact removal. Thus, some implementations enable the diffusion autoencoder to apply locally confined semantic editing. The benefit of the described solution may be that some implementations are superior to methods that require paired training data and provide significant improvements when generalized to previously unseen input images (i.e., in the wild).

[0026] Figure 1 illustrates a computing device 100 that includes an artifact removal model 105, which is trained to remove opaque artifacts using the disclosed techniques. The artifact removal model 105 includes a semantic encoder 110, an artifact removal transform 115, and a semantic decoder 120. The semantic encoder 110 can be a diffusion autoencoder (sometimes referred to as DiffAE) configured to encode an input image 130 into a semantic and stochastic latent space. The input image 130 can be an image captured by a camera included in the computing device 100. The input image 130 can be an image captured by another computing device and transmitted to the computing device 100. The input image 130 can be an image (frame) of a video stream. As used herein, the latent space can be a feature vector represented by the symbol z sem (sometimes referred to as the latent space or semantic latent space). The artifact removal transform 115 can represent a locally selective transform applied to the latent space, as discussed herein. This transform can include a learned linear loss that minimizes the change to the input image, as discussed herein. Once the image has been modified in the latent space, the decoder 120 can be configured to convert the image from the latent space to an output image 140.

[0027] Similar to other generative models, such as generative adversarial networks and normalizing flows, a generative diffusion model such as the artifact removal model 105 can use a Gaussian latent space. Different from other methods, the artifact removal model 105 does not generate an image in one network pass from the Gaussian latent space, but traverses multiple latent spaces spanned by a Markov chain of the Gaussian latent space. Thus, the inference process can be an iterative denoising method starting from pure noise. During training, the Markov chain can be used to generate paired samples of images from a dataset x0 and one of its latent representations x t in the latent representation. The intermediate representation x t can be obtained by sampling t times from a Gaussian distribution:

[0028]

[0029] This process of adding noise follows a noise schedule defined by β t ; t ∈ 0,..., T - 1. The noise schedule can include steps where independent Gaussian noise can be added. Thus, it is equivalent to sampling x t directly from x0 with variance. This results in the following distribution:

[0030]

[0031] The inverse process can be such that the model ∈ θ can be trained to estimate the noise used to sample x0 is parameterized in the following way. Although the inference process can be stochastic, the deterministic technique for reversing the process can be expressed as follows:

[0032]

[0033] The training objective can be a simplified version of the variational lower bound on the log-likelihood of q(x t |x0) with respect to the noise ∈ 1:T added at time step t, resulting in:

[0034]

[0035] A deterministic technique for encoding samples into a Gaussian latent space can be derived using this form. However, the resulting manipulation of the latent may not lead to semantically meaningful changes in the image space. Therefore, a semantic latent z sem can be developed, which encodes an image into a one-dimensional (1D) feature vector that is used as the conditional input to the noise prediction model ∈ θ (x t ,t,z sem ) using the following parameterization:

[0036]

[0037] When substituting it into the inverse process, it becomes the following:

[0038]

[0039] Thus, the encoding process for the Gaussian latent can be expressed as follows:

[0040]

[0041] Classification loss: An autoencoder can be used to manipulate an image using a linear transformation in the latent space. This transformation can be implicitly learned by training a classifier on the semantic latent z sem . To obtain the class probabilities p, the following single fully connected layer is used as follows:

[0042] p(Z sem ) = ∑ i (Z sem,i w cls,i b cls,i )…………………………………………(7)

[0043] For binary labels y (e.g., artifact, glare, no artifact, no glare), the binary cross-entropy of the probability p is calculated as follows:

[0044] L cls= -(y log(p) + (1 - y) log(1 - p)) ……………………………(8)

[0045] The resulting transformation between one class and another in the latent space is given by z sem,l = T θ (z sem ) = w ⊙ s eem given.

[0046] Figure 2 shows a flowchart of a method for determining (training) an opaque artifact removal transformation, such as the artifact removal transformation 115. This transformation can represent the semantic latent space direction for opaque artifact removal. As Figure 1 shown, the flowchart includes a semantic encoder 210, a noise function 215, a linear transformation model 220, a diffusion model 225 (described with respect to Figure 2 and Figure 4 and Figure 5 ), a weighted average 230, a classification loss (BCE) 235, and a loss 240.

[0047] The linear transformation model 220 can be configured as an opaque artifact removal transformation. The linear transformation model 220 can be implemented as a linear transformation of the form T θ as the opaque artifact removal transformation. The linear transformation T θ can be optimized using a classification loss and a linear loss. The classification loss can be based on the labels given to the images 205 (x0) used in training. The training images 205 can be classified as including opaque artifacts (e.g., glare) or not including opaque artifacts (e.g., not including glare or in some cases no glare, slight glare, strong glare). The training images 205 are not paired. In other words, one image is used in each training iteration (e.g., no ground truth image is used); instead, each individual image 205 is labeled as including opaque artifacts or not including opaque artifacts. Since the images 205 are not paired, sufficient training images can be obtained with minimal difficulty. The BCE 235 can be used to optimize the linear transformation and can be determined based on these labels. The linear loss can be configured to penalize the difference between the weighted average 230 of the input images with and without opaque artifacts in the image space and the images reconstructed from the weighted average 230 of the original and transformed semantic latent (sometimes referred to as the latent space or the transformed semantic latent space). In the example of Figure 2 , the L1 loss 240 is calculated on and . The losses (classification loss and linear loss) can be combined as follows:[[]]

[0048] L = L cls + λlin L lin ……………………………………………………………(9)

[0049] The disclosed implementations can be configured to remove opaque artifacts in images 130, 205 while preserving other attributes. Other attributes can include, for example, the identity of a person or the background of an image. The implementations can achieve this by restricting the region of the image on which the transformation occurs. Since this is a change restricted to the region of the opaque artifact, the implementations incorporate this prior information. In the interpretation of how to restrict the region, the input image can be represented as x, the mask with values {0, 1} is represented as m, the pixels that should be affected by the transformation are represented as m ⊙ x, and the pixels that should not be affected by the transformation are represented as (1 - m) ⊙ x.

[0050] For t > 1, in each transformation from x t-1 to x t the global transformation algorithm uses the semantic latent z sem as follows:

[0051]

[0052] The semantic latent for transforming an image into a label for which a classifier is trained can be obtained by where w cls are the weights of the classifier trained to classify the attributes that should be manipulated ( sometimes referred to as the latent space or the transformed semantic latent space).

[0053] Figure 3 Shows a comparison between the global transformation and the regional transformation used in some of the disclosed implementations. Figure 3 Includes the original image 300 used as the input image (e.g., image 130), the output image 305 representing the global transformation, and the output image 310 representing the regional transformation. As Figure 3 shown, the global transformation that produces the output image 305 not only removes glare but also changes other attributes of the image, such as the smile, hair, head shape, etc. To construct a method that better restricts the transformation to the region of interest, the implementations locate the region of interest, such as Figure 3 region 320, and restrict the transformation to that region of interest so that other regions are not affected.

[0054] Figure 4 Shows the global transformation applied to the semantic latent z θ (sometimes referred to as the latent space or the semantic latent space) through the network f of the diffusion model 225. As sem (sometimes referred to as the latent space or the semantic latent space). As Figure 4As shown, the global transformation method uses a transformed semantic latent (sometimes referred to as the latent space or the transformed semantic latent space) to weight the channels of the network f θ . Figure 5 Regional transformations applied to the semantic latent are shown. In contrast to the global transformation, to better confine the transformation to the region of interest (i.e., the region of the input image that includes the opaque artifact, or the masked region), the implementation first extends the per-channel weighting in the full conditional latent vector to per-pixel weighting, and then uses the mask 510 to apply the linear transformation T θ only to the region of interest. Figure 5 Shows the extension of channel weighting to per-pixel weighting and uses the mask 510 to select whether to pick the original semantics for the pixel or the transformed Since it is applied at different scale levels of the UNet, resulting in different numbers of channels, 1×1 convolutions are used to adapt the number of channels of z sem to that of z sem,l = T θ (z sem ) = w ⊙ z sem , so that there is a corresponding number of levels for each channel.

[0055] The original transformation is as follows:

[0056]

[0057] To apply the locally selected transformation, some implementations can extend the per-channel weighting to per-pixel weighting. The implementation can do this by selecting the original (x,y) or the transformed T (z θ ) sem value according to the mask value m (c) for each pixel (x, y) of this channel.

[0058]

[0059] This results in the following regional selective semantic transformation:

[0060]

[0061] Since the transformation is only applied to the region m ⊙ x that may contain opaque artifacts, deformations outside this region related to glare removal do not damage the overall image restoration. Then, this property can be used to find the transformation direction that is linear in the image space. Some implementations can adapt the mask 510 for different scale levels by applying nearest neighbor downsampling.

[0062] The regional transformation method just described spatially constrains the transformation to the eye region, but the glare removal transformation T θ may still result in undesirable deformations (e.g., undesirable deformations in the glasses and the eye region), as Figure 3 shown. This is because other attributes of this region can be related to opaque artifacts. For example, a face with glare is more likely to look upward than a face without glare. To constrain T θ to only remove opaque artifacts, the implementation can utilize the fact that opaque artifact removal (or addition) is a linear operation in the image space. The linear assumption means that given a pair of images with and without opaque artifacts, one can obtain images with varying opaque artifact intensities by considering different linear mixtures of the two images. On the other hand, for a good direction of opaque artifact removal, moving along that direction should also gradually remove the opaque artifacts. Therefore, the implementation can use multiple intensities of a linear transformation that converts one class in the latent space to another class, and apply a loss criterion that implements a linear change in the image space.

[0063] Specifically, the implementation samples α ∈ [0; 1] and obtains an interpolation between glare and non - glare in the semantic latent space as follows:

[0064] z sem = Enc(x0)……………………………………………………(14a)

[0065]

[0066] Using these semantic latent vectors, the implementation can apply one iteration of the diffusion model to obtain the estimated final image using the following settings

[0067]

[0068] The resulting linear loss is the mean absolute difference between the interpolation in the image space and the rendering of the interpolation in the semantic latent space with respect to the weights of the classifier:

[0069]

[0070] Some implementations can be trained using a subset of a dataset that includes, for example, faces with glasses. The implementation can select images with the label "Human Face", and can preprocess and filter the images. Filtering includes rejecting low-quality images, extreme poses, and very bright or dark images. The implementation can apply a detector for glasses to all remaining images, and annotate the images with glasses according to the glare levels "No Glare", "Mild Glare", and "Strong Glare". In some implementations, tens of thousands of images can be used as training input images. In some implementations, for some of the images, synthetic glare can be added only to the lens area.

[0071] In an example implementation, the linear transformation can be trained on images of size 256×256 with a batch size of 21 on a single NVIDIA A100 GPU. The implementation can fix the learning rate to 10 -3 , and train on 500000 samples. For the linear loss, the transformation weight of T θ is 0.3, and the implementation can sample from a uniform distribution For inference, the example implementation can precompute semantic and stochastic conditions for efficiency. For hyperparameter tuning, the example implementation uses a small set of images excluded from the test set, and sets the number of diffusion steps T to 250. For the final evaluation, the implementation can use 1000 diffusion steps.

[0072] Example 1. Figure 6 is a block diagram of a method for generating enhanced images according to an example implementation. As Figure 7 shown, in step S605, an image including at least one opaque artifact is received.

[0073] In step S610, an enhanced image is generated by minimizing at least one opaque artifact using a trained linear transformation model. The trained linear transformation model is trained using: a first estimated image generated based on a semantic latent space using a diffusion model, a second estimated image generated based on a semantic latent space and a noisy image using a diffusion model, a loss that implements a linear transformation in the trained linear transformation model, and using the same training images to generate the semantic latent space and the noisy image.

[0074] Example 2. Figure 7 is a block diagram of a method for training a diffusion model according to an example implementation. As Figure 7As shown, in step S705, an image including a label identifying at least one opaque artifact is received. In step S710, a first latent space or a transformed semantic latent space is generated based on the image using a linear transformation model. In step S715, a noisy image is generated based on the image. In step S720, a first estimated image is generated based on the first latent space using a diffusion model. In step S725, a second estimated image is generated based on the first latent space and the noisy image using a diffusion model. In step S730, the linear transformation model is trained based on the first estimated image, the second estimated image, and a loss implementing the linear transformation in the linear transformation model.

[0075] Example 3. The method as described in any of the above examples may further include generating a second latent space or a semantic latent space by encoding the image using a semantic encoder, and generating the first latent space based on the second latent space using the linear transformation model.

[0076] Example 4. The method as described in any of the above examples may further include generating a third estimated image based on the second latent space and the noisy image using the diffusion model, generating a fourth estimated image based on the first latent space and the noisy image using the diffusion model, and generating the first estimated image as a weighted average of the third estimated image and the fourth estimated image.

[0077] Example 5. The method as described in any of the above examples may further include generating a weighted latent space as a weighted average of the first latent space and the second latent space, and generating the second estimated image based on the weighted semantic latent space and the noisy image using the diffusion model.

[0078] Example 6. The method as described in any of the above examples, wherein the semantic encoder, the linear transformation model, and the diffusion model form an autoencoder. The autoencoder may be configured to modify an image using a linear transformation in the latent space.

[0079] Example 7. The method as described in any of the above examples, wherein the linear transformation model may include a classifier with weights, and the training of the linear transformation model includes modifying the weights.

[0080] Example 8. The method as described in any of the above examples, wherein the linear transformation model may include a classifier with weights, the weights may be per-pixel weights, and the training of the linear transformation model may include modifying the per-pixel weights in a region of the first latent space including the at least one opaque artifact.

[0081] Example 9. The method according to any one of the above examples, wherein the region of the first latent space including the at least one opaque artifact can be identified using a mask.

[0082] Example 10. The method according to any one of the above examples, wherein the linear transformation model can include a classifier with weights, and the loss can be the mean absolute difference between the first estimated image and the second estimated image with respect to the weights.

[0083] Example 11. A method can include any combination of one or more of Examples 1 to 10.

[0084] Example 12. A non-transitory computer-readable storage medium, including instructions stored thereon, which when executed by at least one processor are configured to cause a computing system to execute the method according to any one of Examples 1 to 11.

[0085] Example 13. A device, including components for executing the method according to any one of Examples 1 to 11.

[0086] Example 14. A device, including at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, using the at least one processor, cause the device to at least execute the method according to any one of Examples 1 to 11.

[0087] Figure 8A and Figure 8B Compare the output of the currently described technology with glare removal techniques that rely on paired inputs. In Figure 8A and Figure 8B , column A represents the input image with glare, column B represents applying the RePaint model, column C represents applying the RePaint model with a threshold (applying inpainting only in the glasses area), column D represents applying the IBCLN model retrained for glare in the glasses (rather than reflections in general), column E represents DiffAE trained with "slight glare" vs. "no glare" and "strong glare" vs. "no glare", and column F represents applying the disclosed glare removal model of the present disclosure.

[0088] Example implementations may include a non - transitory computer - readable storage medium that includes instructions stored thereon, which when executed by at least one processor are configured to cause a computing system to perform any of the methods described above. Example implementations may include a device that includes means for performing any of the methods described above. Example implementations may include a device that includes at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the device to perform at least any of the methods described above.

[0089] In accordance with aspects of the present disclosure, implementations of the various techniques and methods described herein can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations thereof. Implementations can be implemented as a computer program product (e.g., a computer program tangibly embodied in an information carrier, a machine - readable storage device, a computer - readable medium, a tangible computer - readable medium) for execution or control of the operation of a data - processing apparatus (e.g., a programmable processor, a computer, or multiple computers). In some implementations, the tangible computer - readable storage medium can be configured to store instructions that, when executed, cause a processor to perform a certain process. A computer program, such as the computer programs described above, can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand - alone program or as a module, component, sub - routine, or other unit suitable for use in a computing environment. A computer program can be deployed to be processed on one computer or on multiple computers distributed at one site or across multiple sites and interconnected by a communication network.

[0090] Although certain features of the described implementations have been illustrated as described herein, many modifications, substitutions, changes, and equivalents will now occur to those skilled in the art. Accordingly, it is to be understood that the appended claims are intended to cover all such modifications and changes that fall within the scope of the implementations. It should be understood that they are presented by way of example only and not by way of limitation, and that various changes in form and detail can be made. Except for mutually exclusive combinations, any part of the devices and / or methods described herein can be combined in any combination. The implementations described herein can include various combinations and / or sub - combinations of the functions, components, and / or features of the different implementations described.

[0091] It should be understood that in the foregoing description, when an element is referred to as being on another element, connected to another element, electrically connected to another element, coupled to another element, or electrically coupled to another element, the element can be directly on the other element, connected or coupled to the other element, or there can be one or more intermediate elements. In contrast, when an element is referred to as being directly on another element, directly connected to another element, or directly coupled to another element, there are no intermediate elements. Although the terms "directly on", "directly connected to", or "directly coupled to" may not be used throughout the detailed description, elements shown as being directly on, directly connected, or directly coupled may be so referred to. The claims (if any) of the present application may be modified to recite the exemplary relationships described in the specification or shown in the figures.

[0092] As used in this specification, the singular forms may include the plural forms unless specifically stated otherwise in context. Spatially relative terms (e.g., above, on top of, upper, below, beneath, under, lower, etc.) are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. In some implementations, the relative terms above and below may include vertically above and vertically below, respectively. In some implementations, the term "adjacent" may include laterally adjacent or horizontally adjacent.

[0093] It should also be noted that the software implementation aspects of the example implementations are typically encoded on some form of non-transitory program storage medium or implemented on some type of transmission medium. The program storage medium can be magnetic (e.g., a floppy disk or a hard disk) or optical (e.g., a compact disc read-only memory or CD ROM), and can be read-only or random access. Similarly, the transmission medium can be a twisted pair, a coaxial cable, an optical fiber, or some other suitable transmission medium known in the art. The example implementations are not limited by these aspects of any given implementation.

[0094] It should also be noted that in some alternative implementations, the indicated functions / actions may not occur in the order indicated in the figures. For example, two figures shown consecutively may actually be executed simultaneously or sometimes in the reverse order, depending on the functionality / behavior involved.

[0095] Finally, it should also be noted that although the appended claims recite specific combinations of the features described herein, the scope of the present disclosure is not limited to the specific combinations claimed hereinafter, but extends to cover any combination of the features or implementations disclosed herein, regardless of whether that specific combination has been specifically recited in the appended claims at this time.

Claims

1. A method, comprising: Receiving an image including at least one opaque artifact; And Generating an enhanced image by minimizing the at least one opaque artifact using a linear transformation model, wherein the linear transformation model is trained using: A first estimated image generated by a diffusion model based on a first latent space, A second estimated image generated by the diffusion model based on a noisy image and the first latent space, Implementing a loss of a linear variation in the linear transformation model, Wherein the same training image is used to generate the first latent space and the noisy image, and the difference between the first estimated image and the second estimated image is compared with the loss.

2. The method according to claim 1, wherein Training the linear transformation model includes: Generating the first latent space by encoding the image using a semantic encoder; and Generating a second latent space using the linear transformation model based on the first latent space.

3. The method according to claim 2, wherein, Training the linear transformation model further includes: Generating a third estimated image using the diffusion model based on the first latent space and the noisy image; Generating a fourth estimated image using the diffusion model based on the second latent space and the noisy image; and Generating the first estimated image as a weighted average of the third estimated image and the fourth estimated image.

4. The method according to claim 2 or claim 3, wherein Training the linear transformation model further includes: Generating a weighted latent space as a weighted average of the first latent space and the second latent space; and Generating the second estimated image using the diffusion model based on the weighted latent space and the noisy image.

5. The method according to any one of claims 2 to 4, wherein The semantic encoder, the linear transformation model, and the diffusion model form an autoencoder.

6. The method according to any one of claims 1 to 5, wherein, The linear transformation model includes a classifier with weights, and The training of the linear transformation model includes modifying the weights.

7. The method according to any one of claims 1 to 5, wherein, The linear transformation model includes a classifier with weights, The weights are per-pixel weights, and The training of the linear transformation model includes modifying the per-pixel weights in a region of the second latent space including the at least one opaque artifact.

8. The method according to claim 7, wherein, The region of the second latent space including the at least one opaque artifact is identified using a mask.

9. The method according to any one of claims 1 to 8, wherein, The linear transformation model includes a classifier with weights, and The loss is the mean absolute difference between the first estimated image and the second estimated image with respect to the weights.

10. A method, comprising: Receiving an image including a label identifying at least one opaque artifact; Generating a first latent space using a linear transformation model based on the image; Generating a noisy image based on the image; Generating a first estimated image using a diffusion model based on the first latent space; Generating a second estimated image based on the first latent space and the noisy image using the diffusion model; and Training the linear transformation model based on the first estimated image, the second estimated image, and a loss implementing the linear transformation in the linear transformation model.

11. The method according to claim 10, further comprising: Generating a second latent space by encoding the image using a semantic encoder; and Generating the first latent space using the linear transformation model based on the second latent space.

12. The method according to claim 11, further comprising: Generating a third estimated image based on the second latent space and the noisy image using the diffusion model; Generating a fourth estimated image based on the first latent space and the noisy image using the diffusion model; and Generating the first estimated image as a weighted average of the third estimated image and the fourth estimated image.

13. The method according to claim 11, further comprising: Generating a weighted latent space as a weighted average of the second latent space and the first latent space; and Generating the second estimated image based on the weighted latent space and the noisy image using the diffusion model.

14. The method according to any one of claims 11 to 13, wherein The semantic encoder, the linear transformation model, and the diffusion model form an autoencoder.

15. The method according to any one of claims 11 to 14, wherein the linear transformation model includes a classifier with weights, and the training of the linear transformation model includes modifying the weights.

16. The method according to any one of claims 11 to 14, wherein the linear transformation model includes a classifier with weights, the weights are per-pixel weights, and the training of the linear transformation model includes modifying the per-pixel weights in the region of the first latent space including the at least one opaque artifact.

17. The method according to claim 16, wherein, The region of the first latent space including the at least one opaque artifact is identified using a mask.

18. The method according to any one of claims 11 to 17, wherein the linear transformation model includes a classifier with weights, and the loss is the mean absolute difference between the first estimated image and the second estimated image with respect to the weights.

19. A non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium including instructions stored thereon, the instructions being configured to cause a computing system to perform the method according to any one of claims 1 to 18 when executed by at least one processor.

20. A device, including means for performing the method according to any one of claims 1 to 18.

21. A device, including: at least one processor; and at least one memory, the at least one memory including computer program code; The at least one memory and the computer program code are configured to, with the at least one processor, cause the device to perform at least the method according to any one of claims 1 to 18.