Linear transformation model trained on unpaired data using a diffusion model
A diffusion autoencoder-based model learns to remove glare and reflections in images using semantic latent space transformations, addressing the limitations of paired data reliance and achieving superior artifact removal with realistic results.
Patent Information
- Application Number
- JP2025526745
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-11
- Filing Date
- 2023-11-11
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2043-11-11
AI Technical Summary
Conventional methods for removing opacity artifacts like glare and reflections from eyeglasses in images rely on paired input-output examples, which are costly, difficult to obtain, and unsuitable for creating pixel-aligned paired data, especially for diverse lens shapes and environmental conditions, leading to suboptimal model performance.
A diffusion autoencoder-based model that learns linear transformations in a semantic latent space to remove opacity artifacts without paired input-output examples, using a novel linear loss and masked transformations to constrain edits to the eye region, ensuring realistic and targeted removal of glare and reflections.
The model effectively removes opacity artifacts while preserving image integrity, outperforming previous methods by providing more realistic outputs and generalizing well to unseen images, even in real-world conditions.
Smart Images

Figure 2025535602000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 383,416, filed November 11, 2022, the entire disclosure of which is incorporated herein by reference.
[0002] The present disclosure relates to image manipulation, and more particularly to a robust method for realistically removing eyeglass glare from an input image. [Background technology]
[0003] Glare and reflections from eyeglasses (opacity artifacts) are common in input images, such as portrait photographs, video conference streams, or other scenes where a subject's face is captured in the image. Unfortunately, these artifacts (glare and reflections) are often unavoidable when capturing images in the presence of strong sunlight, bright lights, nearby screens, etc. Opacity artifacts obscure the subject's eyes, affecting the aesthetics of the portrait and hindering understanding of the subject's facial expression. Computationally removing such artifacts from images has significant value because it improves image quality and expands the situations in which good portrait photographs and good subject-centric videos can be captured. Summary of the Invention
[0004] In some aspects, the techniques described herein relate to methods for removing opacity artifacts (e.g., glare, reflections) from lenses in images. Specifically, these techniques train a deglare model that learns to remove reflections given only binary class labels, i.e., a set of images with and without reflections. Specifically, a diffusion autoencoder is used to learn latent embeddings of input images, and then the embeddings are edited to remove opacity artifacts. Because opacity artifacts are additive in image space, embodiments may include a novel linear loss that uses the additive nature of opacity artifacts to find edit directions. To further constrain edits to remove opacity artifacts without changing other attributes or while minimizing changes to other attributes, embodiments may include a masked transformation in the feature space of the denoising network to restrict edits to the eye region. Embodiments can create pixel-aligned paired data that provides more realistic resulting images than previous approaches that rely on paired data.
[0005] In a general aspect, a device, system, non-transitory computer-readable medium (storing computer-executable program code that can be executed on a computer system), and / or method can perform a process by receiving an image including at least one opacity artifact and generating an enhanced image by minimizing the at least one opacity artifact using a trained linear transformation model, where the trained linear transformation model is trained using a first estimated image generated based on a semantic latent space using a diffusion model, a second estimated image generated based on the semantic latent space and a noisy image using the diffusion model, and a loss that enforces linear change in the trained linear transformation model, where the semantic latent space and the noisy image are generated using the same training image.
[0006] In another general aspect, a device, system, non-transitory computer-readable medium (storing computer-executable program code that can be executed on a computer system), and / or method can perform a process by a method that includes receiving an image including a label that identifies the inclusion of at least one opacity artifact; generating a transformed semantic latent space based on the image using a linear transformation model; generating a noisy image based on the image; generating a first estimated image based on the transformed semantic latent space using a diffusion model; generating a second estimated image based on the transformed semantic latent space and the noisy image using the diffusion model; and training a linear transformation model based on the first estimated image, the second estimated image, and a loss that enforces linear change in the linear transformation model.
[0007] The foregoing illustrative summary, as well as other illustrative objects and / or advantages of the present disclosure and the manner in which they are achieved, are further described below in the detailed description and accompanying drawings. [Brief explanation of the drawings]
[0008] [Figure 1] 1 illustrates a computing device including a glare removal model according to a possible implementation of the present disclosure. [Figure 2] 1 illustrates a data flow diagram of a method for determining a transformation of a glare reduction model according to a possible implementation of the present disclosure. [Figure 3] 1 illustrates the difference between a global transformation and a local transformation according to a possible implementation of the present disclosure. [Figure 4] We show the global transformation applied by the network to the semantic latent space. [Figure 5] 1 illustrates a local transformation applied to a semantic latent space by a network according to a possible embodiment of the present disclosure. [Figure 6] FIG. 1 is a block diagram of a method for generating an enhanced image according to an exemplary embodiment. [Figure 7] FIG. 1 shows a block diagram of a method for training a diffusion model in accordance with at least one embodiment of the present disclosure. [Figure 8A] The output of the currently described technique is compared with the output of a glare reduction technique that relies on paired inputs. [Figure 8B] The output of the currently described technique is compared with the output of a glare reduction technique that relies on paired inputs. DETAILED DESCRIPTION OF THE INVENTION
[0009] The elements in the drawings are not necessarily drawn to scale relative to each other. Like reference numerals indicate corresponding parts throughout the several views.
[0010] Embodiments relate to systems and methods for removing opacity artifacts from input images. Specifically, embodiments relate to training a machine-learned model to remove opacity artifacts in an image independent of paired input images. In other words, a single image may be used in each training iteration (e.g., without using a ground truth image). For example, opacity artifacts may be associated with glare. Accordingly, some embodiments relate to removing glare from an image. For example, some embodiments may relate to training a machine-learned model to remove glare from eyeglasses worn by a subject in an image, where the training technique does not rely on paired input images. Opacity artifacts may include, for example, glare, shadows, and / or image discontinuities. In some embodiments, the opacity artifact may be, for example, a human skin condition such as a rash, hives, vitiligo, eczema, etc. In some embodiments, the opacity artifact may be, for example, an environmental discontinuity such as a tree losing some leaves, a discolored wall, missing some grass, etc. Other opacity artifacts are within the scope of this disclosure. Some implementations can not only fill in missing information, but also change parts of an image without distorting them. For example, the described techniques can be used to change the color of tree leaves (e.g., change style from summer to fall) without distorting the leaves.
[0011] Conventional machine-learned methods for reducing / removing opacity artifacts rely on paired images for pixel-wise supervised learning. In such supervised learning, one input image represents a ground truth image or desired output image (without opacity artifacts), and the other image represents the same image with opacity artifacts. A model is then trained to generate a ground truth image given the input image. However, technical challenges exist, such as the significant impact of the number of paired images used in training on the quality of the model, and the cost and difficulty of obtaining a sufficient number of pairs of real-world (e.g., non-synthetic) images with and without opacity artifacts. This limits the ability to create robust models from such real-world data.
[0012] To address the deficiencies of manually curated image pairs for supervised training, other conventional methods generate synthetic image pairs, for example, using physically based rendering or capturing images with and without glass surfaces. However, these methods are not suitable for creating pixel-aligned paired data for opacity artifacts. For example, modeling eyeglass reflections is difficult due to the wide variety of lens shapes, tints, and coatings that can introduce effects such as refractive distortions, color casts, etc. Furthermore, capturing pixel-aligned pairs of images with and without eyeglass reflections is challenging because human subjects may move between captures, and removing a reflection source such as a bright screen changes the overall scene illumination.
[0013] In contrast, disclosed embodiments include technical solutions having a model that learns to remove opacity artifacts (e.g., glare and reflections) from images without paired input-output examples. Instead of such supervised learning using image pairs, the technical solutions may include some implementations that learn linear transformations in a semantic latent space used by generative approaches in synthesis and restoration. In some implementations, the model (in inference mode) encodes input images into semantic and probabilistic latent spaces (sometimes referred to as latent spaces or semantic latent spaces), applies the learned linear transformations to the semantic latent space (the output is sometimes referred to as latent space or transformed semantic latent space), and decodes the image using the original probabilistic latent. The resulting linear loss and latent masked semantic transformation help preserve the appearance of regions of the image without opacity artifacts while removing only the opacity artifacts, resulting in a more realistic output image. Once trained, the model can be pushed to / included in various client devices for various purposes of removing opacity artifacts from images, photos, and / or videos. For example, models can be pushed / included in smartphone cameras to remove opacity artifacts (e.g., glare, shadows, etc.) from photos, used in webcams to remove opacity artifacts (e.g., glare, shadows, etc.) from video conferencing feeds, etc.
[0014] The diffusion autoencoder may include a diffusion model. The diffusion model may be configured to gradually transform data (e.g., image data) into noise and then train a neural network to learn to invert the noisy data back to its original data type. The incrementing may include reducing the noise in the noisy data by replacing some of the information masked by the noise. In some implementations, new data can be generated by starting with pure noise and incrementing through the diffusion model.
[0015] In some implementations, a diffusion autoencoder (such as DiffAE) may be modified using semantic and probabilistic latent spaces. In some implementations, a diffusion model may be modified using semantic and probabilistic latent spaces. Given a set of unpaired images from two domains, a diffusion autoencoder can transform images from one domain to another by learning latent directions and editing the latent code along those directions. However, because the latent edits are global and often contain some unintended biases in the two domains, such edits unnecessarily change the image. Examples of such distortions include identity, head pose changes, and 3D shape deformations. Due to the additive nature of reflections, implementations include a novel linear loss to ensure that any semantic edits along the latent edit directions can only result in images with varying glare intensities. In other words, the output image may be a weighted blend of an image with opacity artifacts and an image without opacity artifacts.
[0016] This can lead to constrained optimization that penalizes non-linear changes in image space, such as changes in pose or 3D shape. To spatially restrict edits to regions containing opacity artifacts, some implementations may include feature transformations in the diffusion model. While some diffusion autoencoder approaches can apply channel-wise weighting to features in the diffusion model, implementations extend channel-wise weighting to pixel-wise transformations. This ensures application of opacity artifact removal transformations to regions that may contain opacity artifacts and avoids spurious changes in regions that do not contain opacity artifacts. Accordingly, implementations may include a diffusion-based opacity artifact (e.g., reflection, glare) removal method that can learn from unpaired image sets of images with and without opacity artifacts.
[0017] Some implementations may include a linear loss, which constrains the search in the latent space in a direction that does not distort the image. In other words, the linear loss can minimize or eliminate changes to the input image other than opacity artifact removal. Thus, some implementations enable the diffusion autoencoder to apply locally constrained semantic edits. An advantage of the described solution may be that some implementations outperform methods that require paired training data and provide significant improvements when generalizing to previously unseen input images, i.e., in the wild.
[0018] 1 illustrates a computing device 100 including an artifact removal model 105 trained to remove opacity artifacts using the disclosed techniques. The artifact removal model 105 includes a semantic encoder 110, an artifact removal transform 115, and a semantic decoder 120. The semantic encoder 110 may be a diffusion autoencoder (sometimes referred to as DiffAE) configured to encode an input image 130 into a semantic and probabilistic latent space. The input image 130 may be an image captured by a camera included in the computing device 100. The input image 130 may be an image captured by another computing device and sent to the computing device 100. The input image 130 may be an image (frame) of a video stream. As used herein, a latent space refers to a space defined by z sem The image may be a feature vector (sometimes referred to as a latent space or semantic latent space) referred to by the notation . The artifact removal transform 115 may represent a locally selected transform applied to the latent space, as discussed herein. This transform may include a learned linear loss that minimizes changes to the input image, as discussed herein. Once the image is rectified in the latent space, the decoder 120 may be configured to convert the image from the latent space to the output image 140.
[0019] Similar to other generative models such as generative adversarial networks and regularized flows, generative diffusion models such as artifact removal model 105 may use a Gaussian latent space. Unlike other methods, artifact removal model 105 does not generate images in a single network pass from a Gaussian latent space, but traverses multiple latent spaces spanned by a Markov chain of Gaussian latent spaces. Thus, the inference process can be an iterative denoising method that starts from pure noise. During training, a Markov chain is used to generate a dataset x0 and its latent representation x t The intermediate representation x can be used to generate a pair of images. t can be obtained by sampling t times from a Gaussian distribution.
number
[0020] This noise-adding process is called β t ;t∈0,…,T-1. The noise schedule may include steps where independent Gaussian noise can be added. Thus, the variance is used to determine the noise schedule from x0 to x t This is equivalent to sampling directly, which gives us the following distribution:
number
[0021]
number
[0022] The training objective is to measure the noise ε added at time step t. t The log likelihood q(x 1:T This can be a simplified version of the variational lower bound for |x0), which gives
number
[0023]
number
[0024] Reversing this process gives us the following:
number
[0025] Thus, the encoding process of the Gaussian latent can be expressed as:
number
[0026] Classification loss: Using an autoencoder, we can manipulate images using a linear transformation in the latent space. This transformation is then transformed into a semantic latent z sem It can be learned implicitly by training a classifier with To obtain the class probabilities p, a single fully connected layer is used as follows: p(z sem )=Σ i (z sem,i w cls,i b cls,i ) (7)
[0027] For a binary label y (e.g., artifact, glare, no artifact, no glare), the binary cross entropy of probability p is calculated as follows: L cls =-(ylog(p)+(1-y)log(1-p)) (8)
[0028]
number
[0029] 2 shows a flow diagram of a method for determining (training) an opacity artifact(s) removal transform, such as the artifact removal transform 115 of FIG. 1. This transform can represent a semantic latent space direction for opacity artifact removal. As shown in FIG. 2, the flow diagram includes a semantic encoder 210, a noise function 215, a linear transformation model 220, a diffusion model 225 (described with respect to FIGS. 4 and 5), a weighted average 230, a classification loss (BCE) 235, and a loss 240.
[0030]
number
[0031]
number
[0032] For t>1, x t-1 From x t At each transition to , the global transformation algorithm involves the semantic latent z as follows: sem is used.
number
[0033]
number
[0034] Figure 3 illustrates a contrast between global and local transformations used in some disclosed embodiments. Figure 3 includes an original image 300, which serves as an input image (e.g., image 130), an output image 305 representing a global transformation, and an output image 310 representing a local transformation. As shown in Figure 3, the global transformation resulting in output image 305 not only removes glare, but also changes other attributes of the image, such as the smile, hair, and head shape. To develop a method for better confining the transformation to a region of interest, an embodiment locates a region of interest, e.g., region 320 in Figure 3, and confines the transformation to this region of interest, leaving other regions unaffected.
[0035]
number
[0036] The original transformation is:
number
[0037]
number
[0038] This results in the following locally selective semantic transformation:
number
[0039]
number
[0040]
number
[0041] Specifically, in an implementation, we sample α∈[0;1] to obtain an interpolation between glare and non-glare in the semantic latent space as follows: z sem =Enc(x0) (14a)
number
[0042]
number
[0043] The resulting linear loss is the mean absolute difference between rendering the interpolation in image space and the interpolation in semantic latent space, in terms of the classifier weights.
number
[0044] In some implementations, for example, training may be performed using a subset of the dataset that includes faces wearing glasses. In some implementations, images with a "human face" label may be selected and the images may be preprocessed and filtered. Filtering may include rejecting low-quality images, extreme poses, and very bright or very dark images. In some implementations, an eyeglasses detector may be applied to all remaining images, and images with eyeglasses may be annotated according to glare levels: "no glare," "light glare," and "strong glare." In some implementations, tens of thousands of images may be used as training input images. In some implementations, synthetic glare may be added only to the lens region in some images.
[0045]
number
[0046] Example 1. Figure 6 is a block diagram of a method for generating an enhanced image according to an exemplary embodiment. As shown in Figure 7, in step S605, an image containing at least one opacity artifact is received.
[0047] In step S610, an enhanced image is generated by minimizing at least one opacity artifact using a trained linear transformation model, where the trained linear transformation model is trained using a first estimated image generated based on a semantic latent space using a diffusion model, a second estimated image generated based on the semantic latent space and the noisy image using the diffusion model, and a loss that enforces linear change in the trained linear transformation model, where the semantic latent space and the noisy image are generated using the same training image.
[0048] Example 2. Figure 7 is a block diagram of a method for training a diffusion model according to an exemplary embodiment. As shown in Figure 7, in step S705, an image including a label identifying the inclusion of at least one opacity artifact is received. In step S710, a first latent space or a transformed semantic latent space is generated based on the image using a linear transformation model. In step S715, a noisy image is generated based on the image. In step S720, a first estimated image is generated based on the first latent space using a diffusion model. In step S725, a second estimated image is generated based on the first latent space and the noisy image using the diffusion model. In step S730, a linear transformation model is trained based on the first estimated image, the second estimated image, and a loss that enforces linear change in the linear transformation model.
[0049] Example 3. The method according to any of the above examples may further include generating a second latent space or a semantic latent space by encoding the image with a semantic encoder, and generating a first latent space based on the second latent space with a linear transformation model.
[0050] Example 4. The method of any of the above examples may further include generating a third estimated image based on the second latent space and the noisy image using a diffusion model, generating a fourth estimated image based on the first latent space and the noisy image using a diffusion model, and generating the first estimated image as a weighted average of the third estimated image and the fourth estimated image.
[0051] Example 5. The method of any of the above examples may further include generating a weighted latent space as a weighted average of the first latent space and the second latent space, and generating a second estimated image based on the weighted semantic latent space and the noisy image using a diffusion model.
[0052] Example 6. The method of any of the previous examples, wherein the semantic encoder, the linear transformation model, and the diffusion model form an autoencoder. The autoencoder may be configured to rectify the image using a linear transformation in latent space.
[0053] Example 7. The method of any of the preceding examples, wherein the linear transformation model may include a classifier with weights, and training the linear transformation model includes modifying the weights.
[0054] Example 8. The method of any of the preceding examples, wherein the linear transformation model may include a classifier with weights, the weights may be pixel-wise weights, and training the linear transformation model may include modifying the pixel-wise weights in a region of the first latent space that includes at least one opacity artifact.
[0055] Example 9. The method of any of the previous examples, wherein a region of the first latent space containing at least one opacity artifact may be identified using a mask.
[0056] Example 10. The method of any of the above examples, wherein the linear transformation model may include a classifier with weights, and the loss may be the mean absolute difference between the first estimated image and the second estimated image with respect to the weights.
[0057] Example 11. The method may include any combination of one or more of Examples 1-10.
[0058] Example 12. A non-transitory computer-readable storage medium including instructions, the instructions stored on the non-transitory computer-readable storage medium and configured, when executed by at least one processor, to cause a computing system to perform the method of any of Examples 1-11.
[0059] Example 13. An apparatus comprising means for carrying out the method according to any one of Examples 1 to 11.
[0060] Example 14. An apparatus including at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code configured, by the at least one processor, to cause the apparatus to at least perform the method of any of Examples 1-11.
[0061] 8A and 8B compare the output of the presently described technique with the output of a glare removal technique that relies on paired inputs. In Figures 8A and 8B, column A represents an input image with glare, column B represents the application of the RePaint model, column C represents the application of the RePaint model with a threshold (applying restoration only to the eyeglasses region), column D represents the application of the IBCLN model retrained on eyeglasses glare (rather than general reflections), column E represents DiffAE trained on "light glare" vs. "no glare" and "strong glare" vs. "no glare," and column F represents the application of the disclosed glare removal model of the present disclosure.
[0062] Exemplary embodiments may include a non-transitory computer-readable storage medium containing instructions, which when stored on the non-transitory computer-readable storage medium and executed by at least one processor, are configured to cause a computing system to perform any of the methods described above. Exemplary embodiments may include an apparatus including means for performing any of the methods described above. Exemplary embodiments may include an apparatus including at least one processor and at least one memory containing computer program code, the at least one memory and computer program code configured, by the at least one processor, to cause the apparatus to at least perform any of the methods described above.
[0063] According to aspects of the present disclosure, implementations of the various techniques and methods described herein may be implemented in digital electronic circuitry, or computer hardware, firmware, software, or combinations thereof. Implementations may also be implemented as a computer program product (e.g., a computer program tangibly embodied in an information carrier, a machine-readable storage device, a computer-readable medium, a tangible computer-readable medium) for processing by, or controlling the operation of, a data processing apparatus (e.g., a programmable processor, a computer, or multiple computers). In some embodiments, a tangible computer-readable storage medium may be configured to store instructions that, when executed, cause a processor to perform a process. Computer programs such as the computer program(s) described above may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, such as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may be deployed for processing by one computer or by multiple computers at one site or distributed across multiple sites and interconnected by a communications network.
[0064] While certain features of the described embodiments have been described herein, many modifications, substitutions, changes, and equivalents will occur to those skilled in the art. It is therefore to be understood that the appended claims are intended to cover all such modifications and variations that are within the scope of the embodiments. They are presented by way of example only, and not limitation, and it should be understood that various changes in form and detail may be made. Any portion of the apparatus and / or methods described herein may be combined in any combination, except in mutually exclusive combinations. The embodiments described herein may include various combinations and / or subcombinations of the functions, components, and / or features of the different embodiments described.
[0065] When the foregoing description refers to an element being on, connected to, electrically connected to, coupled to, or electrically coupled to another element, it will be understood that the element may be directly on, directly connected to, or directly coupled to another element, or that one or more intervening elements may be present. In contrast, when an element is referred to as being directly on, directly connected to, or directly coupled to another element, there are no intervening elements present. Throughout the detailed description, elements that are shown to be directly on, directly connected to, or directly coupled may be so referenced, even if the terms directly on, directly connected, or directly coupled are not used. The claims of this application may be amended to recite the exemplary relationships, if any, described in the specification or shown in the drawings.
[0066] As used herein, the singular can include the plural unless the context clearly indicates otherwise. Spatially relative terms (e.g., above, above, upper of, below, below, below, and below) are intended to encompass various orientations of the device in use or operation in addition to the orientation shown in the figures. In some embodiments, the relative terms "above" and "below" can include vertically above and vertically below, respectively. In some embodiments, the term "adjacent" can include laterally adjacent or horizontally adjacent.
[0067] It should also be noted that the software implementation aspects of the exemplary embodiments are typically encoded on some form of non-transitory program storage medium or implemented via some type of transmission medium. The program storage medium may be magnetic (e.g., a floppy disk or hard drive) or optical (e.g., compact disk read-only memory, or CD ROM), and may be read-only or random-access. Similarly, the transmission medium may be twisted wire pair, coaxial cable, fiber optics, or some other suitable transmission medium known in the art. The exemplary embodiments are not limited by these aspects of any particular implementation.
[0068] It should also be noted that in some alternative implementations, the functions / acts described may occur out of the order noted in the figures. For example, two figures shown in succession may, in fact, be executed concurrently, or may sometimes be executed in the reverse order, depending on the functions / acts involved.
[0069] Finally, it should also be noted that while the appended claims set forth particular combinations of features described herein, the scope of the present disclosure is not limited to the specific combinations claimed below, but instead extends to encompass any combination of features or embodiments disclosed herein, whether or not that particular combination is expressly recited in the appended claims at this time.
Claims
1. receiving an image including at least one opacity artifact; generating an enhanced image by minimizing the at least one opacity artifact using a linear transformation model, the linear transformation model comprising: a first estimated image generated based on the first latent space using a diffusion model; and a second estimated image generated based on the noisy image and the first latent space using the diffusion model; and a loss that enforces linear change in the linear transformation model; The method, wherein the first latent space and the noisy image are generated using the same training images, and the difference between the first estimated image and the second estimated image is compared to the loss.
2. Training the linear transformation model includes: generating the first latent space by encoding the image using a semantic encoder; and generating a second latent space based on the first latent space using the linear transformation model.
3. Training the linear transformation model further comprises: generating a third estimated image based on the first latent space and the noisy image using the diffusion model; and generating a fourth estimated image based on the second latent space and the noisy image using the diffusion model; and generating the first estimated image as a weighted average of the third estimated image and the fourth estimated image.
4. Training the linear transformation model further comprises: generating a weighted latent space as a weighted average of the first latent space and the second latent space; and generating the second estimated image based on the weighted latent space and the noisy image using the diffusion model.
5. The method according to any one of claims 2 to 4, wherein the semantic encoder, the linear transformation model and the diffusion model form an autoencoder.
6. the linear transformation model includes a classifier with weights; The method of any of claims 1 to 5, wherein the training of the linear transformation model comprises modifying the weights.
7. the linear transformation model includes a classifier with weights; the weights are pixel-wise weights, 6. The method of claim 1, wherein the training of the linear transformation model comprises modifying the pixel-wise weights in a region of the second latent space that includes the at least one opacity artifact.
8. The method of claim 7 , wherein the region of the second latent space containing the at least one opacity artifact is identified using a mask.
9. the linear transformation model includes a classifier with weights; The method according to any of claims 1 to 8, wherein the loss is the mean absolute difference between the first estimated image and the second estimated image with respect to the weights.
10. receiving an image including a label identifying the inclusion of at least one opacity artifact; generating a first latent space based on the image using a linear transformation model; generating a noisy image based on the image; generating a first estimated image based on the first latent space using a diffusion model; generating a second estimated image based on the first latent space and the noisy image using the diffusion model; training the linear transformation model based on the first estimated image, the second estimated image, and a loss that enforces linear change in the linear transformation model.
11. generating a second latent space by encoding the image using a semantic encoder; The method of claim 10 , further comprising: generating the first latent space based on the second latent space using the linear transformation model.
12. generating a third estimated image based on the second latent space and the noisy image using the diffusion model; and generating a fourth estimated image based on the first latent space and the noisy image using the diffusion model; and The method of claim 11 , further comprising: generating the first estimated image as a weighted average of the third estimated image and the fourth estimated image.
13. generating a weighted latent space as a weighted average of the second latent space and the first latent space; The method of claim 11 , further comprising: generating the second estimated image based on the weighted latent space and the noisy image using the diffusion model.
14. The method according to any of claims 11 to 13, wherein the semantic encoder, the linear transformation model and the diffusion model form an autoencoder.
15. the linear transformation model includes a classifier with weights; The method according to any one of claims 11 to 14, wherein the training of the linear transformation model comprises modifying the weights.
16. the linear transformation model includes a classifier with weights; the weights are pixel-wise weights, 15. The method of claim 11, wherein the training of the linear transformation model comprises modifying the pixel-wise weights in a region of the first latent space that includes the at least one opacity artifact.
17. The method of claim 16 , wherein the region of the first latent space containing the at least one opacity artifact is identified using a mask.
18. the linear transformation model includes a classifier with weights; The method according to any of claims 11 to 17, wherein the loss is the mean absolute difference between the first estimated image and the second estimated image with respect to the weights.
19. 19. A non-transitory computer readable storage medium comprising instructions, the instructions being stored on the non-transitory computer readable storage medium and configured to, when executed by at least one processor, cause a computing system to perform a method according to any of claims 1 to 18.
20. Apparatus comprising means for carrying out the method according to any one of claims 1 to 18.
21. at least one processor; at least one memory containing computer program code, Apparatus, wherein said at least one memory and said computer program code are configured, by said at least one processor, to cause said apparatus to at least perform a method according to any of claims 1 to 18.
Citation Information
Patent Citations
Image data restoration device, image data restoration method, and program
JP2019021258A
Image data restoration apparatus and image data restoration method
US20190026866A1