Method and apparatus for determining and using a controllable orientation in a GAN space

By employing gradient directions from an auxiliary network to control GAN latent codes, the method addresses data-intensive and entanglement issues in GAN models, achieving improved realism and control in image generation.

JP7850863B2Active Publication Date: 2026-04-23LOREAL SA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
LOREAL SA
Filing Date
2023-07-27
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing GAN models require large amounts of labeled training data and suffer from artifacts and attribute entanglements, leading to a decrease in user experience in tasks like human face editing.

Method used

Utilize gradient directions of an auxiliary network to control semantics in the GAN latent codes, enabling unraveled control over GAN output semantics using approximately 60 samples acquired through human supervision, and select important latent code channels with a Grad-CAM-based mask.

Benefits of technology

Achieves precise control over GAN output semantics with reduced data requirements, minimizing artifacts and attribute entanglements, resulting in more realistic and interpretable image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007850863000008
    Figure 0007850863000008
  • Figure 0007850863000009
    Figure 0007850863000009
  • Figure 0007850863000010
    Figure 0007850863000010
Patent Text Reader

Abstract

The methods, apparatus, and techniques herein relate to determining directions in a GAN latent space and obtaining untangled control over GAN output semantics, and for enabling use in the generation of synthetic images, such as for use in learning another model or creating augmented reality. The methods, apparatus, and techniques herein utilize, according to embodiments, the gradient directions of an auxiliary network to control semantics within a GAN latent code. It has been shown that about 60 samples can be used as the minimum amount of labeled data, which can be quickly obtained by a human teacher. Also, herein, according to embodiments, it has been shown that more untangled control is achieved by using a mask to select important latent code channels during operation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Cross-reference of related applications] This application claims priority to U.S. Provisional Application No. 63 / 392,905, filed on 28 July 2022, the entire contents of which are incorporated herein by reference. This application also claims priority to French Patent Application No. FR2210849, filed on 20 October 2022, the entire contents of which are incorporated herein by reference.

[0002] This disclosure relates to image processing using deep neural networks and artificial intelligence, including generative adversarial networks (GANs), for purposes such as creating augmented reality. More specifically, this disclosure relates to methods and apparatus for determining and using controllable orientations in a GAN space. [Background technology]

[0003] GANs have been successfully applied to various tasks in the beauty industry, such as simulating makeup effects on human faces and changing hair color, in order to provide realistic augmented reality (AR) technology. Nevertheless, conditional generative models that control output semantics typically require large amounts of labeled training data, which can be costly and time-consuming to acquire. Furthermore, artifacts and attribute entanglements can appear in the current face editing output, which can lead to a decrease in the overall user experience.

[0004] When editing human faces or generating large amounts of synthetic labeled data to train conditional models, it is desirable to control native GAN output semantics with the aim of improving the user experience through artifact removal and attribute disentangling. [Overview of the project]

[0005] The methods, apparatus, and techniques described herein relate to determining directions in the GAN latent space and obtaining unraveled control over the GAN output semantics, enabling their use in generating synthetic images, such as for training another model or creating augmented reality. According to embodiments, the methods, apparatus, and techniques described herein utilize the gradient directions of an auxiliary network to control the semantics in the GAN latent codes. It has been shown that approximately 60 samples can be used as a minimum amount of labeled data, which can be rapidly acquired through human supervision. Furthermore, according to embodiments, it has been shown that selecting important latent code channels using a mask during the operation results in more unraveled control. Various embodiments are shown and described herein, as described in the following descriptions according to embodiments. These and other embodiments will be apparent to those skilled in the art who consider this application as a whole.

[0006] Statement 1: A computing device comprising a processor and a memory device. The memory device stores instructions that, when executed by the processor, cause the computer device to: The instructions cause the computer device to provide a generator and an auxiliary network that share a latent space. The generator is configured to produce a composite image exhibiting human-interpretable semantic attributes, and the auxiliary network comprises a plurality of semantic attribute classifiers, each of which is a semantic attribute classifier controlled to produce a composite image from source images, each semantic attribute classifier is configured to classify the presence of one of the semantic attributes in the image to provide a meaningful direction for the generator to control one of the semantic attributes in the composite image. The instructions cause the computer device to produce the composite image from source images by applying respective semantic attribute controls to control each of the semantic attributes in the composite image. Each of the aforementioned semantic attribute manipulation units corresponds to the significant direction provided by the classifier associated with the semantic attribute.

[0007] Statement 2: The computing device of Statement 1, wherein the significant direction includes the gradient direction, and the instruction causes the computing device to calculate each semantic attribute operation unit from each gradient direction of the classifier associated with the semantic attribute.

[0008] Statement 3: The computing device of Statement 2, which causes the computing device to combine each gradient direction in each classifier associated with the two or more semantic attributes in order to control two or more semantic attributes.

[0009] Statement 4: The computing device in statement 1 or 2, which causes the computing device to unravel each of the semantic attribute manipulation units for application to the generator.

[0010] Statement 5: The computing device of Statement 4, wherein each semantic attribute manipulation unit includes a gradient direction vector calculated from the parameters of the respective semantic attribute classifier associated with each semantic attribute manipulation unit, and unraveling one of the respective semantic attribute manipulation units from another of the respective semantic attribute manipulation units includes removing a significant data dimension of the gradient direction vector associated with one of the respective semantic attribute manipulation units from the gradient direction vector associated with the other of the respective semantic attribute manipulation units.

[0011] Statement 6: The computing device of Statement 5, wherein the important data dimensions to be removed correspond to a threshold, the threshold identifying those data dimensions having an absolute value greater than or equal to the threshold.

[0012] Statement 7: The computing device of Statement 6, wherein if the i-th data dimension for the gradient direction vector associated with one of the respective semantic attribute manipulation units is identified by the threshold, the corresponding i-th data dimension for the gradient direction vector associated with the other of the respective semantic attribute manipulation units is set to zero.

[0013] Statement 8: The instruction causes the computing device to receive an input that identifies at least one of the semantic attributes controlled for the original image, one of the computing devices in any of statements 1 through 7.

[0014] Statement 9: The computing device of Statement 8, wherein the input identifies at least one of the respective semantic attributes, and includes a granular input for determining the amount of the semantic attribute applied when generating the composite image.

[0015] Statement 10: The instruction causes the computing device to perform one or more of the following actions: generate the composite image as a service; provide an e-commerce interface for purchasing a product or service; provide a recommendation interface for recommending a product or service; and provide an augmented reality interface using the composite image to provide an augmented reality experience, one of the computing devices in any of statements 1 through 9.

[0016] Statement 11: The computing device according to any one of Statements 1 to 10, wherein a particular semantic attribute among the plurality of semantic attributes includes one of the following: facial features including age, gender, smile, etc.; posture effects; makeup effects; hair effects; nail effects; cosmetic surgery or dental effects including one of rhinoplasty, facelift, blepharoplasty, implants, otoplasty, teeth whitening, orthodontics, etc.; and orthotic device effects including one of eye orthotics, mouth orthotics, or ear orthotics, etc.

[0017] Statement 12: The computing device, one of statements 1 to 11, wherein the generator is a GAN-based generator, each classifier is a neural network-based binary classifier, and the generator and each neural network-based binary classifier are co-trained using training images that show label semantic attributes to define the significant direction in each classifier.

[0018] Statement 13: A method comprising training a generative adversarial network-based (GAN-based) generator g to generate a composite image from source images in which at least one semantic attribute from a set of semantic attribute definitions is selectively controlled, wherein the generator includes a model that maps latent codes (z) in a latent space (Z) to images (x=g(z)) in an image space (X) in which human-interpretable semantics exist, the training comprising co-training the generator g and an auxiliary network including respective classifiers for each semantic attribute of the set of definitions, each classifier providing significant data directions used to control each of the semantic attributes when generating an updated image, and the training further comprising providing the generator g and the auxiliary network to generate a composite image.

[0019] Statement 14: The method of Statement 13 comprises calling the generator g using at least one control calculated from the parameters of the auxiliary network to generate a plurality of composite images having semantic attributes selected from the definition set, the auxiliary network providing the significant data direction for each of the at least one control.

[0020] Statement 15: The method of Statement 13 or 14, comprising training a further network model using at least some of the plurality of composite images.

[0021] Statement 16: A method comprising generating a composite image (g(z')) from an original image using a generator (g) having a latent code (z), wherein the generator manipulates a target semantic attribute (k) in the composite image, the generation being a significant direction identified from an auxiliary network that classifies each semantic attribute including the target semantic attribute k in z, discovering a significant direction in z with respect to the target semantic attribute k, defining z' by optimizing z with respect to the target semantic attribute k according to the significant direction, and outputting the composite image.

[0022] Statement 17: The method of Statement 16, wherein the auxiliary network includes a set of binary classifiers, one for each semantic attribute that the generator can manipulate, and the auxiliary network is co-learned to share a latent code space Z with the generator g.

[0023] Statement 18: The method of Statement 16 or 17, wherein each of the semantically significant data directions comprises the respective dimensional data vectors obtained from each of the individual classifiers, each vector comprising the direction and rate of the fastest increase in the individual classifier.

[0024] Statement 19: The method, as described in any of Statements 16 to 18, comprises unraveling the data direction by filtering out important dimensions of the semantically significant data direction for each other semantically significant data direction in of Statements 16 to 18, wherein a particular dimension is largely influenced by the magnitude of its absolute gradient.

[0025] Statement 20: The method of Statement 19, wherein the filtering includes evaluating each dimension of each semantically significant data direction unraveled from the target semantically significant data direction, and setting the value of the corresponding dimension in the target semantically significant data direction to zero if a particular dimension exceeds a threshold.

[0026] Statement 21: The latent code z is z' = z + αn z k' It is optimized as follows, where n z k n is a vector representing the direction of the semantically significant data before filtering, and n z k' The method of statement 20, wherein is a vector representing the direction of the target semantically significant data after filtering, and α is a hyperparameter that controls the interpolation direction and step size.

[0027] Statement 22: The method of any of statements 16 to 21, wherein the method includes repeating the discover, define, and optimize operations with respect to the latent code z' in order to further manipulate the semantic attribute k.

[0028] Statement 23: The method according to any one of Statements 16 to 22, comprising: generating a composite image for manipulating a plurality of target semantic attributes; discovering a significant direction in z with respect to each of the target semantic attributes, as identified by the auxiliary network for classifying each of the target semantic attributes in z; and optimizing z according to each of the significant directions.

[0029] Statement 24: The method of any of statements 16 to 23, wherein the method comprises training another network model using the composite image.

[0030] Statement 25: The method according to any one of statements 16 to 24, comprising receiving an input that identifies the target semantic attribute controlled for the original image.

[0031] Statement 26: The method of Statement 25, wherein the input includes fine-grained input for determining the amount of the target semantic attribute applied when generating the composite image.

[0032] Statement 27: The method according to any one of Statements 16 to 26, wherein the method includes providing the generator and the auxiliary network for generating the composite image as a service; providing an e-commerce interface for purchasing a product or service; providing a recommendation interface for recommending a product or service; and providing an augmented reality interface using the composite image for providing an augmented reality experience.

[0033] Statement 28: The method described in any one of Statements 16 to 27, wherein the target semantic attribute includes one of the following: facial features including age, gender, smile, etc.; posing effects; makeup effects; hair effects; nail effects; cosmetic surgery or dental effects including one of rhinoplasty, facelift, eyelid surgery, implants, ear reconstruction, teeth whitening, orthodontics, etc.; and the effects of appliances including one of eye appliances, mouth appliances, ear appliances, etc.

[0034] Statement 29: A method comprising the step of providing an augmented reality (AR) interface to provide an AR experience. The AR interface is configured to generate a composite image from a received image using a generator by applying a respective semantic attribute manipulation unit to control each semantic attribute in the composite image, each of the semantic attribute manipulation units corresponding to a meaningful direction provided by a classifier associated with the semantic attribute. The method comprises the step of receiving the received image and providing the composite image for the AR experience.

[0035] Statement 30: The method of Statement 29, comprising processing the composite image using an effects pipeline to simulate an effect, and providing the composite image having the simulated effect for presentation on the AR interface. The method of Statement 29 or 30 may be combined, for example, with any of the themes (to the extent applicable) of Statements 1 to 13 or 16 to 28.

[0036] Statement 31: A computing device comprising at least one processor and at least one non-transitory memory device storing computer-readable instructions for execution by the at least one processor, wherein the instructions cause the computing device to perform any one of the methods of statements 13 to 30.

[0037] Statement 32: A computer program product comprising at least one non-temporary storage device for storing computer-readable instructions for execution by at least one processor of a computing device, wherein the execution of the instructions causes the computing device to perform any one of the methods of Statements 13 to 30. [Brief explanation of the drawing]

[0038] [Figure 1] Figures 1A, 1B, and 1C provide an overview of the learning-related and / or architecture-related aspects disclosed herein in each embodiment. Figure 1A is an image of a computing device providing a convolutional neural network (CNN) in one embodiment. Figure 1B is an image of the gradient dimensions of the semantic attributes k and m in one embodiment. Figure 1C is a flowchart of the operation of one-step optimization in one embodiment. [Figure 2] Figure 2 is a table of images showing the deciphering of semantic attributes in a facial image according to one embodiment, specifically the deciphering of a smile from glasses. [Figure 3] Figure 3 shows a pseudocode list operation in one embodiment. [Figure 4] Figure 4 is a table of images showing the results of semantic attribute manipulation using a controllable neural network in one embodiment. [Figure 5] Figure 5 is a table of images showing the results of semantic attribute release using a controllable neural network in one embodiment. [Figure 6] Figure 6 is a table of images visualizing the directions found for comparing three controllable neural networks: two conventional controllable neural networks and the controllable neural network in one embodiment of this specification. [Figure 7] Figure 7 is a table of images showing a comparison of attribute decryption results obtained by a conventional controllable neural network and a controllable neural network in one embodiment of this specification. [Figure 8] Figure 8 is a block diagram of a graphical user interface (GUI) for generating a composite image from an original image using a controllable GAN in one embodiment. [Figure 9] Figure 9 is a block diagram of a computer system in one embodiment. [Figure 10]Figure 10 is a flowchart illustrating the operation in each embodiment of this specification. [Figure 11] Figure 11 is a flowchart illustrating the operation in each embodiment of this specification.

[0039]

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046]

[0047]

[0048] This concept is best illustrated through its specific embodiments described herein with reference to the accompanying drawings, where similar reference numerals throughout refer to similar features. It should be understood that, as used herein, the term "inventive invention" is intended to mean the conceptual invention that forms the basis of the embodiments described below, and not merely the embodiments themselves. Furthermore, it should be understood that the general inventive concept is not limited to the exemplary embodiments described below, and the following description should be read in that regard. [Modes for carrying out the invention]

[0049] While GAN models can generate highly realistic images from their latent spaces given random inputs, the generation process is typically a black box, making direct control of the output semantics impossible. Nevertheless, previous studies [VB20, VB20, SZ21, HHLP20, SYTZ20] have shown that significant directions and channels exist within the GAN latent space, and that linearly interpolating in these directions or varying individual channel values ​​results in interpretable transformations, such as adding glasses or a smile to a human face. The objective of the methods, apparatus, and techniques herein is to determine (e.g., learn) directions in the GAN latent space and to obtain unraveled control over the GAN output semantics, enabling, for example, the use of such learned directions to create augmented reality or to generate synthetic images to train a model. This specification focuses on human face editing tasks useful in the beauty industry and other industries. For this purpose, the methods, apparatus, and techniques herein, according to embodiments, utilize the gradient directions of an auxiliary network to control the semantics in the GAN latent code. Approximately 60 samples can be used as the minimum amount of labeled data, and it is shown that this data can be rapidly acquired through human supervision. Furthermore, according to the embodiments, it is shown herein that important latent coding channels can be selected using a Grad-CAM-based mask during operation, resulting in more unraveled control. The following sections review prior research in the relevant field to contextualize the teachings herein.

[0050] [GPAM +The Generative Adversarial Network (GAN) introduced in

[14] currently dominates the field of generative modeling due to its powerful ability to synthesize photorealistic images. Generally, a GAN consists of two networks: a generator that learns the mapping from its latent space to the image space, and a discriminator that distinguishes GAN-synthesized images from real images, and these two networks are jointly trained in an adversarial manner. To improve the fidelity of the output and the stability of the learning, various variants of GANs have been proposed [ACB17, BDS18, CCK + 18, KLA18, KLA + 20], and there is a growing interest in studying the latent spaces of these native GAN models that can control the generation process more finely.

[0051] The potential semantic operations of GANs have been studied by multiple prior works [VB20, VB20, SZ21, HHLP20, SYTZ20, PBH20], and it has been shown that there are significant directions and channels in the GAN latent space / GAN feature space. One of a series of studies has discovered the control of the output semantics of unconditional GANs with explicit teachers [SYTZ20, WLS21]. For example, InterFaceGAN [SYZT20] assumes that for any binary semantics (e.g., male vs. female), there exists a hyperplane in the latent space that functions as a separating boundary, and its normal vector represents a significant direction. A classifier pre-trained on the CelebA dataset [LLWT15] is used to generate pseudo-labels for GAN-synthesized images, and then the separating boundary is learned using state vector machines (SVMs) learned with paired data of GAN latent codes and corresponding semantic labels. [WLS21] discovers channels that control the positioning effect in the GAN activation space with the guidance from a pre-trained semantic discrimination network. However, these teacher-based methods require a large amount of labeled data and pre-trained neural networks, and the learned controls sometimes become intertwined, so the range of controls found by these teacher-based methods is usually limited.

[0052] Another series of studies has discovered such controls in self-supervised or unsupervised ways [HHLP20, CVB21, VB20]. For example, GANSpace [HHLP20] identifies important latent directions by applying principal component analysis (PCA) to vectors within the GAN latent space or feature space, and [VB20] discovers interpretable directions within the GAN latent space by jointly optimizing the directions and a reconstructor that reconstructs these directions and manipulation intensities from images generated based on the manipulated GAN latent codes. These methods do not require labeled data, but they typically require a large-scale manual examination of various manipulation directions and the identification of significant controls. More importantly, unlike techniques taught according to embodiments herein, control over the target semantics is not necessarily guaranteed.

[0053] Gradient-based knowledge improves the interpretability of neural networks [ZKL + 16. SCD + 17] or learning stability [WDB +

[19] is being used to improve [SCD]. +

[17] uses gradient information from the classification output to obtain a localization map, which is visual evidence within the image. [WDB + 19] proposes to improve the stability of the GAN's learning behavior by adding an additional step before the generator and classifier jointly optimize. In this additional step, the latent code is first optimized toward the region that the classifier considers more real, and this direction is obtained by computing the gradient toward the latent code.

[0054] [WDB +In contrast to

[19] , according to embodiments of this specification, control over the GAN output semantics is obtained by utilizing an auxiliary binary classifier gradient, which has not been seen in the literature to date. Accordingly, according to these embodiments, methods, apparatus and techniques for obtaining unraveled control in the GAN latent space are described herein. This method, apparatus and techniques first find semantically significant directions by computing the gradient of a binary classifier that scores various semantics given a latent code input, and then select important dimensions in the latent code for attribute unraveling during the operation. [Determination of semantic significance]

[0055] Figure 1A is an image representation of a computing device 100 in one embodiment, which includes a CNN architecture 102 and an image output (e.g., image 104 including 104A and 104B) stored in the device 100's memory device 106. Figure 1B is an image representation of a portion of the memory device 106 storing the dimensions of the gradients 132,130 for semantic attributes k and m, respectively, in one embodiment. This image shows how the dimension of the gradient for attribute m is unraveled from the corresponding dimension in the gradient for attribute k. Figure 1C is a flowchart of a one-step optimization operation 150 in one embodiment, which is further described below herein.

[0056] The CNN architecture 102 includes a controllable GAN g 108 configured to generate output images g(z)104A and g(z')104B from their respective latent codes z and z' (e.g., parts of the latent space Z of GAN g 108). Images g(z)104A and g(z')104B represent face images with various semantic attributes. In the illustrated example, the face of the individual in image g(z)104A is not smiling, while the face of the same individual in image g(z')104B is smiling.

[0057] The CNN architecture 102 further includes multiple auxiliary networks 110, one for each semantic attribute controlled by the GAN g 108. Each of the multiple auxiliary networks 110 includes its own mapping function f, and these auxiliary networks 110 are trained together with the GAN g 108 to find significant directions for each controlled semantic attribute. In one embodiment, the mapping function f includes a binary classifier. Figure 1A shows, for simplicity, one of the classifiers from each auxiliary network 110, i.e., the mapping function f for semantic attribute k. k In 110A, only the details within the relevant box 112 are shown. In Figure 1A, each of the GAN g 108 and the auxiliary network 110 is pre-trained with training images (e.g., 60 images for each semantic attribute). As shown in box 112, a mapping function f maps the GAN latent code (e.g., z) to a semantic score with respect to k in order to manipulate the semantic attribute k (e.g., smile) with respect to the GAN latent code z. k From the gradient direction n z k 114 is used as input to derive the modified z (e.g., z') of GAN g 108. The formula for deriving the modified z is explained further below.

[0058] Next, refer to Figure 1B. According to one embodiment, the semantic attribute is unraveled. The important dimension for manipulating the semantic attribute m is removed with respect to attribute k. According to one embodiment, the important dimension is n with respect to the semantic attribute m. z m Dimensions with a large gradient at 130. Such dimensions have a gradient direction n with respect to the semantic attribute k. z k At 112, it is masked to a value of 0 (e.g., set or defined), and the gradient direction n z k' 132 is the gradient direction n after masking (for example, after setting important attributes to zero). z kThis shows 112. Dimensional masks, especially those with large gradients, are explained further below.

[0059] Next, referring to Figure 1C, operation 150 shows the pipeline for one-step optimization. The operation begins in 152 with latent code z, such as a face image. In 154, significant directions in z are found with respect to semantic attribute k (e.g., the gradient is obtained using the learned mapping function f). k (Calculated from the parameters). The decision at 156 evaluates whether there is an entanglement with an arbitrary semantic m (where m represents a different semantic attribute from k controlled by the GAN g). If so, branching to 158 via "Yes" reveals significant directions in z with respect to the semantic attribute m (e.g., each learned mapping function f m (Calculated from the parameters). In 160, as illustrated with reference to Figure 1B, and similarly as further described below herein, the operation performs untangling. In one embodiment, the determination of entanglement, as in 156, was identified experimentally. Entanglement was observed by determining which attributes were entangled after editing each attribute using a small meta-validation set of 30 images following the original gradient direction of k. The relationships between the target attribute and any entangled attributes were tabled. Important attributes from each of the identified attributes (i.e., attribute m ≠ k) are excluded by masking. In one embodiment, for untangling, the top channels of the attributes (e.g., the 100 with the largest gradient magnitudes) are joined. And these dimensions are (e.g., n z k' (To generate it) it becomes 0 in the target direction.

[0060] In one embodiment, for example, without prior observational experiments, the gradient n is directly obtained. z m The gradient n depends on the magnitude of the i-th dimension value in the context. z kThe value of the i-th dimension in can be excluded by masking. In the first case, when there is entanglement, excluding those dimensions by masking helps in untangling. In the second case, there is no entanglement, and important dimensions in m are not important in k, so excluding them by masking does not significantly change the direction.

[0061] In step 162, the operation optimizes the latent code z. The operation can be repeated with subsequent instances of the latent code z (e.g., looping back to step 154) until the desired operation result is achieved. Referring again to step 156, in one embodiment, if there are no entanglements, the operation proceeds to step 162 via the "no" branch to optimize the latent code z. As described above, in one embodiment, the operation may mask out important dimensions of any other attributes, even if there are no entanglements. For example, it may be simpler (e.g., programmatically) to simply unravel important dimensions from all m≠k rather than unraveling important dimensions from the gradient of k after determining which attributes of m are entangled.

[0062] Figure 2 is Table 200 of images, including rows 202 and 204, to illustrate the unraveling of smiles from glasses. In both rows 202 and 204, the first (left) column 206 is the original image. The second (center) column 208 is edited with a smaller interpolation distance, and the third (right) column 210 is edited with a larger interpolation distance. The difference is that in the second row, when the distance is larger (i.e., more smiles are added compared to column 208), no glasses are added. In other words, unraveling is applied to the second row. Thus, row 202 shows each face image of the same subject in each image. From left to right, the subject is i) without smile and glasses, ii) with smile but without glasses, iii) with smile and glasses. In row 204, the image shows the same subject as shown in row 202. The image in row 204 shows the application of unraveling, which is not applied in row 202. From left to right, the subjects are: i) smiling and without glasses, ii) smiling but without glasses, and iii) smiling but without glasses.

[0063] During the one-step optimization operation, as explained with reference to Figures 1B and 1C, f 眼鏡 The important dimension in the gradient from is n z 笑顔 It is masked to 0. |n z k | i The term indicates the absolute value of the gradient vector in dimension i. Unlike the original multiclass classification setting of Grad-CAM, where ReLU is selected as the activation, here the absolute value is used, as in the binary classification setting. Dimensions that negatively impact the semantic score also contain significant information.

[0064] More specifically, according to one embodiment, a method is provided for determining (e.g., learning) significant directions in the GAN latent space in order to manipulate the semantics in the GAN output. A well-trained GAN model is (in the GAN latent space Z (e.g., z∈Z, Z⊆R d Learn a mapping g that maps a d-dimensional latent code z in ) to an image x = g(z) in image space X. Human-interpretable semantics exist in image space X and include, for example, age, gender, glasses, smile, head posture, or other semantic observations. A set of scoring functions s1···s for K semantics k Given a latent space Z, the semantic space C1···C k The K mappings up to ⊆R are obtained for each. Here,

number

[0065] s1···s k Given the accuracy and independence of the results, if the k-th semantic attribute of the output image changes, the score c will change. z k While this will change accordingly, other semantic scores are expected to remain largely the same. There is a hypothesis that useful information is embedded in the mapping from the GAN latent space to the semantic space, and that this can be used to find semantically significant directions. In particular, c z k To control this, it is proposed to interpolate the latent code according to the gradient direction of such a mapping function. For simplicity of computation, each mapping function s k (g(z)) is the original s k A neural network f trained on pair samples of GAN latent codes and corresponding semantic labels, generated by (g(z)) (for example, a scoring function for ground truth / human perception). k It is parameterized using [a specific method]. Its direction is calculated as follows:

number

number

[0066] [Unraveling the attributes being manipulated] According to one embodiment, a method and / or technique for minimizing semantic entanglement is provided. Attribute entanglement may (e.g., sometimes) occur during the interpolation of latent codes following the direction found as described above. As used herein, and as will be understood by those skilled in the art, two or more semantic attributes are entangled when, during interpolation, the operation of one semantic attribute affects one or more of the other two or more attributes. While it has been observed that non-symmetric semantics may be altered by interpolation along the original direction, such effects can occasionally be eliminated by randomly excluding the dimensions of the direction vector used for interpolation. Thus, it is assumed that among the d-dimensional direction vectors found, only some dimensions cause changes in symmetric semantic attributes, while the other dimensions represent biases learned from the training data. For example, in the direction of increasing a person's age, glasses appear during the operation because glasses and age are correlated in the data.

[0067] Grad-CAM [SCD] +

[17] is a class-discriminating localization technique that uses gradient information to provide a visual explanation for CNN-based models. According to the methods and / or techniques herein, dimensions are removed by filtering based on the magnitude of the gradient from the semantic scoring function. This results in more unraveled control. In particular, Grad-CAM measures the importance of neurons by:

number

number

[0068] According to the embodiment, nz k It is considered the sole activation map, and the importance of the i-th dimension is calculated as follows:

number

[0069] According to the definition of gradient, n in its i-th dimension z k The value of c is caused by a small change in z in the same dimension. z k This shows the rate of change of L. Intuitively, i k n is large z k The dimension is c z k This has a greater impact on L i k small n z k The dimensions are less relevant. However, such unrelated dimensions are another semantic attribute n z m When calculating the control for (m≠k), the magnitude of the gradient may have become large. Therefore, if these dimensions change even slightly when optimizing the k-th semantic, c z m This can affect and lead to problems of attribute entanglement. Therefore, according to one embodiment, L is not smaller than a certain threshold. i k The dimension in the gradient for any k-th semantic attribute with a certain property is considered important, while any important dimension in predicting a semantic attribute that appears to be coupled to the object from the direction of the object is excluded. In one embodiment, if k is entangled with m and they share important dimensions, even important dimensions of k can be excluded by masking.

[0070] Figure 3 illustrates, in pseudocode form, the operation for calculating a new direction from which the attributes have been unraveled in one embodiment. Line 1 shows that the target semantic is in direction nz k is described as being intertwined with another semantics. In line 2, a scoring function f m for the intertwined semantics m and a threshold of t are established to identify important gradients (i.e., those with large gradient magnitudes). In line 3, the gradient (n z k ) of the intertwined semantics m is determined from their respective functions. In line 4, an investigation is made to compare each of the i-dimensional elements L i m with t. Here, the element L i m =|n z m | i . In line 5, for the gradients identified by the threshold, the corresponding gradient magnitude in the dimension of k is set to 0. That is, n z k [i]=0 in i∈E, and in line 6, the final form of n z k is returned (e.g., provided).

[0071] At each step, to increase the semantic score c z k by one through the unraveling of the m-th semantics, z is updated according to the following equation in which n z k' is recalculated according to operation 300 of the pseudo-code algorithm in FIG. 3. [Equation] For example, it will be understood that the operation can loop and repeat so that z’ is calculated two or more times to further interpolate along the direction of the target semantic attribute until an image with desired characteristics (e.g., g(z’)104B) is generated.

[0072] [Details of Embodiment] According to one embodiment, the GAN portion of the network architecture structure 102 (e.g., GAN g 108) is pre-trained on the FFHQ dataset [KLA18], for example, StyleGAN2 [KLA + The structure of

[20] is adapted. Optimization is performed using StyleGAN2[KLA + This is performed in the W space of

[20] . 400 images are generated by StyleGAN2, from which 30 images are manually selected as positive / negative samples for each target semantic, and 10% of the selected pairs are used for evaluation. Multiple binary classifiers (e.g., each of 110 instances) are simultaneously trained on different target semantics using a multi-label learning method to minimize the sum of all binary cross-entropy losses. According to one embodiment, each classifier (e.g., f k ) includes two fully connected layers with hidden layers equal in size to 16, ReLU activation, and sigmoid activation in its output neurons. For the filtering threshold t, n z m Experiments have confirmed that the channel works well when the gradient magnitude is the 100th largest. For the alpha value, 0.4 was used in all experiments.

[0073] Next, in one embodiment, the dimension important for the semantic attribute m is the gradient vector n z k To remove it, we entangle m with k and the gradient vector n of m. z mThe following is determined. Each dimension of the vector is evaluated using the absolute value of the dimension to determine the 100th largest value (for example, a sort can be performed). This 100th largest value becomes a threshold t depending on m during filtering. If each i-th dimension in the vector m is greater than or equal to the 100th largest value, the corresponding i-th dimension in the vector k is set to 0. For each semantic attribute m intertwined with k, the gradient vector of m is determined and evaluated against t in a similar manner, and then the respective dimensions are masked in the gradient vector of k.

[0074] [Results and Discussion] The evaluation results are based on the disclosed methods, systems, and technologies (for example, StyleGAN2 [KLA18] pre-trained on the FFHQ dataset [KLA18] of network architecture 102 according to the described embodiments). +

[20] The above is presented. Figure 4 is a table of 400 images showing the results of the operation for five attributes: age, posture, gender, glasses, and smile. These qualitative examples are edited using the learning direction learned with 60 samples for each attribute. Figure 5 is a table of 500 images showing the attribute unraveling. Figures 6 and 7 are tables of 600 and 700 images showing the results of comparison with other studies [SYTZ20, WLS21, HHLP20], respectively.

[0075] Table 400 shows the results of operations on five different attributes, i.e., the results of operations on a single attribute. For each group of three samples within a row (e.g., 402), the center image is the original image synthesized with StyleGAN2, and the left / right images correspond to the synthesized results based on the original latent code edited according to the negative / positive direction found by the methods or techniques disclosed herein (i.e., as output from network architecture 102 (via GAN g 108)).

[0076] This approach works well for all attributes, both in negative and positive directions. Specifically, for the age attribute, it was found that it can not only create / remove subtle aging effects such as wrinkles and acne, but also modify facial features to a high degree while maintaining a good sense of the person's identity. This approach can also add glasses with virtually no change to irrelevant semantics. This means that many realistic image samples can be created with glasses, even though such semantic attributes are absent (i.e., not present) in the original FFHQ training data.

[0077] Table 500 shows the results of the attribute deconstruction in Figure 5. For each group of three images in each row (e.g., 502, 504, 506, and 508), the first image on the left (e.g., 502A, 502D, 504A, and 504D) is the original image synthesized by StyleGAN2. The adjacent left and right images (502B / 502C, 502E / 502F, 504B / 504C, and 504E / 504F) correspond to the synthesis results (via GAN g 108) based on the original latent code edited according to the negative / positive direction found by the methods or techniques herein (e.g., using relatively small and large interpolation distances). The first two rows, 502 and 504, show that the aging direction is intertwined with the glasses direction. By excluding key dimensions of the glasses classification from the aging direction during optimization, the two attributes are successfully unraveled (e.g., row 504), and similar aging results are achieved. Similarly, for the last two rows, 506 and 508, the approach of excluding key dimensions as disclosed herein has succeeded in finding a new direction in which the gender direction is less intertwined with the smile.

[0078] While the controls found by the methods or techniques disclosed herein are generally independent, we have found that failures still exist where manipulating one semantic affects another, with the most common entanglements being glasses relative to age and smile relative to gender. This result suggests that by excluding key dimensions that cause changes in the logits of entangled attributes during optimization, we can find a slightly different approach in which irrelevant semantics are less affected and the target semantics are still well manipulated.

[0079] [Comparison] The results from the methods and / or techniques taught in this application can be compared, for example, with InterFaceGAN[SYTZ20] and GANSpace[HHLP20] for three semantic attributes (smile, gender, and age). Attributes not supported by some known methods and implementations are not shown. In [SYTZ20], the SVM boundary used the same training data used to train the GAN g and auxiliary network in Figure 1. In [HHLP20], the officially published orientations were used. Table 600 in Figure 6 visualizes the orientations found by each of the three implementations, with the first two being known and the bottom row showing the results according to the embodiments herein (e.g., using network architecture 102 with GAN g 108 and auxiliary network 110). For each group of three samples in the row, the center image is the original image synthesized by StyleGAN2, and the left and right images correspond to the synthesized results based on the original latent code edited by the orientations found by each method to decrease / increase the target semantic score.

[0080] Overall, the methods disclosed herein are superior to the unsupervised GANSpace method, which generates less realistic images with imprecise changes in the target semantics, and find similarities to InterFaceGAN. Nevertheless, the methods disclosed herein are found to be superior in attribute unraveling by Grad-CAM-inspired channel filtering compared to the conditional manipulation technique proposed in [SYTZ20], which adjusts the target direction by subtracting its projection vector in a direction for editing entangled semantics, several examples of which are shown in Figure 7.

[0081] Table 700 (Figure 7) shows a comparison of attribute unraveling results between the method disclosed herein and a single known conventional method (InterFaceGAN [SYTZ20]). The method disclosed herein unravels attributes better during the operation. In the gender-smile entanglement example in the first two rows, the smile attribute is better preserved when editing gender in the second such row (the method disclosed herein). In the age-glasses entanglement example in the bottom two rows, the method disclosed herein finds unraveled control that better preserves irrelevant semantics, such as a person's facial expression and hairstyle.

[0082] An example of GANS use is described in U.S. Patent Application No. 16 / 683,398, filed November 14, 2019, entitled "System and Method for Augmented Reality by translating an image using Conditional cycle-consistent Generative Adversarial Networks (ccGans)," which is incorporated herein by reference.

[0083] In one embodiment, the disclosed techniques and methods include developer-related methods and systems for defining (by condition, etc.) a CNN architecture including a GAN g and an auxiliary classifier f for classifying semantic attributes, wherein the GAN is controllable by learning to interpolate semantic attributes. User-related methods and systems are also shown, such that a trained generator model (e.g., generator g(x)) is used at runtime to process the original image x for image-to-image transformation to obtain a composite image x'=g(x) having (or not having) a particular semantic attribute.

[0084] Figure 8 is a block diagram of GUI800 in one embodiment. GUI800 can be displayed via a display device (not shown). In one embodiment, the user can specify the target semantic attribute and the amount of the desired attribute. For example, +smile, -age can identify the attribute, and an amount such as a percentage can specify the amount to give to the synthesized image when the generator is invoked. Because the interpolation process is continuous, a pre-trained image classifier can be used to quantify the amount (or change) of the semantic attribute and find the desired manipulation intensity.

[0085] GUI800 shows a plurality of input operators 802 provided for receiving selection inputs for each target semantic attribute configured and learned by the CNN architecture 102. In this embodiment, each input operator (e.g., 802A) represents a slider-type operator for each single attribute, identifying fine-grained inputs. Other types of input operators (radial operators, text boxes, buttons, etc.) may be used for fine-grained inputs or other inputs. Operators can accept percentage values, range selections, scalars, or other values. For example, the age attribute may be associated with an integer range and accept fine-grained inputs approximating age in years or decades. Others may select relative ranges for the abundance of an attribute (e.g., small, medium, large). These relative ranges can be associated with fine-grained values ​​such as 15%, 50%, and 85%, or other values.

[0086] The semantic attribute input is applied when the generator is invoked, depending on the selection of the "Apply" operation unit 804. The output of the generator is controlled by semantic attribute operation units derived from the semantic attribute classifiers of each auxiliary network. The semantic attribute input is applied to the original image x(806) via these attribute operation units to generate a composite image x'808 (e.g., the output image). The composite image is controlled by attribute operation units and can be expressed as x'=g(x). The attribute operation units correspond to the semantic attribute input from the interface operation unit 802. All available semantic attributes have operation units, but the user can choose not to change the attributes, and therefore the generator does not need to interpolate along the direction identified within the relevant semantic operation unit.

[0087] The original image 806 can be identified for use via the "original image" operation unit 810 (for example, it can be uploaded, copied from storage, acquired via camera input to obtain a selfie, etc.). The resulting composite image x'808 can be saved via the "save image" operation unit 812.

[0088] In one embodiment, before applying the controls when the generator is invoked, the attributes are untangled so that the composite image is generated with minimal entanglement. In one embodiment, untangling can be enabled or disabled via one or more controls (not shown). For example, in one embodiment, age and glasses can be selectively untangled, smile and gender can be selectively untangled, or both such entanglements can be untangled.

[0089] In the illustrated embodiment, a separate semantic attribute manipulation unit is provided for each learned attribute, but in one embodiment (not shown), fewer manipulation units are provided (for example, for only one attribute or only two attributes). Separate manipulation units are provided for individual attributes, but in one embodiment (not shown), a single manipulation unit can be provided for combined attributes (e.g., age and gender). The multi-attribute manipulation unit is computed by vector operations. To create a face with fewer smiles, more glasses, and an older (higher age) face, the GUI is configured to receive inputs for each direction, i.e., -smile, +glasses, and +age. According to one embodiment of the operation of the computing device, the inputs are associated with each gradient vector (of the relevant semantic attribute manipulation units), and the vectors are added together and normalized. The generator interpolates (linearly) along the computed direction (i.e., the combined direction) to produce an output image.

[0090] In one embodiment, the GUI is provided by a first computing device, and the generator is provided by another computing device remotely located relative to the first computing device. This other computing device may provide a controllable GAN generator as a service. The first computing device providing the GUI is configured to communicate input images and semantic attribute direction inputs (e.g., percentage values ​​or other inputs) to the remotely located computing device providing (e.g., running) the generator. Such a remotely located computing device may provide an application programming interface (API) or other interface for receiving the source image (or selection) and semantic attribute direction inputs. The remotely located computing device can calculate direction vectors and invoke the generator by applying semantic attribute manipulation units. In an alternative embodiment, the computing device providing the GUI may be the same as the computing device providing the generator.

[0091] Those skilled in the art will understand that in one embodiment, in addition to embodiments of a developer (used for, for example, learning time) and a target computing device (used for, for inference time), embodiments of a computer program product are disclosed in which, when instructions stored in a non-temporary storage device (e.g., memory, CD-ROM, DVD-ROM, disk, etc.) are executed, the computing device is made to perform one of the method embodiments disclosed herein. In one embodiment, the computing device includes a processor (e.g., a microprocessor (e.g., CPU, GPU, multiple any identical ones), a microcontroller, etc.) that executes computer-readable instructions, such as those stored in a storage device. In one embodiment, the computing device includes (e.g., "purpose built") circuitry that performs the function of the instructions without the need to read such instructions.

[0092] Furthermore, embodiments relating to e-commerce systems are also shown and described. In one embodiment, a user's computing device is configured as a client computing device with respect to the e-commerce system. The e-commerce system stores, for example, a computer program for such a client computing device. Thus, the e-commerce system has a computer program product as a component, which stores instructions for configuring such a client computing device (e.g., its processing unit) when executed by the client computing device. These and other embodiments will be apparent.

[0093] Figure 9 is a block diagram of a computer system 900. In one embodiment, the computer system 900 includes a plurality of computing devices, which in one embodiment include servers, developer computers (PCs, laptops, etc.), mobile devices such as smartphones and tablets, etc. A network model learning environment 902 is shown, which includes hardware and software for defining and configuring a GAN-based CNN architecture 102 by conditioning, etc. The architecture 102 includes a GAN-based generator 108 (e.g., generator g) and an auxiliary network 110 for the classifier (e.g., a plurality of mapping functions f). In one embodiment, the GAN and the mapping functions are jointly conditioned, and the latent space z of the GAN is updated based on the result of the semantic function as described above.

[0094] In one embodiment, a CNN 102, which combines a generator 108 and an auxiliary network 110, is provided for use on a target device such as one of the mobile devices 910, 912, or other devices such as 913 of the system 900. Mobile device 910 is, for example, a typical user computing device of a consumer user. It is understood that such a user may use other forms of computing devices such as a desktop computer, workstation, etc. Device 913 represents a computing device (and therefore a training data generation device) for pre-generating training data. In this embodiment, such a computing device employs generator 108 to generate additional image data. This additional image data can be easily labeled with semantic attributes and can be used to train a network model (e.g., in a supervised manner). The form factor of device 913 can be a server, laptop, desktop, etc., and does not have to be a consumer-type mobile device such as a tablet or smartphone.

[0095] In one embodiment, the network model learning environment 902 employs a GAN model (generator 108) that has been pre-trained at least partially for an image task (e.g., face image generation). The generator 108 is pre-trained, for example, by using an image dataset 914 stored in a data server 916. In one embodiment, the generator 108 is a model developed "in-house". In one embodiment, the generator 108 is publicly available, for example, through an open-source license. The dataset is similarly developed and available. Depending on the image task and the type of (e.g., supervised) network architecture, the training is supervised, and the dataset is annotated according to such training. In other scenarios, the training is unsupervised, and the data is defined accordingly. In one embodiment, the GAN model (generator 108) is further conditioned, for example, using face images labeled for semantic attributes, in the form shown in Figure 1A.

[0096] In one embodiment, the generator 108 and auxiliary network 110 are incorporated into the augmented reality (AR) application 920. Although not shown, in one embodiment, the application is developed using an application developer computing device for a specific target device having specific hardware and software, particularly an operating system configuration. In one embodiment, the AR application 920 is a native application configured for execution in a specific native environment, such as one defined for a specific operating system (and / or hardware). In one embodiment, the AR application 920 takes the form of a browser-based application, configured to run, for example, in the browser environment of the target device.

[0097] In one embodiment, the AR application 920 is distributed to user devices such as mobile devices 910 and 912 (e.g., downloaded on the user device). Native applications are often distributed via an application distribution server 922 (e.g., a "store" operated by a third-party service), but this is not mandatory.

[0098] In one embodiment (not shown), the AR application 920 does not include the CNN architecture 102 itself (it does not include the generator and auxiliary network). Rather, the application is configured with an interface for communicating with a remote device that provides these components as a service (not shown) (e.g., as a cloud-based service). Storing and running the generator and auxiliary network requires a large amount of resources and may be too large / demanding for some computing devices. Other reasons may also influence the paradigm of the AR application.

[0099] In one embodiment, the AR application 920 is configured to provide the user with an augmented reality experience (e.g., via an interface). For example, processing by the generator 108 applies an effect to an image. The mobile device includes a camera (not shown) for capturing images (still or moving images, whether they are, for example, selfies). In one embodiment, the effect is applied to the image (e.g., moving image) in real time (and displayed on the mobile device's display) to simulate the effect on the user when the video is being captured. As the camera position changes, the effect is applied according to the image of the video being captured, simulating augmented reality. As is understood, real-time operation is constrained by processing resources. In one embodiment, the effect is delayed and not simulated in real time, which may affect the augmented reality experience.

[0100] In one embodiment, computing devices are coupled to communicate over one or more networks (e.g., 922), including wireless networks and others, public networks and others.

[0101] As an example, and not an limitation, the e-commerce system 924 is web-based and provides a browser-based AR application 920A as a component of the e-commerce services provided by the e-commerce system 924. The e-commerce system 924 includes a configured computing device and a data store 8926 (e.g., a database or other configuration). The data store 926 stores data about products, services and related information (e.g., technologies for applying the products). The data store 926 or other data storage device (not shown) stores recommendation rules or forms of other product and / or service recommendations, etc., to help the user select from available products and services. Products and services are presented through a user experience interface displayed on the user's (mobile) computing device. It will be understood that the e-commerce system 924 is a simplified representation.

[0102] In one embodiment, a browser-based AR application 920A (or AR application 920) provides an augmented reality customer experience, such as simulating products, technologies, or services offered or promoted by an e-commerce system 924. In this embodiment, it will be understood that the AR application 920 is also configured to provide e-commerce services, such as through a connection to the e-commerce service 924.

[0103] Examples, but not limited to, products include beauty (e.g., makeup) products, anti-aging or rejuvenation products, and services include beauty, anti-aging or rejuvenation services. Services include treatments or other procedures. Products or services relate to parts of the human body such as the face, hair or nails. In one embodiment, a computing device (such as a mobile device 912) configured in this way provides a face effect unit 912A including a processing circuit. This processing circuit is configured to apply at least one face effect to an original image and generate one or more virtual instances of the original image with the effect applied on the computing device's e-commerce interface, facilitated by an e-commerce system. In one embodiment, the face effect unit 912A generates the original image with the effect applied using a generative adversarial network (GAN)-based generator(g) and auxiliary network 110 as described herein. In one embodiment, the computing device provides a user experience unit 912B including a processing circuit. This processing circuit determines at least one product or service from a data store 926 and generates one or more virtual instances recommended on the e-commerce interface for purchasing the product or service. In one embodiment, at least one product is associated with each facial effect, and the facial effect unit applies each facial effect to provide a virtual trial experience.

[0104] In one embodiment, the user experience unit 912B is configured to display a graphical user interface (e.g., browser-based or otherwise) and to cooperate with the computing device 912 and the e-commerce system 924. In one embodiment, the e-commerce system 924 is thus configured to provide an AR application for execution by a client computing device such as a mobile device (e.g., 912) and to cooperate to provide e-commerce services to the client computing device (e.g., 8912) to facilitate recommendations (of products / services) for AR simulation and to facilitate purchases.

[0105] Accordingly, any computing device, in particular a mobile device, provides a computing device for transforming an image from a first domain space to a second domain space. This computing device includes a storage unit that stores a generative adversarial network (GAN)-based generator (g), which is configured to produce an image controlled for semantic attributes. In one embodiment, the computing device includes a processing unit configured to receive a source image, receive an input for identifying at least one controlled semantic attribute (e.g., including an input for refining the semantic attribute (e.g., in percentages)), provide the image to the generator g to obtain a composite (e.g., a new) image corresponding to the semantic attribute input, and provide the new image for display (e.g., via an AR application 920).

[0106] In one embodiment, the generator is configured to synthesize specific semantic attributes, including facial features such as age, gender, and smile; posture effects; makeup effects; hair effects; nail effects; cosmetic surgery or dental effects, including one of rhinoplasty, facelift, eyelid surgery, implants, ear surgery, teeth whitening, and orthodontics; and appliance effects, including one of eye appliances, mouth appliances, and ear appliances.

[0107] In one embodiment, the composite output of generator g (e.g., a composite image or a new image) is provided to a second network model (e.g., as an unillustrated component of AR application 920 or 920A) for, for example, a face editing task. The second network model may include a second generator or feature detector and simulator to apply effects (e.g., to generate further new images). These effects can be displayed in an AR interface. In one embodiment, the processing unit is configured to provide a composite image (or further new images defined therefrom) in an augmented reality interface to simulate effects applied to an image. In one embodiment, the effects include any of the following: makeup effects, hair effects, nail effects, cosmetic or dental effects, prosthetic effects, or other simulated effects applied. In one embodiment, the source image includes an applicable part of a subject (e.g., face, hair, nails, or body part), such as a user of the device. For example, the input image may be processed by a generator to simulate aging, and the effects pipeline may process the composite image showing aging to apply makeup and / or hair effects to the composite image.

[0108] The network model learning environment 902 provides a computing device configured to perform methods such as those involving conditioning a GAN-based generator and an auxiliary network of classifiers that share the same latent space. Embodiments of the computing device aspects of the network model learning environment 902, or any related embodiments in, for example, the generator or model, will be understood to be applicable to embodiments of the learning method with appropriate adaptation.

[0109] Figure 10 is a flowchart of operation 1000 in one embodiment of this specification. In one embodiment, this operation provides a method such as learning (e.g., by conditioning). In step 1002, the operation is performed through conditioning by a generative adversarial network (GAN) based generator g 108, along with training of an auxiliary network 110. The training updates the latent space Z. In one embodiment, the generator g 108 is trained with labeled images having respective semantic attribute labels, along with training of the auxiliary network 110. Co-learning trains the auxiliary network 110 using the latent space Z, and each gradient of the auxiliary network 110 (e.g., if each semantic attribute has a corresponding binary classifier) ​​is useful in controlling the operation of the GAN-based generator g when generating images.

[0110] In step 1004, the operation provides a generator g 108 and an auxiliary network 110 for use in a computing device to generate an image. The embodiments of the associated computed device and computer program product will be evident as in other embodiments.

[0111] Figure 11 is a flowchart of operation 1100 according to one embodiment of this specification. Operation 1100 is performed by a computing device such as devices 910, 912, 913, or 924.

[0112] The operation in 1102 provides a generator and an auxiliary network that share a latent space. The generator is configured to produce a composite image exhibiting human-interpretable semantic attributes. The auxiliary network includes multiple semantic attribute classifiers, each of which is controlled (by the generator) when generating the composite image from the source images. Each semantic attribute classifier is configured to classify the presence of one (relevant) semantic attribute in the image, thereby providing a meaningful direction for controlling one of the semantic attributes in the generator's composite image.

[0113] When calling the generator to produce a composite image from the source images, the operation in 1104 applies a respective semantic attribute manipulation unit to control each semantic attribute in the composite image produced from the source images. Each semantic attribute manipulation unit corresponds to the significant direction provided by the classifier associated with the semantic attribute.

[0114] A practical application of generator g includes generating and providing more readily available (larger amounts of) labeled synthetic data for use in training (or further training) one or more network models. For example, such data can help improve the results of a face editing model. For instance, in a hair segmentation task, a segmentation model may initially be trained on unbalanced data with unevenly distributed hair colors, which results in unsatisfactory segmentation results for minority hair color groups. To address this problem and create a more balanced dataset, generator g and / or techniques described herein can be used to edit existing labeled data for hair colors and generate rich data for minority groups. Thus, the use of generator 108 (and network 110) may include generating labeled synthetic images for training and modeling, and further including training the model using the labeled synthetic images.

[0115] Practical implementations may include any or all of the features described herein. These and other embodiments, features, and various combinations may be expressed as methods, apparatus, systems, means for performing functions, program products, or other methods that combine the features described herein. Several embodiments have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the processes and technologies described herein. In addition, other steps may be provided or steps may be removed from the described processes, and other components may be added to or removed from the described systems. Accordingly, other embodiments are within the scope of the following claims.

[0116] Throughout this specification and the claims, the words “comprise” and “contain,” and their variations, mean “including, but not limited to,” and are not intended to exclude (and do not exclude) other components, integers, or steps. Throughout this specification, unless the context otherwise indicates, the singular form includes the plural form. In particular, where the indefinite article is used, this specification should be understood to mean both the singular and the plural, unless the context otherwise indicates.

[0117] Features, integers, characteristics, or groups described in relation to a particular aspect, embodiment, or example of the present invention should be understood to be applicable to any other aspect, embodiment, or example, unless they are incompatible. All features disclosed herein (including any appended claims, abstract, and drawings) and / or all steps of any method or process so so disclosed can be combined in any combination, except for combinations in which at least some of such features and / or steps are mutually exclusive. The present invention is not limited to the details of any aforementioned example or embodiment. The present invention extends to any novel one or any novel combination of features disclosed herein (including any appended claims, abstract, and drawings), or any novel one or any novel combination of steps of any method or process disclosed. [References] [ACB17] Martin Arjovsky, Soumith Chintala, and Leon Bottou. Wasserstein generative adversarial networks. In International conference on machine learning, pages 214-223. PMLR, 2017. [BDS18] Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096, 2018. [CCK +18] Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo. Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In Proceedings of the IEEE conference on computer vision and pattern recog-nition, pages 8789-8797, 2018. [CVB21] Anton Cherepkov, Andrey Voynov, and Artem Babenko. Navigating the gan parameter space for semantic image editing. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 3671-3680, 2021. [GPAM + 14] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. Advances in neural information processing systems, 27, 2014. [HHLP20] Erik Harkonen, Aaron Hertzmann, Jaakko Lehtinen, and Sylvain Paris. Ganspace: Dis-covering interpretable gan controls. arXiv preprint arXiv:2004.02546, 2020. [HRU +17] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equi-librium. Advances in neural information processing systems, 30, 2017. [KLA18] Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. arXiv preprint arXiv:1812.04948, 2018. [KLA + 20] Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 8110-8119, 2020. [LLWT15] Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of the IEEE international conference on computer vision, pages 3730-3738, 2015. [PBH20] Antoine Plumerault, Herve Le Borgne, and Celine Hudelot. Controlling generative models with continuous factors of variations. arXiv preprint arXiv:2001.10238, 2020. [SCD + 17] Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision, pages 618-626, 2017. [SYTZ20] Yujun Shen, Ceyuan Yang, Xiaoou Tang, and Bolei Zhou. Interfacegan: Interpreting the disentangled face representation learned by gans. IEEE transactions on pattern analysis and machine intelligence, 2020. [SZ21] Yujun Shen and Bolei Zhou. Closed-form factorization of latent semantics in gans. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 1532-1540, 2021. [VB20] Andrey Voynov and Artem Babenko. Unsupervised discovery of interpretable directions in the gan latent space. In International Conference on Machine Learning, pages 9786- 9796. PMLR, 2020. [WDB + 19] Yan Wu, Jeff Donahue, David Balduzzi, Karen Simonyan, and Timothy Lillicrap. Logan: Latent optimisation for generative adversarial networks. arXiv preprint arXiv:1912.00953, 2019. [WLS21] Zongze Wu, Dani Lischinski, and Eli Shechtman. Stylespace analysis: Disentangled con-trols for stylegan image generation. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 12863-12872, 2021. [ZKL + 16] Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2921-2929, 2016.

Claims

1. A method comprising generating a composite image (g(z')) from an original image using a generator (g) having a latent code (z), The generator manipulates the target semantic attribute (k) in the synthesized image, The above generation is, A significant direction identified from an auxiliary network that classifies each semantic attribute, including the target semantic attribute k, in z, and discovering a significant direction in z with respect to the target semantic attribute k. With respect to the aforementioned target semantic attribute k, z' is defined by optimizing z according to the significance direction, From the dimension of the significant direction at z in the target semantic attribute k, the important dimensions of the respective significant directions in each other semantic attribute (m, m≠k) that are intertwined with the target semantic attribute k are removed by filtering, thereby unraveling the direction. This includes outputting the aforementioned composite image, The auxiliary network shares the generator g with the latent space Z, where z belongs to Z. A method in which a particular dimension is considered important based on the magnitude of the absolute value of the gradient in that dimension.

2. The method according to claim 1, wherein the auxiliary network includes a set of binary classifiers, one for each semantic attribute that the generator can operate on, and the auxiliary network is co-learned to share the latent space Z with the generator g.

3. The method according to claim 1, wherein each of the aforementioned significant directions includes the respective dimensional data vectors obtained from each individual classifier, and each vector includes the direction and rate of the fastest increase in the individual classifier.

4. The method according to claim 1, wherein the filtering includes evaluating each dimension of each significance direction unraveled from the significance direction of the target semantic attribute k, and setting the value of the corresponding dimension of the target semantic attribute k in the significance direction to zero if a particular dimension exceeds a threshold.

5. The latent code z is z' = z + αn z k' Optimized as, Here, n z k This is a vector representing the significant direction of the target semantic attribute k before filtering, n z k' This is a vector representing the significant direction of the target semantic attribute k after filtering, The method according to claim 4, wherein α is a hyperparameter that controls the interpolation direction and step size.

6. The method according to claim 1, further comprising repeating the discover, define, and optimize operations with respect to a latent code z' in order to further manipulate the semantic attribute k.

7. This involves generating a composite image by manipulating multiple target semantic attributes, To discover the significant direction at z for each of the target semantic attributes, as identified by the auxiliary network that classifies each of the target semantic attributes at z, The method according to claim 1, comprising optimizing z according to each of the significant directions.

8. The method according to claim 1, further comprising training another network model using the aforementioned synthesized image.

9. The method according to claim 1, comprising receiving an input that identifies at least one of the semantic attributes controlled for the original image.

10. To provide the generator and the auxiliary network for generating the composite image as a service, To provide an e-commerce interface for purchasing products or services, To provide a recommendation interface for recommending a product or service, To provide an augmented reality experience, an augmented reality interface using the aforementioned synthesized image is provided, The method according to claim 1, comprising any one or more of the following.

11. The method according to claim 1, wherein the target semantic attribute includes age, gender, facial features including a smile, posture effects, makeup effects, hair effects, nail effects, cosmetic surgery or dental effects including one of rhinoplasty, facelift, eyelid surgery, implants, ear reconstruction, teeth whitening, orthodontics, and the effects of orthotic devices including one of eye or mouth or ear orthotic devices.

12. A method comprising the step of providing an augmented reality (AR) interface to provide an AR experience, The AR interface is configured to generate the composite image from the received image using a generator by applying a semantic attribute manipulation unit to control each semantic attribute in the composite image, each of the semantic attribute manipulation units corresponds to a significant direction provided by a classifier associated with the semantic attribute, and the generator and the classifier share a latent space. The generation of the composite image involves using a latent code z' defined by optimizing the latent code z of the latent space according to the significance direction of each of the semantic attributes, and unraveling the direction by filtering out important dimensions of the significance direction of other semantic attributes intertwined with each of the semantic attributes from the dimension of the significance direction of each of the semantic attributes, wherein a particular dimension is considered important based on the magnitude of the absolute value of the gradient in that dimension. The method includes the step of receiving the received image and providing the composite image for the AR experience.

13. The method according to claim 12, comprising processing the composite image using an effect pipeline to simulate the effect, and providing the composite image having the simulated effect for presentation on the AR interface.

14. The computing device includes at least one processor and at least one non-temporary storage device that stores computer-readable instructions for execution by the at least one processor, wherein the instructions cause the computing device to perform the method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • System for and method of generating composite image using local edition

    JP2021111372A

  • Systems and methods for augmented reality using a model of generative image transformations with conditional cycle consistency

    JP2022519003A