Try-on using Inverted GAN

LOHO addresses the challenge of hair style transfer by optimizing the latent space of GANs using two-stage optimization and gradient orthogonalization, allowing for controlled attribute manipulation and achieving high-quality, photorealistic results even under misalignment.

JP7696438B2Active Publication Date: 2025-06-20LOREAL SA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023553744
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-03-02
Filing Date
2022-03-03
Publication Date
2025-06-20
Estimated Expiration
2042-03-03

AI Technical Summary

Technical Problem

Existing methods for hair style transfer using GANs struggle with controlled manipulation of attributes while maintaining photorealism, especially under source-target hair misalignment.

Method used

The Latent Optimization of Hairstyles via Orthogonalization (LOHO) approach, which uses GAN inversion to optimize the latent space by decomposing hair into three attributes (perceptual structure, appearance, and finer style) and applying two-stage optimization and gradient orthogonalization to separately optimize these attributes.

Benefits of technology

LOHO enables controlled manipulation of hair attributes, achieving realistic hair style transfer even under misalignment, with improved photorealism and quality of synthesized images, as evidenced by lower Frechet Inception Distance (FID) scores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696438000017
    Figure 0007696438000017
  • Figure 0007696438000018
    Figure 0007696438000018
  • Figure 0007696438000019
    Figure 0007696438000019
Patent Text Reader

Abstract

Hairstyle transfer is challenging due to differences in hair structure in source and target hairs. Hairstyle Latent Optimization via Orthogonalization (LOHO) is an optimization-based approach that uses GAN inversion to fill hair structural details in the latent space during hairstyle transfer. Hair is decomposed into three attributes: perceptual structure, appearance, and finer style, with losses tuned to model each of these attributes independently. Two-stage optimization and gradient orthogonalization allow for a separated latent space optimization of the three hair attributes. Using LOHO for latent space manipulation, we can manipulate hair attributes individually or jointly and transfer desired attributes from the hairstyle to synthesize novel photorealistic images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] 《Cross - Reference》 This application claims the priority of U.S. Provisional Application No. 63 / 155,842, filed on March 3, 2021, the entire content of which is incorporated herein by reference. This application also claims the priority of French Patent Application No. 2201829, filed on March 2, 2022, the entire content of which is incorporated herein by reference.

[0002] This application relates to computer processing for image processing, improvement of image processing and neural networks, and more particularly to systems, methods, and techniques for try - on of styles, particularly hair styles, using reverse GANs (reverse generative adversarial networks).

Background Art

[0003] Computer processing of images using neural networks has opened up new means for the simulation of effects. The progress of generative adversarial networks (GANs) enables the synthesis of both conditional [15, 32] and unconditional

[19] photorealistic images. In parallel, recent research has achieved impressive latent space manipulation by learning disentangled feature representations

[26] , enabling the manipulation of photorealistic global and local images.

[0004] However, achieving a controlled manipulation of the attributes of synthetic images while maintaining photorealism remains an unsolved problem.

Summary of the Invention

[0005] Style transfer, including hair style transfer, is difficult due to differences in the structure of source and target objects such as hair. In one embodiment, Latent Optimization of Hairstyles via Orthogonalization (LOHO) is an optimization-based approach that uses GAN inversion to fill in the details of the hair structure in the latent space during hair style transfer. Hair is decomposed into three attributes: perceptual structure (e.g., shape), appearance, and finer style, and includes losses adjusted to model each of these attributes independently. Two-stage optimization and gradient orthogonalization enable the optimization of separate latent spaces for the three hair attributes. By using LOHO for latent space manipulation, users can manipulate hair attributes individually or jointly, transfer desired attributes from a hair style, and synthesize new photorealistic images. In one embodiment, the LOHO approach (e.g., two-stage optimization and gradient orthogonalization for the optimization of a latent space with separated style attributes) can be generalized to the transfer of other styles such as clothing.

Brief Description of the Drawings

[0006]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

[0007]

[0008]

[0009]

[0010]

[0011]

[0012]

[0013]

[0014]

[0015]

[0016]

[0017]

DETAILED DESCRIPTION OF THE INVENTION

[0018] Various embodiments are detailed herein with respect to the transfer of hairstyles. It will be understood by those skilled in the art that other styles of transfer tasks are described herein and may be implemented using techniques, methods, and apparatuses adapted to other styles of transfer tasks.

[0019] In one embodiment, a user can perform semantic and structural editing on their portrait image using fine-grained control. As a particular challenging and commercially attractive example, the transfer of hairstyles is evaluated and described herein, and the user can transfer hair attributes from multiple independent source images to manipulate their portrait image. In one embodiment, Latent Optimization of Hairstyles via Orthogonalization (LOHO) is a two-stage optimization process in the latent space of a generative model such as a generative adversarial network (GAN) [12,18]. An exemplary technical contribution is that the transfer of attributes is controlled by orthogonalizing the gradients of the transferred attributes so that the application of one attribute does not interfere with other attributes.

[0020] Previous research on the transfer of [

[30] ] hair styles has used complex pipelines of GAN generators to achieve realistic transfer of hair appearance, each specialized for a particular task such as hair synthesis or background inpainting. However, using a pre-trained inpainting network to fill the holes left by misaligned hair masks results in blurry artifacts. According to one embodiment, to generate a more realistic synthesis from the transferred hair shape, the prior distribution of a single pre-trained GAN for generating faces is called upon to fill in the missing shape and structure details.

[0021] LOHO achieves realistic hair style transfer even under the aforementioned source-target hair misalignment. LOHO directly optimizes the extended latent space and noise space of the pre-trained StyleGANv2 [

[20] ]. Using a carefully designed loss function, the LOHO approach decomposes hair into three attributes, namely perceptual structure (e.g., shape), appearance, and finer style. Each of these attributes is then modeled individually, thereby enabling better control over the synthesis process. Furthermore, LOHO adopts a two-stage optimization to significantly improve the quality of the synthesized images, with each stage optimizing a subset of the losses in the objective function. Some of the losses are sequentially optimized due to their similar design and are not jointly optimized under the LOHO approach. Finally, LOHO uses gradient orthogonalization to explicitly separate the hair attributes during the optimization process.

[0022] Figures 1A, 1B, 1C, 1D, and 1E are images 100, 102, 104, 106, and 108 for showing transfer samples of hairstyles synthesized using LOHO according to one embodiment. For a given portrait image 100 and 106 in Figures 1A and 1D, LOHO can manipulate hair attributes based on a plurality of input conditions. The inserted images (e.g., 102A, 104A, and 18A) represent target hair attributes in the order of appearance and finer styles, structures, and shapes. LOHO can convey appearance and finer styles (e.g., shown in Figure 1B) and perceptual structures (e.g., shown in Figure 1B) without changing the background.

[0023] Furthermore, LOHO can change multiple hair attributes simultaneously and independently (e.g., as shown in Figure 1C).

[0024] According to the features of the LOHO approach, the following are provided: · A new approach for performing hair style transfer by optimizing the extended latent space and noise space of StyleGANv2. · An objective function including multiple losses for modeling each important hair style attribute. · A two-stage optimization strategy that leads to a significant improvement in the photorealism of the synthesized images. · Introduction of gradient orthogonality to a general method for jointly optimizing attributes in a latent space without interference. The effectiveness of gradient orthogonality was qualitatively and quantitatively demonstrated. · Evaluation using the calculated Frechet Inception Distance (FID) score for hair style transfer on in-the-wild portrait images in a real environment. The FID is used to evaluate the generative model by calculating the distance between the Inception

[29] features of real and synthetic images within the same domain. The calculated FID score indicates that the framework and related methods and techniques according to the embodiments may be superior to the state-of-the-art (SOTA) hair style transfer results. 《Related Research》

[0025] Generative adversarial networks. Generative models, especially GANs, have been very successful across a variety of computer vision applications such as data augmentation for discriminative tasks like image-to-image translation [15, 32, 40], video generation [34, 33, 9], and object detection

[24] . GANs [18, 3] transform a latent code into an image by learning the distribution underlying the training data. A more recent architecture, StyleGANv2

[20] , has set a benchmark for generating realistic human faces. However, training such networks requires a significant amount of data, significantly raising the barrier to training SOTA-GANs for specific use cases such as hair style transfer. As a result, methods built using pre-trained generators are becoming the de facto standard for performing various image manipulation tasks. In one embodiment, StyleGANv2

[20] is utilized as a representative pre-trained face synthesis model, and an optimization approach for using the pre-trained generator for controlled attribute manipulation is outlined.

[0026] Embedding in the latent space. Understanding and manipulating the latent space of GANs via inversion has become an active area of research. GAN inversion involves embedding an image into the latent space of a GAN such that the synthetic image resulting from that latent embedding is the most accurate reconstruction of the original image. I2S [1] is a framework that can reconstruct an image by optimizing the extended latent space W + + of a pre-trained Style-GAN

[19] . The sampled embedding W + is the concatenation of 18 different 512-dimensional w vectors, one for each layer of the StyleGAN architecture. I2S++ [2] further improves the reconstruction quality of the image by additionally optimizing the noise space N. Furthermore, by including a semantic mask in the I2S++ framework, users can perform tasks such as inpainting and overall editing of images. Recent methods [13, 27, 41] learn an encoder that directly maps an input from the image space to the latent space W + . In one embodiment, LOHO follows GAN inversion in that it optimizes the W + space and the noise space N of the recent StyleGANv2 to perform semantic editing of hair on portrait images. In one embodiment, LOHO further utilizes the GAN inversion algorithm for the simultaneous manipulation of spatial local attributes such as hair structure from multiple sources while preventing interference between competing objectives of different attributes.

[0027] Transfer of hairstyles. Hair is a difficult part in human face modeling and synthesis. Previous studies on hair modeling involve capturing hair geometry [8,7,6,35] and using this hair geometry downstream for interactive hair editing. However, these methods cannot capture the main visual factors, thereby degrading the quality of the results. Recent studies [16,23,21] have shown progress regarding the use of GANs for hair generation, but these methods do not allow intuitive control over the synthesized hair. MichiGAN

[30] proposed a conditional synthesis GAN that enables controlled manipulation of hair. MichiGAN separates hair into four attributes by specifying intentional mechanisms and representations, and generates SOTA results for hair appearance changes. Nevertheless, MichiGAN has difficulty handling hair transfer scenarios with arbitrary shape changes.

[0028] This is because MichiGAN implements shape changes using a separately trained inpainting network to fill the "holes" generated during the hair transfer process. In contrast, aspects of the method herein call on the prior distribution of a pre-trained GAN to "fill" in the latent space rather than in pixel space. Compared to MichiGAN, aspects of the method herein generate more realistic synthetic images in difficult cases where the shape of the hair changes. <Methodology> 《Background》

[0029] The objective function proposed by Image2StyleGAN++ (I2S++) [2] is:

Equation

[0030] The change of the I2S++ objective function in Equation 1 is used by [2] to improve image reconstruction, image crossover, image inpainting, local style transfer, and other tasks. For hair style transfer, it is desirable to perform both image crossover and image inpainting. To transfer a certain hair style to another person, crossover is required, and the remaining areas where the hair of the original person was painted over are needed. 《Framework》

[0031] Figure 2 shows a two-stage network framework with background blending (inpainting) 200 for LOHO according to one embodiment. The network framework 200 represents an inference time framework as opposed to a training framework. The GAN generator 202 comprises a pre-trained GAN for style transfer. In stage 1 (206), starting from the "mean" face 204 (I G ), the network framework 200 reconstructs the target identity of (I1 (208)) and the target perceptual structure of the hair from (I2 (210)). In stage 2 (212), while maintaining the perceptual structure via gradient orthonormalization (GO), the network framework 200 transfers the finer style and appearance of the target hair from I3 (214). Finally, I3 is blended with the background of I1.

[0032] For the hair style transfer problem, three portrait images of a person are provided: I1, I2, and I3 (208, 210, and 214). Consider transferring the attributes of the hair shape and structure of person 2 (of I2), and the appearance and finer style attributes of the hair of person 3 (of I3) to person 1 (of I1). M1 f (208A) is used as the binary face mask of I1, and M1 h , M2 h and M3 h (not shown) are used as the binary hair masks of I1, I2, and I3. Next, M2 h is separately dilated and eroded by about 20% to generate the dilated version M2 h,d and the eroded version M2 h,e (210A). M2 h,ir ≡ M2 h,d - M2 h,e is the ignore region (e.g., background without face, without hair) that requires inpainting. In this embodiment, M2 h,ir is not optimized. Instead, StyleGANv2 (GAN generator 202) is called to inpaint the relevant details within this region. This feature enables the network framework 200 to perform the transfer of the hair shape in situations where the hair shapes of person 1 and person 2 are misaligned.

[0033] In two stages 206 and 212, the segmentation network 218 is used to process the synthetic image (as the input to stage 1 206 and after it is refined and provided as the input to stage 2 212) and the input images (I1, I2, and I3) to define the respective segmentation masks. In two stages 206 and 212, according to one embodiment, a pre-trained CNN 220 (e.g., VGG

[28] ) for face image processing is used to extract high-level features as further described.

[0034] In Embodiment 200, the background of I1 is not optimized. Therefore, in order to recover the background, in 216, in this embodiment, the background of I1 is soft-blended with the foreground (foreground, hair and face) of the synthetic image I G Specifically, in this embodiment, GatedConv

[36] (not shown) is used to inpaint the masked foreground region of I1, and then blending is performed. 《Objective》

[0035] A loss is used to monitor the relevant regions of the synthetic image in order to perform hair style transfer. To keep the notation simple, I G ≡G(W + +N) is used as the synthetic image, and M G f (204A) and M G h are used as the corresponding face region and hair region.

[0036] Reconstruction of Identity. In order to reconstruct the identity of Person 1, in one embodiment, the Learned Perceptual Image Patch Similarity (LPIPS)

[39] loss is used. LPIPS is a perceptual loss based on human similarity judgment and is therefore well-suited for face reconstruction. To calculate the loss, a pre-trained VGG

[28] 220 is used to extract high-level features

[17] for both. The features are extracted from all five blocks of VGG220 in one embodiment and summed to form a face reconstruction objective:

Equation

[0037] Reconstruction of hair shape and structure. To recover the hair information of Person 2, monitoring is performed via the LPIPS loss. However, M2 h If naively used as the target hair mask, the generator 202 may synthesize hair in undesirable regions of I G . This is especially true when the target face region and the hair region are not properly aligned. To solve this problem, the eroded mask M2 h,e is used to impose a soft constraint on the target placement of the synthesized hair. M2 h,e is combined with M2 h,ir , and the generator can process misaligned pairs by inpainting relevant information in non-overlapping regions. To calculate the loss, features from blocks 4 and 5 of VGG220 are extracted corresponding to the hair regions of I2, I G to form a perceptual structure objective for the hair:

Equation

[0038] Transfer of hair appearance. The hair appearance refers to the overall consistent color of the hair, independent of the hair shape and structure. As a result, it can be transferred from samples of different hair shapes. To transfer the appearance of the target, in one embodiment, 64 feature maps are extracted from the shallowest layer of VGG (relu1_1) such that they best explain the color information. Then, average-pooling is performed within the hair regions of each feature map to discard the spatial information and capture the global appearance. R64×1 The estimated value of the average appearance A is [Number] obtained by where φ(x) represents 64 VGG feature maps of the image x, and y represents the associated hair mask. Finally, the squared L2 distance is calculated to give the hair appearance objective: [Number]

[0039] More detailed transfer of hair. In addition to the overall color, hair also includes finer details such as wisp styles and shading changes between hair strands. Such details cannot be captured by just the appearance loss that estimates the overall average. Therefore, a better approximation is needed to calculate the various finer styles between hair strands. The Gram matrix

[10] captures finer hair details by calculating the second-order association between high-level feature maps. In one embodiment, the Gram matrix is calculated after extracting features from the {relu1_2; relu2_2; relu3_3; relu4_4} layers of VGG. [Number] where γ l represents the feature map extracted from layer l in R HW×C and g l represents the Gram matrix of layer l. Here, C represents the number of channels, and H and W are the height and width respectively. Finally, the squared L2 distance is calculated as follows. [Number]

[0040] Noise Map Regularization. When explicitly optimizing the noise map \(n\in N\), the actual signal may be inserted into the noise map by the optimization. To prevent this, in one embodiment, a regularization term for the noise map

[20] is introduced. For each noise map larger than \(8\times8\), in one embodiment, a pyramid down network is used to reduce the resolution to \(8\times8\). The pyramid network averages the neighborhoods of \(2\times2\) pixels at each step. Additionally, in one embodiment, the noise map is normalized to have zero mean and unit variance, generating a noise objective:

Equation

[0041] Combining all the losses, the overall optimization objective is as follows.

Equation

[0042] Two-stage Optimization. Considering the similar nature of the losses \(L\) r , \(L\) a and \(L\) s , optimizing all the losses jointly from the start is assumed to cause the hair information of person 2 to compete with the hair information of person 3, resulting in undesirable synthesis. To alleviate this problem, the overall objective is optimized in two stages. In stage 1, only the target identity and the perceptual structure of the hair are reconstructed, i.e., in Equation 8, \(\lambda\) a and \(\lambda\) sIt is set to zero. In stage 2, stage 1 provides a better initialization for the stage, thereby converging the model.

[0043] However, this technique itself has drawbacks. It is that there is no monitoring to maintain the perceived structure of the reconstructed hair after stage 1. This lack of monitoring allows StyleGANv2 to call the prior distribution to inpaint or remove hair pixels, thereby canceling the initialization of the perceived structure found in stage 1. Therefore, L r needs to be included in the second stage of optimization.

[0044] Gradient orthogonality. L r captures, by design, all the hair attributes of person 2, namely the perceived structure, appearance, and finer style. As a result, the gradient of L r competes with the gradient corresponding to the appearance and finer style of person 3. This problem is addressed by manipulating the gradient so that the appearance and finer style information is removed. More specifically, the perceptual structure gradient of L r is projected onto a vector subspace orthogonal to its appearance and finer style gradients. This allows the appearance and finer style of person 3's hair to be transferred while maintaining the structure and shape of person 2's hair.

[0045] Latent space W + assuming optimization of, the computed gradients are as follows.

Number

Number

Number

[0046] Dataset. In one embodiment, the Flickr-Faces-HQ (FFHQ) dataset

[19] containing 70,000 high-quality images of human faces was used. Flickr-Faces-HQ has significant variations in ethnicity, age, and hair style patterns. In one embodiment, tuples of images (I1, I2, I3) were selected based on the following constraints: (a) at least 18% of the pixels in each image within the tuple should contain hair, and (b) the respective face regions of I1 and I2 should be somewhat aligned. To enforce these constraints, in one embodiment, a Graphonomy segmentation network

[11] was used to extract hair and face masks, and 2D-FAN [4] was used to estimate 68 2D facial landmarks. For all, the intersection over union (IoU) and pose distance (PD) on the combination of I1 and I2 were calculated using the corresponding face masks and facial landmarks. Finally, in one embodiment, the selected tuples were distributed into three categories of "easy", "moderate", and "difficult" such that both the following IoU and PD constraints were satisfied as shown in Table 1.

Table 1

[0047] Training parameters. In one embodiment, the Adam optimizer

[22] was used with an initial learning rate of 0.1 and annealed using the cosine schedule

[20] . In one embodiment, the optimization was performed in two stages, each stage consisting of 1000 iterations. Based on ablation studies, in one embodiment, 40 appearance loss weight λ a , 1.5×10 4 finer style loss weight λ s and 1×10 5 noise regularization weight λ n were selected. And the remaining loss weights were set to 1. 《Effect of Two-Stage Optimization》

[0048] Figure 3 is a four-column image array 300 showing the effect of two-stage optimization according to one embodiment. In the image array 300, the first column (300A) shows the reference image, the second column (300B) shows the identity (e.g., Person 1), the third column (300C) shows the composite image when the losses are optimized together, and the fourth column (300D) shows the composite image via two-stage optimization + gradient orthogonality.

[0049] Optimizing all the losses together in the objective function branches the framework. During identity reconstruction, hair transfer fails (Column 300C in Figure 3). The structure and shape of the synthesized hair are not preserved, causing undesirable results. On the other hand, performing two-stage optimization clearly improves the synthesis process, resulting in the generation of realistic images that match the provided references. Not only is the identity reconstructed, but the hair attributes are also transferred according to the desired requirements. 《Effect of Gradient Orthogonality》

[0050] Figure 4 is an image array 400 showing the effect of gradient orthogonalization (GO) according to an embodiment. The first row (400A) shows four reference images (from left to right) that show identity, target hair appearance and finer style, target hair structure and shape (mask). The second row (400B) shows two image pairs, e.g., i) (a) and (b), and ii) (c) and (d), each containing the respective composite images and their corresponding hair masks for the non-GO method and the GO method. FIGS. 5A and 5B are graphs 500 and 502 showing the effect of GO according to an embodiment. Graphs 500 and 502 each show the LPIPS hair reconstruction loss (GO vs non-GO) for the iterations and trends of

Number

[0051] Two variants (embodiments) of the framework are compared: non-GO and GO. GO involves manipulating the gradient of L r whereas non-GO leaves the gradient of L r untouched. Non-GO cannot maintain the target hair shape and causes an increase in L r after 1000 iterations (FIGS. 4, 5A, 5B) in stage 2 of the optimization. Appearance and finer style losses that are invariant in position do not contribute to the shape. On the other hand, GO maintains the target hair shape using the reconstruction loss in stage 2. As a result, the IoU is calculated between M2 h and M G h and increases from 0.857 (non-GO) to 0.932 (GO).

[0052] Regarding gradient disentanglement, over time g R2 and (g A2 + g S2) and the similarity between them decreases, indicating that the embodiment of the framework with GO can unravel the hair shape of Person 2 from its appearance and more delicate style (Figs. 5A, 5B). This unraveling enables seamless transfer of the appearance and more delicate style of Person 3's hair to the synthetic image without causing divergence in the model. Here, the GO version of the framework is used for comparison and analysis. 《Comparison with SOTA》

[0053] Transfer of hairstyle. The GO version of this framework was compared with the SOTA model MichiGAN. MichiGAN includes separate modules for (1) the appearance of hair, (2) the shape and structure of hair, and (3) estimating the background. The appearance module bootstraps the generator with its output feature map and replaces the randomly sampled latent code in a conventional GAN

[12] . The shape and structure module outputs a hair mask and an orientation mask and unnormalizes each SPADE ResBlk

[25] in the backbone generation network. Finally, the background module progressively blends the output of the generator with the background information. Regarding training, MichiGAN follows a pseudo-supervised regime. Specifically, features (estimated by the modules) from the same image are fed into MichiGAN to reconstruct the original image. At test time, FID is calculated for 5000 images at a resolution of 512 pixels randomly sampled from the test split of FFHQ.

[0054] To ensure that the results are comparable, the FID score

[14] for LOHO was calculated following the above procedure. In addition to calculating the FID for the entire image, in one embodiment, the score was calculated depending only on the synthesized hair and face regions where the background was masked. Achieving a low FID score on the masked image means that the LOHO model can actually synthesize realistic hair and face regions. This embodiment is called LOHO-HF. Since the background inpainter module of MichiGAN is not publicly available, in one embodiment, GatedConv

[36] is used to inpaint the features related to the masked hair region.

[0055] Quantitatively, LOHO outperforms MichiGAN, achieving an FID score of 8.419, while MichiGAN achieves 10.697 (Table 2). This improvement indicates that the LOHO optimization framework can synthesize high-quality images. LOHO-HF achieves an even lower score of 4.847, demonstrating the excellent quality of the synthesized hair and face regions. 5000 images were randomly sampled uniformly from the test set of FFHQ. Note that the symbol "↓" indicates that the smaller the numerical value, the better the result.

Table 2

[0056] Figure 6 is an image array 600 showing a qualitative comparison between MichiGAN and LOHO according to one embodiment. In each of the six rows showing six respective examples, the first column (narrow) (600A) shows a reference image, the second column (600B) shows an identity person, the third column (600C) shows the output of MichiGAN, and the second column shows the output of LOHO (zoomed in for better visual comparison). In the first and second rows, the examples show that while MichiGAN "copy-pastes" the target hair attribute, LOHO blends the attributes, thereby synthesizing a more realistic image. In the third and fourth rows, the examples show that LOHO handles misaligned examples better than MichiGAN. In the fifth and sixth rows, examples are shown where LOHO transfers the correct style information.

[0057] Qualitatively, better results can be synthesized for examples where the method according to LOHO is difficult. LOHO blends the target hair attribute naturally with the target face as shown in the image array 600. MichiGAN simply copies the target hair onto the target face, causing a lighting mismatch between the two regions. LOHO can handle pairs with various degrees of misalignment, while MichiGAN cannot do this as it relies on blending background and foreground information in pixel space rather than in the latent space. Finally, LOHO transfers relevant style information comparable to MichiGAN. In fact, due to the addition of a style objective that optimizes second order statistics by matching the Gram matrix, LOHO can synthesize hair with various colors even when the original person regarding the hair shape has a uniform hair color, as in the bottom two columns (the fifth to sixth columns) of Figure 6.

[0058] Quality of identity reconstruction. LOHO was also compared with two recent image embedding techniques: I2S[1] and I2S++[2]. I2S is in the latent space W +Introduce a framework that can reconstruct high-quality images by optimizing. I2S also shows how the latent distance calculated between the potential code W of the optimized style * and the average face's W^ is related to the quality of the synthesized image. In addition to I2S, I2S++ optimizes the noise space N to reconstruct images with high PSNR values and SSIM values. Therefore, to evaluate the ability of LOHO to reconstruct high-quality and target identities, similar metrics are calculated on the face region of the synthesized image. Since inpainting in the latent space is an essential part of the results of LOHO, it is compared with the performance of I2S++ for inpainting images with a resolution of 512 pixels.

[0059] The model (LOHO) can achieve equivalent results despite performing the difficult task of hair style transfer (Table 3). I2S shows that the acceptable latent distance for valid human faces is in [30:6; 40:5], indicating that LOHO is within that range. Furthermore, the PSNR score and SSIM score of LOHO are better than those of I2S++, proving that LOHO reconstructs identities that satisfy local structural information.

Table 3

[0060] According to an embodiment, the LOHO framework and related techniques can edit the attributes of portrait images in a real environment. In this setting, after an image is selected, attributes are edited individually by providing a reference image. For example, the structure and shape of the hair can be changed while leaving the appearance of the hair and the background unedited. According to an embodiment, the LOHO framework and related techniques calculate non-overlapping hair regions and spatially fill in the details of the associated background. Following an optimization process, the synthesized image is blended with the inpainted background image. The same applies to changing the appearance and finer style of the hair. LOHO separates the hair attributes and enables them to be edited individually and together, thereby yielding desirable results. Accordingly, FIG. 7 shows an image array 700 of examples representing individual attribute edits, and FIG. 8 shows an image array 800 of examples representing multiple attribute edits. The image array 700 includes an example of appearance and finer style (left example) in a first sub-array 700A and an example of shape (right example) in a second sub-array 700B. The results in FIG. 7 show that the model can edit individual hair attributes without interfering with each other. In FIG. 8, the image array 800 represents results showing that the LOHO framework and related techniques according to an embodiment can edit hair attributes together without interfering with each other. <Limit>

[0061] FIGS. 9A and 9B are image arrays 900 and 902 showing examples of misalignment according to an embodiment. According to an embodiment, the LOHO framework and related techniques are susceptible to extreme cases of misalignment (FIG. 9). In this study, such cases are classified as difficult. They cause the framework and related techniques to synthesize unnatural hair shapes and structures. A GAN-based alignment network [38,5] can be used to transfer hair pose or alignment across difficult samples.

[0062] Figures 10A and 10B are image arrays 1000 and 1002 showing examples of hair detail carry - over according to one embodiment. This may be due to incomplete segmentation of the hair in graphonomy

[11] . This problem can be mitigated using a more sophisticated segmentation network [37, 31]. 《Application to the Real World》

[0063] Figure 11 is a diagram of a computer network 1100 showing a development computing device 1102, a website computing device 1104, a cloud computing device 1105, an application delivery computing device 1106, and respective edge computing devices, namely a smartphone 1108 and a tablet 1110, according to one embodiment. The computing devices are coupled via a communication network 1112. The computer network 1100 is simplified. For example, the website computing device 1104, the cloud computing device 1105, and the application delivery computing device 1106 are exemplary devices of their respective website, cloud, and application delivery systems. The communication network 1112 can include multiple wired and / or wireless networks that can include private networks and public networks.

[0064] In this embodiment, the development computing device 1102 is coupled to a data store 1114 (which can include a database) that configures (and can include training) a network framework 1116 and stores one or more data sets for testing and the like. According to one embodiment, the network framework 116 includes a GAN generator and is configured for two - stage optimization to perform style transfer, particularly hair style transfer.

[0065] The data store 1114 can store software, other components, tools, etc. to assist in development and implementation. In another embodiment (not shown), the dataset is stored in the storage device of the development computing device 1102.

[0066] The development computing device 1102 is configured to define a network framework 1116 according to the embodiments described herein. For example, the development computing device 1102 is configured to configure the network framework of FIG. 2. In one embodiment, the development computing device 1102 is configured to incorporate a pre-trained GAN configured for style transfer, such as StyleGANv2 or a variant thereof, as shown in FIG. 2.

[0067] In one embodiment, the development computing device 1102 defines a network framework 1116 for execution on one server computer device accessible via the website computing device 1104 or the website computing device 1104.

[0068] In one embodiment, the development computing device 1102 defines a network framework 1116 for execution on the cloud computing device 1105. The development computing device 1102 (or another one not shown) incorporates an interface to the network framework into the application 1120A, such as for a website and / or application 1120B for the application delivery computing device (e.g., 1106) for delivery to respective edge devices such as the smartphone 1108 and the tablet 1110.

[0069] The present embodiment of FIG. 11 does not show storing and executing the network framework itself on an edge device such as a tablet or smartphone, and the optimization process within such a framework requires significant processing resources. Execution on an edge device with typical resources for such devices takes a (relatively) long time (about 10-20 minutes) for a single style transfer. It is also possible to execute the framework on a home PC, game console, or other generally consumer-oriented device with a similar runtime. However, since the runtime is not currently considered sufficient to be interactive (even so), FIG. 11 shows a more practical use case where the network framework 1116 is provided by a remote server (e.g., a website or cloud device). In this paradigm, identity and style attribute images are submitted to the server, and the user waits for a response (e.g., a composite image incorporating the identity and the transferred style).

[0070] In one embodiment, an application delivery computing device 1106 provides an application store service (an example of an e-commerce service) and delivers an application for execution on a target device that runs a supported operating system (OS). Examples of application delivery services by computing devices include Apple's App Store (registered trademark) for iPhone (registered trademark) or iPad (registered trademark) devices running iOS (registered trademark) or iPadOS (registered trademark) (both trademarks of Apple Inc., Cupertino, CA). Another exemplary service via applicable computing devices is Google Play (registered trademark) (a trademark of Google LLC, Mountain View, CA) for smartphones and tablet devices from various sources running the Android (registered trademark) OS (a trademark of Google LLC, Mountain View, CA). In this embodiment, smartphone 1108 receives application 1120A from web site computing device 1104, and tablet 1110 receives application 1120B from application delivery computing device 1106.

[0071] In the current paradigm for both website and application delivery examples, the network framework 1116 is not communicated to the edge device. The edge device provides access to the network framework 1116 that is executed on behalf of the edge device (via respective application interfaces). For example, a website computing device executes the network framework 1116 for application 1120A, and a cloud computing device executes it for application 1120B. Applications 1120A and 1120B are each configured for hair style try-on (an effect simulation application) in respective embodiments and provide a virtual and / or augmented reality experience. The operation is further described below in this specification.

[0072] FIG. 12 is a block diagram of a representative computing device 1200. The computing devices of FIG. 11 are similarly configured according to their respective needs and functions. The computing device 1200 includes a processing unit 1202 (e.g., one or more processors, such as a CPU and / or GPU, or other processors, etc., including at least one processor in one embodiment), a storage device 1204 that stores computer-readable instructions (and data) (in one embodiment, at least one storage device that can include memory). The computer-readable instructions (and data), when executed by the processing unit (e.g., a processor), cause, for example, a method to be executed on the computing device. The storage device 804 can include any of a memory device (e.g., RAM, ROM, EEPROM, etc.), a solid state drive (e.g., a semiconductor memory device / IC that can define flash memory), a hard disk drive or other type of drive and tape, disks (e.g., CD-ROM, etc.) and other storage media. Additional components can include a communication unit 1206 for coupling the device to a communication network via wired or wireless means, an input device 1208, and an output device 1210 that can include a display device 1212. In some examples, the display device is a touch screen device that provides an input / output device. The components of the computing device 1200 are coupled via an internal communication system 1214 that can have external ports for coupling to additional devices.

[0073] In some examples, the output device includes a speaker, bell, light, audio output jack, fingerprint reader, etc. In some examples, the input device includes a keyboard, button, microphone, camera, mouse or pointing device, etc. Other devices (not shown) can include a positioning device (e.g., GPS).

[0074] The memory device may include, in one example, an operating system 1216, a user application 1218 (which may be one of applications 1120A or 1120B), a browser 1220 (a type of user application) for browsing websites and executing executables such as application 1120A to access the GAN generator 1116 received from the website, and data 822 for storing images and / or video frames from a camera or data received by other means.

[0075] In FIG. 11, the (data) items to be communicated (described later) are shown adjacent to each communication connection between the computing device and the communication network. Items positioned adjacent to a particular computing device are received by that device, and items positioned closer to the communication network are communicated from each computing device to another device as described hereinafter in this specification.

[0076] Continuing to refer to FIG. 11, in one example, a user of the smartphone 1108 uses a browser to visit a website provided by the website computing device 1104. The smartphone 1108 receives an application 1120A (e.g., a web page and related code and / or data) that provides access to the network framework 1116. In this example, the application is an effect simulation application such as a hair style try-on on an application that provides a virtual and / or augmented reality experience. The user uses a camera to acquire a still image or a video image (e.g., a self-shot image), and this source image is communicated by an application for processing in the network framework 1116 as image I11122. (When provided as a video, a single image (e.g., a still image) can be extracted therefrom.)

[0077] The user selects reference images (e.g., images I2 and I3) from the memory 1124 via a graphical user interface or the like provided by the application 1120A, which represent i) the shape and structure of the hair (image I2), ii) the appearance of the hair (image I3), and (iii) the finer style of the hair (image I3) to be tried on. Each of i), ii), and iii) includes respective hairstyle attributes.

[0078] The effect (attribute characteristics) of the hairstyle trial is generated and / or the resulting image (I G ) is transferred to 1226 using the network framework 1116 while maintaining the identity represented by the image I11122. The obtained image 1226 (I G ) is returned to the smartphone 1108 and displayed via its display device. In one embodiment, I G is displayed on the graphical user interface for comparison with I1. In one embodiment, I G is displayed on the graphical user interface for comparison with all of I1, I2, and I3.

[0079] In one embodiment, the resulting image 1226 is stored in a memory device. In one embodiment, the resulting image 1226 is shared (communicated) via any of social media, text messages, emails, etc.

[0080] In one embodiment, the website computing device 1104 enables services (such as e-commerce services) and facilitates the purchase of hair products or hair styling products, such as one or more products, associated with a reference image virtually tried on via the application 1120A. In one embodiment, the website computing device 1104 provides a recommendation service for recommending hair products or hair styling products. In one embodiment, the website computing device 1104 provides a recommendation service for recommending services (such as hair or hair styling services). Hair or hair styling products can include shampoos, conditioners, oils, serums, vitamins, minerals, enzymes, and other hair or scalp treatment products; coloring agents; sprays, gels, waxes, mousses, and other styling products for application to the hair; hair or scalp tools or implements such as combs, brushes, hair dryers, curling wands, straightening wands, flat irons, scissors, razors, rollers, massage tools, etc.; and accessories including clips, hair ties, scrunchies, bands, etc. Hair or hair styling services can include cutting, coloring, styling, straightening, or other hair and scalp treatments, hair removal, hair replacement / wig services, and consultations for them.

[0081] In one embodiment, application 1120A provides an interface for engaging with the user in a conversational manner to obtain hairstyle, lifestyle, and / or user data that may include image I11122. In one embodiment, the data is analyzed and recommendations are generated. The recommendations can include the selection of reference images from memory 1124. Pairs of reference images (e.g., a particular recommendation I2 having a particular recommendation I3) may be presented for all recommended hairstyles. In some cases, the recommended images I2 and I3 may be the same image, such as a single image showing both the recommended hair style and configuration and the hair appearance and finer style.

[0082] In one embodiment, application 1120A provides an interface for receiving user-provided reference images I2 and I3. For example, the user can place (or generate via a camera) an example of a hairstyle on smartphone 1108 and store it. The user can upload the reference images (collectively 1128) to website 1104 for use in generating a try-on of the hairstyle represented in result image 1126.

[0083] Continuing to refer to FIG. 11, in one example, a user of the tablet 1110 uses a browser to visit a website provided by the application delivery computing device 1106. The tablet 1110 receives an application 1120B that provides access to the network framework 1116. The application 1120B is configured in one example to be similar to the application 1120A. The user uses a camera to obtain a still image or a video image (1130) (e.g., a self - taken image), and this image is used as the image I1 and communicated for processing by the GAN generator in the cloud computing device 1105. In this embodiment, the user of the tablet 1110 also uploads images I2 and I3 (collectively 1132). The images 1132 can be recommended by the application 1120B or arranged by the user. The resulting image 1134 is communicated via the display device of the tablet 1110, displayed, stored in a storage device, and can be shared (communicated) via social media, text messages, email, etc.

[0084] The application 1120B is configured to provide the tablet 1110 with one or more interfaces to services for recommending and / or promoting the purchase of products and / or services that may be associated with a hairstyle in one embodiment.

[0085] In one example, the application 1120B is a photo - gallery application. The hairstyle effect is applied to a user image (an example of the image I1) from the participant's camera using the framework 1116. The application 1120B can facilitate the user's selection of the images I2 and I3. For example, from a data store related to the photo - gallery application (e.g., the storage device of the tablet 1110) or from the Internet or other data stores (e.g., via a recommendation service).

[0086] Therefore, in one embodiment, the network framework 1116 is configured to perform a hair style transfer to generate a composite image including the identity from the first image, the first hair style attribute from the second image, and at least one second hair style attribute from the third image. The network framework 1116 is configured for an editable hair style transfer, a) to provide a hair shape and structure relaxation feature, and b) to provide a hair appearance and finer style, thereby enabling the selection of hair attributes to be transferred. In one embodiment, the network framework performs a transfer for relaxing style attributes from each other using two-stage optimization. In an embodiment, the network framework generates a composite image (I G ), and in the first stage of optimization, the identity from the face of the first image (I1) is reconstructed into the face region of I G , and the hair shape and structure attributes from the hair region of the second image (I2) are reconstructed into the hair region of I G , respectively. Further, in the second stage, the network framework transfers each of the hair appearance attributes and finer style attributes from the hair region of the third image (I3) to the hair region of I G reconstructed in the first stage. In one embodiment, inpainting fills the background region, such as from the background of I1.

[0087] In one embodiment, the network framework is configured to perform gradient orthogonality in two-stage optimization to relax at least one style attribute represented by I2 and at least one style attribute represented by I3.

[0088] In one embodiment, when the style to be transferred is a hair style, at least one style attribute represented by I2 is a hair shape and structure attribute, and at least one style attribute represented by I3 is i) an appearance attribute and ii) a finer style attribute.

[0089] In one embodiment, the two-stage optimization is the identity reconstruction loss (L f) Shape and structural reconstruction loss of hair (L r ) Appearance loss (L a ) and optimize losses including finer style loss (L s ). In one embodiment, L f and L r are optimized in the first stage without optimizing L A and L s , and L f , L r , L a and L s are optimized in the second stage, and L r is optimized through gradient orthogonalization to avoid conflicts between the appearance and finer style attributes of I2 and those of I3.

[0090] Although the embodiments herein are mainly described with reference to the transfer of hair styles, for multiple style attributes, the transfer of other styles may be performed. According to one embodiment, a method for performing style transfer using artificial intelligence (AI) is provided, and the style includes multiple style attributes. The method includes processing a plurality of images including a first image (I1), a second image (I2), and a third image (I3) using an AI network framework having two-stage optimization for generating a synthetic image (I G ) including an identity represented by the first image (I1), a style determined from at least one style attribute represented by the second image (I2), and at least one style attribute represented by the third image (I3). In this manner, the network framework is configured to optimize the latent space of the GAN to perform style transfer while resolving at least one style attribute represented by I2 and at least one style attribute represented by I3. In one embodiment, I G includes an identity region, a style region, and a background region, and in the first stage, according to the objective function for optimizing the latent space, the network framework reconstructs the identity represented by I1 into the identity region of I G and at least one style attribute represented by I2 into IG configured to reconstruct into the style region of. In one embodiment, in the second stage according to the target function, the network framework transfers each of at least one style attribute represented by I3 into the style region of I G respectively configured to transfer respectively.

[0091] In embodiments such as hair style transfer, when the identity of I1 is unique among I1, I2, and I3, complete hair style transfer becomes possible. In embodiments such as hair style transfer, when the shape and structure of the hair of I2 are unique among I1, I2, and I3, hair style transfer related at least to the shape and structure becomes possible. In embodiments such as hair style transfer, when the appearance of the hair of I3 is unique among I1, I2, and I3, hair style transfer related at least to the appearance becomes possible. In embodiments such as hair style transfer, when the finer style of I3 is unique among I1, I2, and I3, hair style transfer related at least to the finer details of the hair becomes possible.

[0092] In one embodiment, the method uses a segmentation network to process I G , I1, I2, and I3 respectively, defines a respective hair (style) mask and face (identity) mask for each image, and defines respective target masks for transferring the style using one selected from such masks.

[0093] According to one embodiment, the GAN generator first generates I G as a mean image for receiving style transfer.

[0094] In one embodiment, it is reconstructed using high-level features extracted by processing I1 using a neural network encoder with pre-trained identity. In one embodiment in hair style transfer, it is reconstructed using features from a block after being generated by processing I2 using a neural network encoder with pre-trained hair shape and structure. In one embodiment in hair style transfer, the hair region of I2 is an eroded hair region that imposes a soft constraint on the placement of the target of the synthesized hair. In one embodiment in hair style transfer, the hair appearance is transferred using the overall appearance determined from the features extracted in the first block by processing I3 using a neural network encoder with pre-trained hair appearance, and the overall appearance is determined regardless of spatial information. In one embodiment in hair style transfer, a style finer than the hair style is transferred according to a high-level feature map extracted by processing I3 using a neural network encoder with pre-trained hair appearance.

[0095] In one embodiment in hair style transfer, the hair appearance includes color, and the finer style includes finer details including either the style of the bundle and the shading change between hair strands.

[0096] In one embodiment, the method includes providing an interface to a (e-commerce) service for purchasing products and / or services related to style transfer.

[0097] In one embodiment, I G is provided for display within a graphical user interface for comparison with I1.

[0098] In one embodiment, the method includes providing an interface to a service configured to recommend products and / or services related to style transfer.

[0099] In one embodiment of hair style transfer, the method includes providing an interface for receiving I1, providing a storage of reference images showing respective style attributes such as hair shape and structure and hair style including hair appearance and finer styles, providing a selection interface for receiving an input for defining I2 from one of the reference images, and providing a selection interface for receiving an input for defining I3 from one of the reference images. In one embodiment (e.g., in hair style transfer), it includes providing an interface for receiving one or both of I2 and I3 from sources other than the storage of reference images.

[0100] In one embodiment, a network framework is configured to perform a virtual try-on of a hair style on an identity image (e.g., I1), and provide a plurality of reference images (e.g., I2 and I3) representing different hair style attributes for simulating the hair style on the identity, and a network framework is configured to perform an optimization to unwind different hair style attributes to provide realistic synthesized hair when incorporating the identity and the hair style into a composite image (e.g., I G ) representing the virtual try-on of the hair style, and a network framework is configured to provide the composite image for presentation. A computer device is provided that includes a network framework configured as described above. In one embodiment, a circuit is configured to provide an interface for purchasing hair or hair style products, services, or both, and an interface for generating recommendations for such products, services, or both. <Conclusion>

[0101] According to an embodiment, the introduction of LOHO, an optimization framework that performs hair style transfer on portrait images, takes a step in the direction of spatially dependent attribute manipulation using a pre-trained GAN. By manipulating the latent space of a representation model trained on a more general task such as face synthesis, it has been shown that developing an algorithm that approaches a specific synthesis task such as hair style transfer is effective in completing many downstream tasks without collecting a large training dataset. The GAN inversion approach can more effectively solve problems such as realistic hole filling than a feed-forward GAN pipeline with access to a large training dataset.

[0102] Practical implementations can include any or all of the features described herein. These and other aspects, features and various combinations can be represented as a method for performing functions, a device, a system, a means and other ways of combining the features described herein. Some embodiments have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the processes and techniques described herein. In addition, other steps can be provided or steps can be excluded from the described processes, and other components can be added to or removed from the described systems. Accordingly, other aspects are within the scope of the claims.

[0103] Throughout the description and claims of this specification, the words "comprise" and "contain" and their variations mean "including but not limited to", and are not intended to exclude other components, integers or steps. Throughout this specification, the singular form includes the plural unless the context requires otherwise. In particular, when an indefinite article is used, it should be understood that both the singular and the plural are intended unless the context requires otherwise.

[0104] It should be understood that features, integers, properties or groups described in connection with a particular aspect, embodiment or example of the invention are applicable to any other aspect, embodiment or example, except where incompatible therewith. All of the features disclosed herein (including any accompanying claims, abstract and drawings) and / or all of the steps of any method or process so disclosed may be combined in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive. The invention is not limited to the details of any of the foregoing examples or embodiments. The invention extends to any novel one or any novel combination of features disclosed herein (including any accompanying claims, abstract and drawings) or any novel one or any novel combination of steps of any method or process disclosed. References - incorporated herein by reference in their entirety. [1]Rameen Abdal, Yipeng Qin, and Peter Wonka. Image2stylegan:How to embed images into the stylegan latent space? In 2019 IEEE / CVF International Conference on Computer Vision (ICCV), 2019. [2] R. Abdal, Y. Qin, and P. Wonka. Image2stylegan++: How to edit the embedded images? In 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8293-8302, 2020. [3]Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale GAN training for high fidelity natural image synthesis. In International Conference on Learning Representations, 2019. [4]Adrian Bulat and Georgios Tzimiropoulos. How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks). In International Conference on Computer Vision, 2017. [5]Egor Burkov, Igor Pasechnik, Artur Grigorev, and Victor Lempitsky. Neural head reenactment with latent pose descriptors. In IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. [6]Menglei Chai, Linjie Luo, Kalyan Sunkavalli, Nathan Carr, Sunil Hadap, and Kun Zhou. High-quality hair modeling from a single portrait photo. ACM Transactions on Graphics, 34:1-10, 10 2015. [7]Menglei Chai, Lvdi Wang, Yanlin Weng, Xiaogang Jin, and Kun Zhou. Dynamic hair manipulation in images and videos. ACM Transactions on Graphics (TOG), 32, 07 2013. [8]Menglei Chai, Lvdi Wang, Yanlin Weng, Yizhou Yu, Baining Guo, and Kun Zhou. Single-view hair modeling for portrait manipulation. ACM Transactions on Graphics, 31, 07 2012. [9]Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A Efros. Everybody dance now. In IEEE International Conference on Computer Vision (ICCV), 2019.

[10] L. A. Gatys, A. S. Ecker, and M. Bethge. Image style transfer using convolutional neural networks. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2414-2423, 2016.

[11] Ke Gong, Yiming Gao, Xiaodan Liang, Xiaohui Shen, Meng Wang, and Liang Lin. Graphonomy: Universal human parsing via graph transfer learning. In CVPR, 2019.

[12] Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2, NIPS'14, page 2672-2680, 2014.

[13] Shanyan Guan, Ying Tai, Bingbing Ni, Feida Zhu, Feiyue Huang, and Xiaokang Yang. Collaborative learning for faster stylegan embedding, 2020.

[14] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30, pages 6626-6637. Curran Associates, Inc., 2017.

[15] P. Isola, J. Zhu, T. Zhou, and A. A. Efros. Image-to-image translation with conditional adversarial networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5967-5976, 2017.

[16] Youngjoo Jo and Jongyoul Park. Sc-fegan: Face editing generative adversarial network with user's sketch and color. In The IEEE International Conference on Computer Vision (ICCV), October 2019.

[17] Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In European Conference on Computer Vision, 2016.

[18] Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. In International Conference on Learning Representations, 2017.

[19] T. Karras, S. Laine, and T. Aila. A style-based generator architecture for generative adversarial networks. In 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019.

[20] Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of StyleGAN. In Proc. CVPR, 2020.

[21] Vladimir Kim, Ersin Yumer, and Hao Li. Real-time hair rendering using sequential adversarial networks. In European Conference on Computer Vision, 2018.

[22] Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015.

[23] Cheng-Han Lee, Ziwei Liu, Lingyun Wu, and Ping Luo. Maskgan: Towards diverse and interactive facial image manipulation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.

[24] J. Li, X. Liang, Y. Wei, T. Xu, J. Feng, and S. Yan. Perceptual generative adversarial networks for small object detection. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1951-1959, 2017.

[25] T. Park, M. Liu, T. Wang, and J. Zhu. Semantic image synthesis with spatially-adaptive normalization. In 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2332-2341, 2019.

[26] Stanislav Pidhorskyi, Donald A Adjeroh, and Gianfranco Doretto. Adversarial latent autoencoders. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), 2020. [to appear]

[27] Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or. Encoding in style: a stylegan encoder for image-to-image translation. arXiv preprint arXiv:2008.00951, 2020.

[28] Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations, 2015.

[29] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2818-2826, 2016.

[30] Zhentao Tan, Menglei Chai, Dongdong Chen, Jing Liao, Qi Chu, Lu Yuan, Sergey Tulyakov, and Nenghai Yu. Michigan: Multi-input-conditioned hair image generation for portrait editing. ACM Transactions on Graphics (TOG), 39(4):1-13, 2020.

[31] A. Tao, K. Sapra, and Bryan Catanzaro. Hierarchical multi-scale attention for semantic segmentation. ArXiv, abs / 2005.10821, 2020.

[32] T. Wang, M. Liu, J. Zhu, A. Tao, J. Kautz, and B. Catanzaro. High-resolution image synthesis and semantic manipulation with conditional gans. In 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 8798-8807, 2018.

[33] Ting-Chun Wang, Ming-Yu Liu, Andrew Tao, Guilin Liu, Jan Kautz, and Bryan Catanzaro. Few-shot video-to-video synthesis. In Advances in Neural Information Processing Systems (NeurIPS), 2019.

[34] Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, and Bryan Catanzaro. Video-to-video synthesis. In Conference on Neural Information Processing Systems (NeurIPS), 2018.

[35] Yanlin Weng, Lvdi Wang, Xiao Li, Menglei Chai, and Kun Zhou. Hair interpolation for portrait morphing. Computer Graphics Forum, 32, 10 2013.

[36] J. Yu, Z. Lin, J. Yang, X. Shen, X. Lu, and T. Huang. Freeform image inpainting with gated convolution. In 2019 IEEE / CVF International Conference on Computer Vision (ICCV), pages 4470-4479, 2019.

[37] Yuhui Yuan, Xilin Chen, and Jingdong Wang. Object-contextual representations for semantic segmentation. In Computer Vision - ECCV 2020, pages 173-190, 2020.

[38] E. Zakharov, A. Shysheya, E. Burkov, and V. Lempitsky. Few-shot adversarial learning of realistic neural talking head models. In 2019 IEEE / CVF International Conference on Computer Vision (ICCV), pages 9458-9467, 2019.

[39] Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018.

[40] J. Zhu, T. Park, P. Isola, and A. A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 2242-2251, 2017.

[41] Jiapeng Zhu, Yujun Shen, Deli Zhao, and Bolei Zhou. In-domain gan inversion for real image editing. In Proceedings of European Conference on Computer Vision (ECCV), 2020. <Others> <Means> The method of Technical Idea 1 is a method for performing style transfer, where the style includes a plurality of style attributes, and a plurality of images including a first image (I 1 ), a second image (I 2 ), and a third image (I 3 ) are processed using an artificial intelligence (AI) network framework comprising a generative adversarial network (GAN) generator and a two-stage optimization for generating a synthetic image (I 1 ) including the identity represented by the first image (I 2 ), the style determined from at least one style attribute represented by the second image (I 3 ), and at least one style attribute represented by the third image (I G ). The network framework is configured to optimize the latent space of the GAN generator to perform the style transfer while unraveling the at least one style attribute represented by I 2 and the at least one style attribute represented by I 3 . The method of Technical Idea 2 is, in the method described in Technical Idea 1, where I G includes an identity region, a style region, and a background region, and in a first stage according to an objective function for optimizing the latent space, the network framework reconstructs the identity represented by I 1 into the identity region of I G , and reconstructs the at least one style attribute represented by I 2 into the style region of I G . The method of Technical Idea 3 is, in the method described in Technical Idea 2, where in a second stage by the objective function, the network framework transfers at least one style attribute represented by I 3 to the style region of I G . The method of Technical Idea 4 is, in the method described in Technical Idea 3, where the network framework is configured to inpaint the background region following the style transfer. The method of Technical Idea 5 is, in the method described in any one of Technical Ideas 1 to 4, where the network framework is configured to perform gradient orthogonality in the two-stage optimization to unravel the at least one style attribute represented by I 2 and the at least one style attribute represented by I 3 . The method of Technical Idea 6 is, in the method described in any one of Technical Ideas 1 to 5, where the style is a hair style, and I 2 The at least one style attribute represented by is a hair shape and structure attribute, and the at least one style attribute represented by I3 is i) an appearance attribute and ii) a finer style attribute. The method of Technical Idea 7 is the method described in Technical Idea 6, wherein the GAN generator is defined from a pre-trained GAN configured for style transfer. The method of Technical Idea 8 is the method described in Technical Idea 6 or 7, wherein the two-stage optimization is the identity reconstruction loss (L f ), the loss of hair shape and structure reconstruction (L r ), the appearance loss (L a ), and the finer style loss (L s ) to optimize the loss including. The method of Technical Idea 9 is the method described in Technical Idea 8, wherein L f and L r are optimized in the first stage without optimizing L a and L s , L f 、L r 、L a and L s are optimized in the second stage, and L r is optimized via gradient orthogonality to avoid competition between the appearance and finer style attributes of I 2 and those attributes of I 3 . The method of Technical Idea 10 is a method of transferring a hairstyle to a synthetic image (I G ), generating a synthetic image (I G ) by a network framework equipped with a generative adversarial network (GAN) generator, the network being configured to perform two-stage optimization to optimize the latent space of the GAN, and in the first stage of the two-stage optimization, the network framework reconstructs the identity from the face of the first image (I 1 ) to the face region of I G , and the hair shape and structure attributes from the hair region of the second image (I 2 ) to the hair region of I G , respectively, and in the second stage of the two-stage optimization, the network framework transfers each of the hair appearance attributes and finer style attributes from the hair region of the third image (I 3 ) to the hair region of I G reconstructed in the first stage. The method of Technical Idea 11 is the method described in Technical Idea 10, wherein the GAN generator is defined from a pre-trained GAN for processing face images for style transfer. The method of Technical Idea 12 is the method described in Technical Idea 10 or 11, wherein the two-stage optimization performs optimization using an objective function composed of an identity reconstruction loss (L f ), a loss of hair shape and structure reconstruction (L r ), an appearance loss (L a ), and a finer style loss (L s ) at each stage. The method of Technical Idea 13 is, in the method described in Technical Idea 12, where L f and L r are optimized in the first stage without optimizing L a and L s , L f 、L r 、L a and L s are optimized in the second stage, and L ris optimized via gradient orthogonality to avoid conflicts between the appearance and finer style features of I 2 and those of I 3 . The method of Technical Idea 14 is, in the method described in any of Technical Ideas 10 to 13, where the background area of I G after the transfer of the hairstyle is preferably inpainted from the background area of I 1 . The method of Technical Idea 15 is, in the method described in any of the above Technical Ideas, where the network framework is configured to provide editable hairstyle transfer, a) hair shape and structure relaxation features, and b) hair appearance and finer style, thereby enabling the selection of hair attributes to be transferred. The method of Technical Idea 16 is, in the method described in Technical Idea 15, where the identity of I 1 is unique between I 1 、I 2 and I 3 , thereby performing a complete hairstyle transfer, and the hair shape and structure of I 2 is unique between I 1 、I 2 and I 3 , thereby performing at least a hairstyle transfer related to shape and structure, and the hair appearance of I 3 is unique between I 1 、I 2 and I 3 , thereby performing at least a hairstyle transfer related to appearance, and the finer style of the hair of I 3 is unique between I 1 、I 2 and I 3 , thereby performing at least a hairstyle transfer related to the finer details of the hair. The method of Technical Idea 17 is, in the method described in any of the above Technical Ideas, where each of I 1 ~I 3 is a portrait image, and I 2 and I 3 are reference images for the hairstyle attributes to be transferred to the identity represented by I 1 . The method of Technical Idea 18 is, in the method described in any of the above Technical Ideas, where I G 、I 1 、I 2 and I 3 each define a respective hair (style) mask and face (identity) mask for each image using a segmentation network, and using one of the selected such masks, define a respective target mask for transferring the style. The method of Technical Idea 19 is, in the method described in any of the above Technical Ideas, where the GAN generator uses I as the average image for receiving the style transfer G is generated first. The method of technical idea 20 is a method described in any of the above-described technical ideas, wherein the identity is I using a pre-trained neural network encoder 1 is reconstructed using high-level features extracted by processing. The method of technical idea 21 is the method described in technical idea 20, in hair style transfer, using the pre-trained neural network encoder I 2 The shape and structure of the hair are reconstructed using the features from the block after being generated by processing. The method of technical idea 22 is the method described in technical idea 21, in hair style transfer, I 2 's hair region is an eroded hair region that imposes a soft constraint on the target placement of the synthesized hair. The method of technical idea 23 is the method described in any of technical ideas 20 to 22, in hair style transfer, using the pre-trained neural network encoder I 3 The appearance of the hair is transferred using the overall appearance determined from the features extracted in the first block by processing, and the overall appearance is determined regardless of spatial information. The method of technical idea 24 is the method described in any of technical ideas 20 to 23, in hair style transfer, using the pre-trained neural network encoder I 3 A finer style is transferred according to the high-level feature map extracted by processing. The method of technical idea 25 is the method described in any of the above-described technical ideas, in hair style transfer, the appearance of the hair includes color, and the finer style includes finer details including either the style of the bundle and the shading change between hair strands. The method of technical idea 26 includes providing an interface to an e-commerce service for purchasing products and / or services related to style transfer. The method of technical idea 27 includes providing an interface to a service configured to recommend products and / or services related to style transfer. The method of technical idea 28 is the method described in any of the above-described technical ideas, wherein the I G is I 1 is provided for display within a graphical user interface for comparison with. The method of technical idea 29 is the method described in any of the above-described technical ideas, and includes: providing an interface for receiving I; providing storage of reference images indicating respective style attributes such as hair styles including hair shape and structure, hair appearance, and finer styles; providing a selection interface for receiving an input for defining I from one of the reference images; and providing a selection interface for receiving an input for defining I from one of the reference images. 1providing an interface for receiving I, providing storage of reference images indicating respective style attributes such as hair styles including hair shape and structure, hair appearance, and finer styles, providing a selection interface for receiving an input for defining I from one of the reference images, and providing a selection interface for receiving an input for defining I from one of the reference images. 2 providing a selection interface for receiving an input for defining I from one of the reference images 3 providing a selection interface for receiving an input for defining I from one of the reference images The method of technical idea 30 is the method described in the method of technical idea 29, and includes providing an interface for receiving one or both of I and I from outside the storage of the reference images. 2 I 3 I The computing device of technical idea 31 includes a processor and a storage device that stores computer-executable instructions that, when executed by the processor, execute the method described in any of the above claims. The computing device of technical idea 32 includes a processor and a storage device that stores computer-executable instructions, and includes a network framework configured to execute transfer of a hair style. The network framework includes a generative adversarial network (GAN) generator configured to generate a composite image (I) including hair attributes in which the identity from the face of the first image (I) is transferred from a reference image. The hair attributes include i) hair shape and structure, ii) hair appearance, and iii) finer hair styles. The network framework is configured to optimize a latent space to disentangle i) hair shape and structure, which are the hair attributes, from ii) hair appearance and iii) finer hair styles. 1 I G I The computing device of technical idea 33 is the computing device described in technical idea 32, where the reference image includes a second image (I) and a third image (I), I and I are each composed of portrait images, and the network framework uses the hair shape and structure extracted from I and the hair appearance and finer hair styles extracted from I. 2 I 3 I 1 、I 2 I 3 I 2 I 3 I The computing device of Technical Idea 34, in the computing device described in any of Technical Ideas 32 to 33, when the instructions are executed, causes the computing device to repaint the once-generated I 1 with the background of I G on it 。 The computing device of Technical Idea 35, in the computing device described in any of Technical Ideas 32 to 34, the GAN generator is trained using two-stage optimization and gradient orthogonality such that the optimization of the latent space enables unraveling the hair attributes. The computing device of Technical Idea 36, in the computing device described in any of Technical Ideas 32 to 35, the appearance of hair includes color, and the fine style of hair includes finer details including either the style of the bundle between hair strands and the shading change. The computing device of Technical Idea 37, in the computing device described in any of Technical Ideas 32 to 36, when the instructions are executed, causes the computing device to operate to provide an interface to an e-commerce service for purchasing products and / or services associated with the hairstyle. The computing device of Technical Idea 38, in the computing device described in any of Technical Ideas 32 to 37, when the instructions are executed, causes the computing device to operate to provide an interface to a service configured to recommend products and / or services related to the hairstyle to the computing device. The computing device of Technical Idea 39, in the computing device described in any of Technical Ideas 32 to 38, when the instructions are executed, causes the computing device to provide an interface to receive I 1 , provides storage of reference images indicating respective hair attributes, and operates to provide a selection interface to receive an input for selecting at least one reference image to define hair attributes for transfer of the hairstyle. The computing device of Technical Idea 40, in the computing device described in any of Technical Ideas 32 to 39, when the instruction is executed, operates to provide an interface for uploading the reference image to the computing device. The computing device of Technical Idea 41 includes a processing circuit. When the processing circuit operates, it provides a network framework for performing virtual hair style try-on on an identity image and a plurality of reference images representing different hair style attributes for simulating the hair style on the identity. The network framework performs an optimization for unraveling the different hair style attributes to provide realistic synthetic hair when incorporating the identity and hair style into a synthetic image representing the virtual hair style try-on, and is configured to provide the synthetic image for presentation. The computing device of Technical Idea 42, in the computing device described in Technical Idea 41, when the circuit operates, operates at least one of providing an interface for purchasing a product, a service, or both associated with the hair style, and providing an interface for generating a recommendation associated with the hair style. The method of Technical Idea 43 performs the try-on using a network framework configured to perform virtual hair style try-on on an identity image and a plurality of reference images representing different hair style attributes for simulating the hair style on the identity, and perform an optimization for unraveling the different hair style attributes to provide realistic synthetic hair when incorporating the identity and hair style into a synthetic image representing the virtual hair style try-on, and provides the synthetic image for presentation. The method of Technical Idea 44, in the method described in Technical Idea 43, includes at least one of providing an interface for purchasing a product, a service, or both associated with the hair style, and providing an interface for generating a recommendation associated with the hair style.

Claims

1. A method for performing style transfer using artificial intelligence (AI), wherein the style includes a plurality of style attributes, A plurality of images including a first image (I 1 ), a second image (I 2 ), and a third image (I 3 ), are processed using an AI network framework comprising a generative adversarial network (GAN) generator and two-stage optimization for generating a synthetic image (I G ) including the identity represented by the I1, the style determined from at least one of the style attributes represented by the I2, and at least one of the style attributes represented by the I3, The AI network framework is configured to optimize the latent space of the GAN model to perform the style transfer while unraveling at least one of the style attributes represented by the I 2 and at least one of the style attributes represented by the I 3 . A method characterized by that.

2. The I G includes an identity area, a style area, and a background area, In a first stage according to an objective function for optimizing the latent space, the AI network framework reconstructs the identity represented by the I 1 into the identity area of the I G , and reconstructs at least one of the style attributes represented by the I 2 into the style area of the I G . The method according to claim 1, characterized by that.

3. In a second stage by the objective function, the AI network framework 3 unravels each of at least one of the style attributes represented by the I GThe method according to claim 2, characterized in that it transfers to the style area described above.

4. The AI network framework is the I 2 At least one of the style attributes represented by and the I 3 The method according to any one of claims 1 to 3, characterized in that it is configured to perform gradient orthogonality in the two-stage optimization in order to unwind at least one of the style attributes represented by and at least one of the style attributes represented by.

5. The style is a hairstyle, The I 2 At least one of the style attributes represented by is a hair shape and structure attribute, The I 3 The method according to any one of claims 1 to 4, characterized in that at least one of the style attributes represented by is i) an appearance attribute and ii) a finer style attribute.

6. The AI network framework is configured to transfer editable hairstyles, a) unwind features of hair shape and structure, and b) provide the appearance and finer style of the hair, thereby enabling selection of hair attributes to transfer. The method according to claim 1.

7. The I 1 The identity of is unique among the I 1 The I 2 And the I 3 Thereby performing a complete hairstyle transfer, The I 2 The shape and structure of the hair of are unique among the I 1 The I 2 And the I 3 Thereby performing a transfer of a hairstyle related to at least shape and structure, The I 3 The appearance of the hair of is the I 1 The I 2 And the I 3unique between them, thereby performing the transfer of the hairstyle related at least to the appearance, said I 3 the finer style of the hair is the I 1 said I 2 unique between said I and said I3, thereby performing the transfer of the hairstyle related at least to the finer details of the hair. The method according to claim 6, characterized in that.

8. The I G said I 1 said I 2 and said I 3 each define a respective hair (style) mask and a face (identity) mask for each image using a segmentation network, and using one of the selected such masks, define a respective target mask for transferring the style. The method according to claim 1, characterized in that.

9. The GAN generator first generates the I as an average image for receiving the transfer of the style, G the identity is reconstructed using the high-level features extracted by processing the I using a pre-trained neural network encoder, 1 In hair style transfer, the shape and structure of the hair are reconstructed using the features from the blocks generated by processing the I using the pre-trained neural network encoder. The method according to claim 1, characterized in that. 2

10. In hair style transfer, the appearance of the hair includes color, and the finer style includes finer details including either the style of the bundle and the shading change between hair strands. The method according to claim 1, characterized in that. **Claim 11**: The method according to claim 1, characterized by including any one of providing an interface to an e-commerce service for purchasing a hair product or a hair styling product associated with the I2 and the I3 for generating the IG, a hair service or a hair styling service associated with the I2 and the I3 for generating the IG, or both; or providing an interface configured to recommend the hair product or the hair styling product, the hair service or the hair styling service, or both. **Claim 12** The I G is provided for display within a graphical user interface for comparison with the I 1 according to claim 1. **Claim 13**: Providing an interface for receiving the I 1 ; providing storage of reference images showing respective style attributes such as hair styles including hair shape and structure and hair appearance and finer styles; providing a selection interface for receiving an input for defining the I from one of the reference images; and providing a selection interface for receiving an input for defining the I from one of the reference images, the method according to claim 1. 2 The I 3 **Claim 14** A computing device comprising a processor and a storage device storing computer-readable instructions executed by the processor, wherein a hair style is transferred via an artificial intelligence (AI) network framework, and a synthetic image (I 1 ) in which hair attributes from a reference image different from the I1 are transferred to the identity from the face in the first image (I GIt is provided with a generative adversarial network (GAN) generator that generates The hair attributes include: i) the shape and structure of the hair, ii) the appearance of the hair, and iii) the finer style of the hair. The AI network framework is configured to optimize the latent space to disentangle the hair attributes, namely: i) the hair shape and structure, ii) the hair appearance, and iii) the finer style of the hair. A computing device characterized by this. **Claim 15** The reference image includes a second image (I 2 ) and a third image (I 3 ). The I 1 , the I 2 and the I 3 are each composed of portrait images. The AI network framework uses the shape and structure of the hair extracted from the I 2 and the appearance of the hair and the finer style of the hair extracted from the I 3 . The computing device according to claim 14, characterized by this. **Claim 16** The GAN generator is trained using two-stage optimization and gradient orthogonality so that the optimization of the latent space enables the disentangling of the hair attributes. The computing device according to claim 14 or 15, characterized by this. **Claim 17** When the computer-readable instructions are executed, the computing device is caused to Provide one or both of an interface to an e-commerce service for purchasing products and / or services associated with the hairstyle and an interface for recommending hair products or hairstyle products, hair services or hair styling services or both. The computing device according to any one of claims 14 to 16, characterized by this. **Claim 18** When the computer-readable instructions are executed, the computing device is caused to provide an interface for receiving the 1 , and provide storage of the reference images indicating the respective hair attributes, provide a selection interface for receiving an input for selecting at least one of the reference images to define the hair attributes for the transfer of the hairstyle. The computing device according to any one of claims 14 to 17, characterized in that it results in

19. A computing device comprising a circuit, When the circuit operates, the computing device is caused to provide a network framework for performing a virtual hairstyle try-on on an identity image and a plurality of reference images representing different hairstyle attributes for simulating a hairstyle on the identity, The network framework performs an optimization for unraveling the different hairstyle attributes to provide realistic synthesized hair when incorporating the identity and hairstyle into a composite image representing the virtual hairstyle try-on, A computing device characterized in that it results in providing the composite image for presentation.

20. When the circuit operates, the computing device is caused to provide at least one of: an interface for recommending a hair product or hairstyle product associated with the reference image used for the virtual hairstyle try-on, a hair service or hair styling service associated with the reference image used for the virtual hairstyle try-on, or both; and an interface for generating a recommendation for a hair or hairstyle. The computing device according to claim 19, characterized in that it results in

Citation Information

Patent Citations

  • Method of training generative adversarial network (GAN), method of generating images using GAN, and computer-readable storage medium

    JP2020191093A

  • Image fusion method, model training method, and related device

    WO2020173329A1