High-Resolution Controllable Facial Aging with Space-Aware Conditional GAN

The method addresses the limitations of existing facial aging techniques by using ethnic-specific aging information and weak spatial supervision in a patch-based training approach, resulting in high-resolution, realistic, and controllable facial aging results.

JP7690500B2Active Publication Date: 2025-06-10LOREAL SA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022580297
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-11
Filing Date
2021-06-29
Publication Date
2025-06-10
Estimated Expiration
2041-06-29

AI Technical Summary

Technical Problem

Existing approaches for facial aging in image processing yield poor-quality results with limited aging options, failing to capture the continuous and nuanced nature of aging, and often distort individual variations and expression wrinkles.

Method used

A computing device and method for controllably converting high-resolution face images by training a model with ethnic-specific aging information and weak spatial supervision, using an aging map to present ethnic-specific aging information, and employing patch-based training to minimize computational resources.

Benefits of technology

The solution achieves state-of-the-art high-resolution facial aging results, providing fine-grained control over the aging process, maintaining identity, and producing realistic images that accurately simulate continuous aging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007690500000007
    Figure 0007690500000007
  • Figure 0007690500000008
    Figure 0007690500000008
  • Figure 0007690500000009
    Figure 0007690500000009
Patent Text Reader

Abstract

Computing devices and methods are provided for controllably transforming facial images, including high-resolution images, to simulate continuous aging. Ethnicity-specific aging information and weak spatial supervision are used to guide the defined aging process by training a model including a GAN-based generator. An aging map presents the ethnicity-specific aging information as a skin sign score or apparent aging value. The scores are located within the map in relation to the location of each of the facial skin sign zones associated with the skin sign. In particular, patch-based training, which involves location information to distinguish similar patches from different parts of the face, is used to train on high-resolution images while minimizing resource usage.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference

[0001] This application claims priority and / or benefit to U.S. Provisional Application No. 63 / 046,011, filed Jun. 30, 2020, entitled "High-Resolution Controllable Face Aging with Spatially-Aware Conditional GANs", and French Patent Application No. 2009199, filed Sep. 11, 2020, entitled "High-Resolution Controllable Face Aging with Spatially-Aware Conditional GANs", the entire contents of each of which are hereby incorporated by reference herein if and to the extent permitted.

Technical Field

[0002] The present disclosure relates to image processing, and more particularly, to high-resolution controllable face aging using spatially-aware conditional generative adversarial networks (GANs).

Background Art

[0003] Facial aging is an image synthesis task that requires transforming a reference image to give the impression of a person of a different age while maintaining the identity and important facial features of the subject. When done correctly, this process can be used in a variety of areas, from predicting the future appearance of missing persons to entertainment and educational uses. Focusing on achieving high-resolution facial aging can be a useful step for capturing fine details of aging (fine lines, pigmentation, etc.). In recent years, GANs

[14] have enabled a learning-based approach for this task. However, the results are often of poor quality and provide only limited aging options. General models such as StarGAN

[10] cannot produce convincing results without additional fine-tuning and modification. This is due in part to the choice of reducing aging to the true or apparent age [1]. Also, current approaches treat aging as a stepwise process and divide aging in bins (30 - 40, 40 - 50, 50+ etc.) [2, 16, 28, 30, 32].

[0004] In reality, aging is a continuous process that can take many forms depending on genetic factors such as facial features and ethnicity, as well as lifestyle choices (smoking, hydration, sun damage, etc.) or behavior. In particular, expression wrinkles can be promoted by habitual facial expressions and may be prominent on the forehead, upper lip, or corners of the eyes (crow's feet). Furthermore, aging is subjective as it depends on the cultural background of the person assessing the age. These factors require a more nuanced approach to dealing with aging.

[0005] Existing approaches and datasets for facial aging result in distorted results towards the average, and individual variations and expression wrinkles are often not favorably seen or overlooked in the overall pattern such as facial hypertrophy. Furthermore, they provide little or no control over the aging process and are difficult to scale to large images, thus hindering their use in many real-world applications.

Summary of the Invention

[0006] According to the technical method of this specification, respective embodiments for a computing device and a method, etc. for controllably converting a face image including a high-resolution image and simulating continuous aging are provided. In one embodiment, an aging process defined by training a model with ethnic-specific aging information and weak spatial supervision is induced. In one embodiment, an aging map presents ethnic-specific aging information as a skin science score or an apparent aging value. In one embodiment, the scores are arranged in the map in relation to the respective positions of the face's skin sign zones related to skin signs. In one embodiment, patch-based training is used to train high-resolution images while minimizing the use of computing resources, especially in relation to position information for distinguishing similar patches from different parts of the face.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

[0008] Drawings containing face images are masked for the purpose of presentation in this disclosure and are not masked during actual use.

[0009]

[0010]

[0011]

[0012]

[0013]

[0014]

[0015]

[0016]

[0017]

[0018]

[0019]

[0020]

DETAILED DESCRIPTION OF THE INVENTION

[0021] According to the technical approach of this specification, each embodiment is a system and method aimed at obtaining a high-resolution face aging result by creating a model that can individually convert local aging signs. FIG. 1 is an array 100 of high-resolution faces showing two faces in each of the continuously processed rows 102 and 104 according to one embodiment.

[0022] In one embodiment, a curated high-resolution dataset is used in connection with novel techniques (combinations) to generate detailed state-of-the-art aging results. Clinical aging signs and weak spatial supervision enable fine-grained control over the aging process.

[0023] In one embodiment, a patch-based approach is introduced to enable inference on high-resolution images while keeping the computational cost of training the model low. As a result, the model can provide state-of-the-art aging results at a scale four times larger than conventional methods. Related research

[0024] Conditional Generative Adversarial Networks (Conditional GANs)

[14] leverage the principle of adversarial loss to enforce that samples generated by a generative model are indistinguishable from real samples. This approach has yielded impressive results, particularly in the area of image generation. GANs can be extended to generate images based on one or several conditions. The resulting conditional GANs are trained to generate images that satisfy both realism and the conditional criteria.

[0025] The conditional GAN for image-to-image translation from unpaired images is a powerful tool for the image-to-image translation

[18] task, where an input image is given to the model to synthesize a translated image. StarGAN

[10] introduced an approach that uses additional conditions to specify the desired translation to be applied. They proposed feeding the input conditions to the generator in the form of a feature map

[10] concatenated to the input image, but the new approach uses more complex mechanisms such as AdaIN

[20] or its 2D extension SPADE

[22] to condition the generator in a more optimal way. While previous techniques require pixel-aligned training images in different domains, recent studies such as CycleGAN

[34] and StarGAN

[10] introduced cycle-consistency loss to enable unpaired training between discrete domains. This has been extended to enable translation between continuous domains

[23] . Facial aging

[0026] To age a face from one image, conventional approaches use training data of either one [2, 16, 30, 32, 33] or multiple images [26, 28] of the same person along with the age of the person when the image was taken. Using longitudinal data with multiple images of the same person introduces significant time-dependent constraints on dataset collection and is less flexible.

[0027] Age is typically classified (e.g., grouped) into separate age groups (20 - 30, 30 - 40, 40 - 50, 50+ etc.) [2, 16, 30, 32], which simplifies framing the problem but limits control over the aging process and does not allow training to leverage the ordered nature of the groups. The disclosure of

[33] addresses this limitation by considering age as a continuous value. Aging is not aimed at different aging of different skin types, and different populations look for different signs of aging. Focusing on apparent age as a guide for aging freezes the subjective perspective. Such an approach cannot align with the population's perspective without additional age-estimation data from that perspective.

[0028] To improve the quality and level of detail of the generated images,

[32] uses an attention mechanism from the generator of

[23] . However, the generated samples are low-resolution images that are too coarse for real-world applications. Working at this scale hides several difficulties in generating realistic images, such as skin texture, fine lines, and overall sharpness of details. Approach Problem statement

[0029] In one embodiment, the goal is to train a model that can generate a realistic high-definition (e.g., 1024×1024) aged face by continuously controlling fine-grained aging signs using a single unpaired image to generate a smooth transformation between the original image and the transformed image. This is a more intuitive approach since aging is a continuous process and aging group bins do not explicitly enforce a logical order.

[0030] In one embodiment, the use of ethnic-specific skin atlases [4-7,13] incorporates the ethnic aspects of clinical aging signs. These atlases define a number of clinical signs such as wrinkles under the eyes, eyelid drooping in the lower part of the face, and the density of pigmented spots on the cheeks. Each sign is linked to a specific zone on the face and scored on an ethnicity-dependent scale. Using these labels in addition to age enables a more complete representation of aging and allows images to be transformed using various combinations of clinical signs and scores.

[0031] In one embodiment, FIGS. 2A-2D are images showing each of the aging sign zones (a)-(d) (202, 204, 206, 208) of the face 212 shown in FIG. 2E. Other sign zones are used but not shown. FIG. 2E also shows the aging map 210 of the face 212. The aging map 210 is constructed, according to one embodiment, from the relevant aging sign scores for all zones of the face 212. The zones (a)-(d) shown in the figure will be understood. FIGS. 2A-2D are shown enlarged with respect to the face 212 of FIG. 2E. In one embodiment, the skin signs represent "age", "forehead wrinkles", "nasolabial fold", "wrinkles under the eyes", "inter-brow wrinkles", "inter-ocular wrinkles", "wrinkles at the corners of the lips", "upper lip", and "lower face drooping". In one embodiment, other skin signs are used where there is sufficient training data etc.

[0032] In the aging map 210, the luminance of each pixel represents the normalized score of the local clinical sign (e.g., under the corner of the lip (a), under the eye wrinkles (b), nasolabial fold wrinkles (c), inter-ocular wrinkles (d), etc.). If the aging sign score is not available (not defined), an apparent aging value is used.

[0033] In other words, in one embodiment, the aging target is passed to the network in the form of an aging map (e.g., 210) of a specific face image (e.g., 212). To do so, face landmarks are calculated and relevant zones for each aging sign (see FIGS. 2A-2D) are defined. Each zone (e.g., the forehead, not shown as a zone in FIGS. 2A-2D) is then filled with the score value of the corresponding sign (e.g., forehead wrinkles). In FIGS. 2A-2D of this embodiment, the skin aging sign values for the applicable regions are (a): 0.11, (b): 0.36, (c): 0.31, and (d): 0.40. In one embodiment, the apparent age is used (via an estimator) or the actual age if available to fill in the blanks where clinical signs are not defined. Finally, a coarse mask is applied to the background of the image.

[0034] In one embodiment, the skin aging sign values (and the apparent age if used) are normalized on a scale of 0 to 1.

[0035] It is ideal to process the entire image at once, but training the model using 1024×1024 images requires significant computational resources. In one embodiment, a patch-based training approach is used, where only a part of the image is used to train the model during training, and the corresponding patch portion of the aging map is trained. Patch-based training not only reduces the context of the task (i.e., global information), but also reduces the computational resources required to process high-resolution images in large batches, as recommended in [8]. For small patches of 128×128, 256×256, or 512×512 pixels, a large batch size is used. In one embodiment, the training samples random patches each time the image is seen in the training process (about 300 times in such training).

[0036] The main drawback of patch-based training is that small patches that may look the same (e.g., forehead and cheek) must age differently (e.g., horizontal and vertical wrinkles respectively). Referring to FIGS. 3A and 3B, in one embodiment, to avoid wrinkles determined from the arithmetic mean on these ambiguous zones, the generator comprises two patches from each of a horizontal gradient location map 300 and a vertical gradient location map 302. Wrinkles by arithmetic mean are not natural in appearance. This allows the model to know the position of the patches in order to potentially distinguish ambiguous zones. Network architecture

[0037] In one embodiment, the training process is based on the StarGAN

[10] framework. The generator is a fully convolutional encoder-decoder derived from

[11] which has SPADE

[22] residual blocks in the decoder to incorporate an aging map and a location map. This allows the model to utilize the spatial information present in the aging map and use it at multiple scales within the decoder. To avoid learning unnecessary details, an attention mechanism from

[23] is used to cause the generator to transform the image only when necessary. The discriminator is a modified version of

[10] and produces outputs for the WGAN [3] objective (given for image i and aging map a in Equation 1), the estimation of the patch coordinates, and the low-resolution estimation of the aging map a.

Number

[0038] In one embodiment, FIGS. 4 and 5 respectively show patch-based training workflows 400 and 500, FIG. 4 shows the training of the generator (G) 402, and FIG. 5 shows the training of the discriminator (D) 502 of the GAN-based model.

[0039] Referring to FIG. 4, the generator (G) 402 includes an encoder section 402A and a decoder section 402B. The decoder section 402B is configured using SPADE residual blocks and accommodates the map and its position again. The workflow operation 400 crops patches from each of the image I (404), the aging map A (406), and the position maps X and Y (408, 410) to obtain the image patch I p (412), the aging map patch A p (414), and the position map X p and Y p (416, 418). The generator 402 converts the image patch I p 412 according to the map 414 and the positions (maps 416, 418) via the SPADE configuration 420 to generate the image Δp422. As described above, the patch size can be 128×128, 256×256, or 512×512 pixels for training a 1024×1024 image.

[0040] The attention mechanism 424 of

[23] is used to convert the image (patch 412) only when it is necessary to give the result G(I p │A p )426 to the generator 402.

[0041] Referring to FIG. 5 and the workflow operation 500, the discriminator (D) 502 generates a real / fake output 504, an estimated position (x, y) 506 of the patch, and an estimated aging map (508). These outputs (504, 506, and 508) are penalized by the WGAN objective, position, and aging map loss functions (510, 512, and 516), respectively. The position and the re-map loss functions will be further described.

[0042] The result 426 is the result G(G(I p │A p))518 is used in cycle GAN-based model training to generate. The cycle consistency loss 520 ensures that the transformation maintains the main features of the original image patch 412. Aging map

[0043] In one embodiment, to avoid imposing a penalty on the model (e.g., the generator) for failing to place the bounding boxes with pixel-precision, the aging map is blurred to calculate the discriminator regression loss on the downsampled 10×10 map with smoothed edges. This formulation allows for packing information in a more compact and meaningful way than individual uniform feature maps [10,28,32,33]. This approach requires multiple feature maps only when there is a large overlap between signs (e.g., forehead pigmentation and forehead wrinkles). In one embodiment, the common case of small overlap is to have only one aging map with the average value of the two signs within the overlap zone. If the zones overlap too much (e.g., forehead wrinkles VS forehead pigmentation), in one embodiment, the aging map includes a two-layer aging map (i.e., one aging map for wrinkles and, in this case, one aging map for pigmentation).

[0044] Considering the image patch i and the aging map patch a, the loss is given by Equation 2.

Equation

[0045] In one embodiment, two orthogonal gradients (position maps 416, 418) are used to assist the generator 402 in applying the aging transformation associated with a given patch (e.g., 412). The X, Y coordinates of patch 412 can be given to the generator 402 as two numbers instead of a linear gradient map, but doing so breaks the full convolution property and thus hinders the use of the model on full-scale images. Considering an image patch i located at coordinates (x, y) and an aging map patch a, the loss is given by Equation 3.

Number

[0046] In one embodiment, the model is β 1 = 0, β 2 = 0.99, a learning rate of 7×10 -5 for G, and an Adam

[21] optimizer of 2×10 -4 for D. According to the two-time-scale update rule

[17] , both models are updated at each step. Additionally, the learning rate is linearly decayed to zero for both G and D over the course of training. To enforce cycle-consistency, the perceptual loss of

[31] is used with λ Cyc = 100. In the regression task, λ Loc = 50 is used to predict the (x, y) coordinates of the patch, and λ Age = 100 is used to estimate the downsampled aging map. The discriminator is penalized with the original gradient penalty shown in

[15] with λ GP = 10. The complete loss objective function is given by Equation 4:

Number

[0047] For inference, in one embodiment, a trained (generator) model G can be optimized for stability, such as by determining an exponential moving average

[29] over the parameters of G to define an inference model G. The trained generator can be used directly on a 1024×1024 image, regardless of the size of the patches used during training, for the complete convolutional nature of the network and the use of continuous 2D aging maps.

[0048] In one embodiment, the target aging map is created manually. In one embodiment, face landmarks and target scores are used to construct the target aging map.

[0049] In one embodiment, the application is configured to facilitate a user entering target aging into an application interface, and the application defines an aging map (and optionally a position map) with the target aging as an aging map value.

[0050] In one embodiment, instead of absolute age, it is easier for the user to enter an age difference (e.g., a delta value such as subtracting 3 years or adding 10 years). In an embodiment, the application then analyzes the received image to determine an apparent age or skin signature value, and then defines an aging map for that analysis to modify the apparent age / skin signature value to conform to the user request. The application is configured to use that map to define a modified image showing the aged image.

[0051] In one embodiment, a method (e.g., a computing device method) is configured as follows:

[0052] Receive a user-provided "selfie" image;

[0053] Analyze an image to generate a "current" skin signature value; automated skin signature analysis is shown and described in U.S. Patent Publication No. 2020 / 0170564 A1, published on June 4, 2020, titled "Automated Image-Based Diagnosis Using Deep Learning," the entire content of which is incorporated herein by reference;

[0054] Present (via a display device) to the user an annotated selfie showing the user's analyzed skin signatures overlaid on the face zones associated with each signature;

[0055] Receive (via a graphical or other user interface) user input to adjust one or more signature scores. By way of example, the input is a skin signature adjustment value (e.g., a target or a delta). By way of example, the input is a product and / or service selection associated with a zone (or two or more). The product and / or service is associated with a skin signature score adjustment value (e.g., a delta).

[0056] Define an aging map using the current skin signature score and the skin signature score adjustment value;

[0057] Use the map in a generator G to define a modified image; and

[0058] Present (e.g., via a display device) to the user a modified image showing how the user would look after using the product and / or service, by way of example; Experiment Experimental setup

[0059] Most face aging datasets [9, 24, 25] suffer from a lack of diversity with respect to ethnicity

[19] and focus on low-resolution images (up to 250×250 pixels). This is not sufficient to capture the details related to skin aging. Furthermore, they often cannot normalize the pose and expression of the face (smiling, frowning, raising eyebrows), resulting in prominent wrinkles unrelated to aging (mostly crow's feet wrinkles, nasolabial folds, forehead wrinkles, and under-eye wrinkles). Finally, the lack of fine-grained information regarding aging signs leads to other approaches capturing unwanted correlated features such as facial obesity, as observed in datasets such as IMDB-Wiki

[25] . These effects can be observed in Figure 6.

[0060] Figure 6 shows an array of images 600 including the original images in the first column 602 and the aged images in the remaining columns, showing a comparison between previous aging approaches and the approach of the current teachings herein. Images by previous aging approaches are presented in rows 604, 606, 608, and 610 in the order of

[28] ,

[16] ,

[26] , and [2], respectively. Images by the approach of the current teachings herein are presented in row 612.

[0061] Previous approaches operate on low-resolution images and suffer from a lack of dynamic range of wrinkles, especially for expression wrinkles (column 604). They are also prone to unwanted correlated features such as color shifts and artifacts (606, 608, and 610) and facial hypertrophy (610).

[0062] To address these issues, the model according to the present teachings was tested on two carefully selected high-resolution datasets using manually generated aging maps or uniform aging maps, emphasizing rejuvenation / aging. FFHQ

[0063] The experiments were conducted using the FFHQ dataset

[20] . In one embodiment, a simple heuristic was applied to select a subset of the better-quality dataset in order to minimize problems in illumination, pose, and facial expression. To do this, facial landmarks were extracted from all faces and used to remove all images where the head was tilted excessively left, right, up, or down. Additionally, to limit artificial "crow's feet", all images where the mouth was open and under the eye wrinkles were removed. Finally, HOG

[12] feature descriptors were used to remove images where the hair covered the face. This selection reduced the dataset from over 70,000 (70k+) images to over 10,000 (10k+) images. Due to the extreme diversity of the FFHQ dataset, the remaining images are still not perfect, especially with regard to illumination color, direction, and exposure.

[0064] To obtain the scores of individual aging signs on these images, in one embodiment, an aging sign estimation model based on the ResNet

[27] architecture trained on the high-quality standardized dataset described below (i.e., 6000 high-resolution 3000×3000 images) was used. Finally, a ground truth aging map was generated using landmarks as a basis for rough boundary lines. The model was trained on 256×256 patches randomly selected on a 1024×1024 face. High-quality standardized dataset

[0065] To obtain better performance, in one embodiment, a dataset of 6000 high-resolution (3000×3000) images of centered faces was collected across most ages, genders, and ethnicities (African, Caucasian, Chinese, Japanese, and Indian). The images were labeled using ethnicity-specific skin atlases [4 - 7, 13] and scored for signs that cover most of the face (apparent age, forehead wrinkles, crow's feet, under-eye wrinkles, upper lip wrinkles, lip corner wrinkles, and lower eyelid drooping of the face). Results FFHQ dataset

[0066] Despite the complexity of the dataset and without ground truth aging values, the patch-based model can continuously transform individual wrinkles on the face.

[0067] Figure 7 is an array 700 of images showing original (column 702), rejuvenated (column 704), and aging (column 706) images of six faces of different ages and ethnicities from the FFHQ dataset using embodiments of the present teachings herein. Figure 7 shows how the model was able to transform different wrinkles despite the complexity of patch-based training, large variations in lighting in the dataset, and the imbalance between clinical signs / grades of age, with the majority of young subjects having few wrinkles. Figure 8 is an array of images 800 showing model results in group 802 where skin sign values are not defined and group 804 where skin sign values are defined according to one embodiment. When the signs are not defined, the map is filled with aging values. This helps the model learn global features such as graying of hair (group 802). By using individual clinical signs in the aging map, all signs can be aged, but maintaining the appearance of the hair intact (group 804) emphasizes the contrast the model has for individual signs and allows the face to be aged in a controllable way that would not be possible with the sole label of age. High-quality standardized dataset

[0068] In more standardized images, with better coverage across ethnicity and aging, this model shows state-of-the-art performance (Figs. 1, 9), demonstrating no details, realism, or visible artifacts. As an example, Fig. 9 is an array 900 of images that continuously shows the aging over time of four faces in rows 902, 904, 906, and 908, according to one embodiment. Zones such as sagging in the forehead and lower face are not altered. Supplementary age information used to fill in the gaps can be seen in thinning or graying of the eyebrows.

[0069] The aging process using the teachings of this specification succeeds along the continuous spectrum of the aging map, enabling the generation of realistic images for diverse sets of significant sign severity values. As shown in the examples of Figs. 10A - 10F, in one embodiment, this realistic and continuous aging using defined aging maps is shown on the same face. Fig. 10A shows an image 1002 of a face before aging is applied. Fig. 10B shows an image 1004 of the face aged via the aging map, which rejuvenates all signs except laugh lines, the corners of the lips, and the wrinkles under the eye on the right part of the face. Fig. 10C shows an image 1006 where the map ages only the lower part of the face, and Fig. 10D shows an image 1008 where the map ages only the upper part of the face. Fig. 10E shows image 1010, where the map is defined to age only the wrinkles under the eyes. Fig. 10F shows an image 1012 of a map defined to age the face asymmetrically, i.e., the right wrinkles under the eyes and the left laugh lines. Evaluation metric

[0070] To be considered successful, the task of facial aging needs to meet three criteria: the image must be realistic, the identity of the subject must be maintained, and the face must be aged. These are achieved during training, thanks to the WGAN objective function, the cycle consistency loss, and the aging map estimation loss, respectively. Essentially, no single metric could guarantee that all criteria were met. For example, the model could leave the input image unchanged without altering it and still succeed in terms of realism and identity. Conversely, the model could succeed in aging but fail in terms of realism and / or identity. If one model is not better than another in all metrics, a trade-off can be selected.

[0071] Experiments on FFHQ and high-quality standardized datasets showed no problems in maintaining the identity of the subjects. In one embodiment, it was chosen to focus on the realism and aging criteria for quantitative evaluation. Since the approach herein does not depend only on age but focuses on age as a combination of age labels, the accuracy of the target age is not used as a metric. Instead, the Frechet Inception Distance (FID)

[17] is used to evaluate the realism of the images, and the Mean Average Error (MAE) is used to evaluate the accuracy of the target aging signs.

[0072] To do so, half of the dataset is used as a reference for real images, and the rest is used as the images to be converted by the model. The aging maps used to convert these images are randomly selected from the ground truth labels to ensure the distribution of the images generated according to the original dataset. The values of the individual scores were estimated for all the generated images using a dedicated aging signature estimation model based on the ResNet

[27] architecture. As a criterion for the FID score, the FID is calculated between the two halves of the real image dataset. Note that the size of the dataset prevents the calculation of the FID for the recommended size of 50,000 or more (50k+)[17,20], and thus leads to an overestimation of the value. This can be seen when only calculating the FID between real images, and a baseline FID of 49.0 is given. The results are shown in Table 1.

Table 1

[0073] In one embodiment, when trained without clinical signs, creating a uniform aging map using only age still gives the model satisfactory results, and the FID and MAE with respect to the reference of the estimated age are low. Therefore, Table 2 shows the Frechet Inception Distance and the mean average error for models with clinical signs and with age only.

[0074]

Table 2

[0075] However, when comparing an aged face with an aging-only approach, in the aging-only model, some wrinkles do not seem to fully exhibit their dynamics. This is due to the fact that in order to reach the limit age of the dataset, it is not necessary to maximize all aging signs. In fact, the oldest 150 people (65 - 80 years old) in the standardized dataset show a median standard deviation of 0.18 for the normalized aging signs, highlighting many possible combinations of aging signs in the elderly. This is a problem with the age-only model as it provides only one approach to aging the face. Signs such as forehead wrinkles, for example, are highly dependent on the subject's facial expression and are an essential part of the aging process. By looking only at the age of the subjects in the dataset, the distribution of these clinical aging signs cannot be controlled.

[0076] In contrast, in one embodiment, using an aging map, the aged face provides much more control over the aging process. By controlling the individual signs of aging, it is possible to select whether to apply these expression wrinkles or not. Along the natural extension of this effect, there is skin pigmentation that is seen as a sign of aging in some Asian countries. The age-based model cannot generate ages for these countries without the need to re-estimate age from a regional perspective. This is different from the approach disclosed herein that can provide a customized face aging experience tailored to the perspectives of different countries when trained using all relevant aging signs, where all are a single model and without additional labels. Ablation experiment

[0077] Effect of Patch Size: When training the model, in one embodiment, for a given target image resolution (1024×1024 pixels in the experiment), the size of the patches used for training can be selected. The larger the patch, the larger the context the model has to execute the aging task. However, for the same computing power, larger patches result in a smaller batch size, which hinders training [8]. Experiments were conducted using patches of 128×128, 256×256, and 512×512 pixels. FIG. 11 shows an array 1100 of images representing rejuvenation and aging results on a 1024×1024 image of a face according to the teachings of this specification. The array 1100 includes a first array 1102 of images for two respective faces and a second array 1104 of images. The array 1102 shows the rejuvenation results for the first face, and the array 1104 shows the aging results for the second face. Rows 1106, 1108, and 1110 show the results using different patch sizes. Row 1106 shows a 128×128 patch size, row 1108 shows a 256×256 patch size, and row 1110 shows a 512×512 patch size.

[0078] FIG. 11 shows that in one embodiment, all patch sizes can manage aging a high-resolution face with varying degrees of realism. The smallest patch size is most troubled by the lack of context and produces inferior results compared to the other two with visible texture artifacts. The 256×256 patch gives convincing results with only minor visible imperfections when compared to the 512×512 patch. These results suggest the application of this technology to larger resolutions, such as 512×512 patches on 2048×2048 images. Position Map:

[0079] To see the contribution of the position map, in one embodiment, the model was trained with and without them. As expected, the effect of the position map is more prominent for small patch sizes with high ambiguity. FIG. 12 shows that for a small patch size and without position information, the model cannot distinguish similar patches from different parts of the face. FIG. 12 shows an array 1200 of images showing the effect of aging over time in two arrays 1202 and 1204 according to two (patch-trained) models according to the teachings of this specification. In FIG. 12, the face aged with the smallest patch size without using the position map is shown in array 1202, and the face aged with the smallest patch size using the position map is shown in array 1204. In each array, the aged face is shown with the difference from the original image. When trained (patches) without using the position map, the model cannot add wrinkles consistent with the position and generates general diagonal ripples. This effect appears less for larger patch sizes because the position of the patch is not ambiguous. The position map eliminates the presence of diagonal texture artifacts and allows horizontal wrinkles to appear, especially on the forehead. Spatialization of Information:

[0080] The proposed use of the aging map according to the teachings of this specification was compared to a baseline approach for formatting conditions, i.e., all signature scores were given as individual uniform feature maps. Since not all signatures are present in a particular patch, especially when the patch size is small, most of the processed information is not useful to the model. The aging map represents an easy way to give the model the labels present in the patch, in addition to the spatial spread and position. FIG. 13 highlights the effect of the aging map. FIG. 13 shows an array 1300 of images showing the aging effect, where the first array 1302 shows the aging using a model (patch) trained using a uniform feature map, and the second array 1304 shows the aging using a model (patch) trained using the aging map according to the teachings of this specification.

[0081] For small or medium patches (e.g., 128×128 or 256×256 pixels), the model struggles to create realistic results. The aging map helps to reduce the complexity of the problem. Thus, FIG. 13 shows, in the array 1302 of three images and the array 1304 of three images, the faces aged at a large patch size (e.g., 512×512 having individual uniform condition feature maps (array 1302) and the proposed aging map (array 1304)), along with the difference from the original image in each respective array. The patch size does not need to be twice the size of the original image size (e.g., 800×800 is large even if it is not the full size of a 1024×1024 image). The aging map helps to make the training more efficient due to the more densely spatially distributed information and produces more realistic aging. This difference highlights the small unrealistic wrinkles for the baseline technique.

[0082] Alternatively, in one embodiment, a different approach is used as shown in StarGAN, whereby the model is given all signature values for each patch, even those that are not present within the patch. Application

[0083] In one embodiment, the disclosed techniques and methods include developer-related methods and systems for defining a model (e.g., through conditioning) having a generator for image-to-image conversion to provide age simulation. The generator exhibits continuous control (across multiple age-related skin signs) to produce a smooth conversion between an original image and a converted image (e.g., of a face). The generator is trained using unpaired individual training images, each of which has an aging map identifying facial landmarks associated with respective age-related skin signs, providing weak spatial supervision to induce the aging process. In one embodiment, the age-related skin signs represent ethnicity-specific dimensions of aging.

[0084] In one embodiment, a GAN-based model having a generator for image-to-image conversion for age simulation is incorporated into a computer-implemented method (e.g., an application) or a computing device or computing system to provide virtual reality, augmented reality, and / or modified reality experiences. The application is configured to facilitate a user to capture a self-captured image (or video) using a camera-equipped smartphone or tablet terminal, etc., and the generator G applies a desired effect for playback or other presentation by the smartphone or tablet terminal.

[0085] In one embodiment, the generator G taught in this specification is configured to be loaded and executed on a generally available consumer-oriented smartphone or tablet terminal (e.g., target device). Exemplary configurations include devices having the following hardware specifications: Intel® Xeon® CPU E5-2686v4@2.30GHz, profiled with only 1 core and 1 thread. In one embodiment, the generator G is loaded and configured to be executed on a computing device having more resources, including other devices having a server, desktop, gaming computer, or multiple cores and executing with multiple threads. In one embodiment, the generator G is provided as a (cloud-based) service.

[0086] Those skilled in the art will understand that in one embodiment, in addition to the developer (e.g., used during training time) and target (used during inference time) computing device aspects, a computer program product aspect is disclosed where instructions are stored in a non-transitory storage device (e.g., memory, CD-ROM, DVD-ROM, disk, etc.) to configure a computing device to execute any of the method aspects disclosed herein.

[0087] FIG. 14 is a block diagram of a computer system 1400 according to one embodiment. The computer system 1400 includes a plurality of computing devices (1402, 1406, 1408, 1410, and 1450) including a server, a developer computer (such as a PC or laptop), and typical user computers (such as PCs, laptops, smartphones, and tablet terminals, smaller form factor (personal) mobile devices, etc.). In an embodiment, the computing device 1402 provides a network model training environment 1412 including hardware and software to define a model for image-to-image conversion that provides continuous aging according to the teachings herein. The components of the network model training environment 1412 include model training components 1414 for defining and configuring a model including a generator G 1416 and a discriminator D 1418 by adjustment or the like. The generator G is useful for defining a model for use in inference to perform image-to-image conversion, while the discriminator 1418D is, as is well known, a configuration for training.

[0088] In this embodiment, the adjustment is performed according to the training workflows of FIGS. 4 and 5. The workflow uses patch training of high-resolution images (e.g., pixel resolution of 1024×1024 or higher). The training uses skin sign values or apparent age for each zone of the face where such skin signs are located. High-density spatialized information about these features is provided, such as by using an aging map. In this embodiment, the position of the patch is provided, for example, using position information, to avoid ambiguity and distinguish similar patches from different parts of the face. In this embodiment, position information is supplied using a gradient position map of (x, y) coordinates in the training image to achieve full convolution processing. In an embodiment, the model and discriminator have a form, provide an output, and are adjusted using the objective functions (e.g., loss functions) described above herein.

[0089] In this embodiment, when training uses patches, an aging map, and a position map, further components of the environment 1412 are an image patch (I p ) maker component 1420, an aging map (A p ) maker component 1422, and a position map (X p , Y p ) maker component 1424. Other components are not shown. In this embodiment, a data server (e.g., 1404) or other form of computing device stores an image dataset 1426 of (high-resolution) images for training and other purposes, and is coupled through one or more networks typically represented as network 1428, and network 1428 couples any of the computing devices 1402, 1404, 1406, 1408, and 1410. Network 1428 is, by way of example, wireless or otherwise, public or otherwise. It will also be understood that system 1400 is simplified. At least some of the services may be implemented by two or more computing devices.

[0090] Once trained, the generator 1416 may be further defined as desired and provided as an inference-time model (generator G IT ) 1430. According to the techniques and methods herein, in an embodiment, the inference-time model (generator G IT 1430) is made available for use in various ways. In one method in one embodiment as shown in FIG. 14, the generator G IT 1430 is provided as a service (SaaS) provided via the cloud server 1408, as cloud service 1432 or other software. A user application such as an augmented reality (AR) application 1434 is defined for use with cloud service 1432 that provides an interface to the generator G IT 1430. In one embodiment, the AR application 1434 is provided for delivery from an application delivery service 1436 provided by the server 1406 (e.g., via download).

[0091] Although not shown, in one embodiment, the AR application 1434 is developed using an application developer computing device for a specific target device having specific hardware and software, particularly an operating system configuration, etc. In one embodiment, the AR application 1434 is a native application configured for execution in a specific native environment, such as one defined for a specific operating system (and / or hardware). Native applications are often distributed via an application distribution service 1436 configured as an e-commerce "store" operated by a third-party service, but this is not necessary. In one embodiment, the AR application 1420 is a browser-based application configured to execute, for example, in the browser environment of the target user device.

[0092] The AR application 1434 is provided for delivery (e.g., download) by a user device such as the mobile device 1410. In one embodiment, the AR application 1434 is configured to provide an augmented reality experience (e.g., via an interface) to the user. For example, an effect is applied to the image by the processing of the estimated time generation unit 1430. The mobile device has a camera (not shown) for capturing an image (e.g., the captured image 1438), which is, in one embodiment, a still image including a self-captured image. Using an image processing technique that provides a transformation from image to image, an effect is applied to the captured image 1438. An aging image 1440 is defined and displayed on the display device (not shown) of the mobile device 1410 to simulate the effect on the captured image 1438. The position of the camera is further changed in response to the captured image(s) to simulate augmented reality, and an effect can be applied. It will be understood that the captured image defines the source or original image, and the aging image defines the transformed or transformed image or the image to which the effect is applied.

[0093] In the current cloud service paradigm of the present embodiment of FIG. 14, the captured image 1438 is provided to the cloud service 1432 and processed by the generator G IT 1430 to perform image-to-image conversion with continuous aging degradation to define the aging image 1440. The aging image 1440 is communicated to the mobile device 1440 for display, storage, sharing, etc.

[0094] In one embodiment, the AR application 1434 provides an interface (not shown) for operating the AR application 1434, for example, a graphical user interface (GUI) that can be voice-enabled. The interface is configured to enable image capture, communication with the cloud service, and display, storage, and / or sharing of the converted image (e.g., the aging image 1440). In one embodiment, the interface is configured to enable the user to provide input to the cloud service, for example, configured to define an aging map. As described above, in one embodiment, the input includes the target age. As described above, in one embodiment, the input includes an age delta. As described above, in one embodiment, the input includes a product / service selection.

[0095] In the embodiment of FIG. 14, an AR application 1434 or another application (not shown) provides access to a computing device 1450 that provides an e-commerce service 1452 (e.g., via communication). The e-commerce service 1452 includes a recommendation component 1454 for providing (personalized) recommendations for products, services, or both. In an embodiment, such products and / or services are anti-aging or anti-wrinkle products and / or services, etc. In an embodiment, such products and / or services are related to, for example, specific skin signs. The captured image from the device 1410 is provided to the e-commerce service 1452. Skin sign analysis is performed by a skin sign analyzer model 1456 or the like using deep learning according to one embodiment. Image processing using the trained model analyzes the skin (e.g., a facial zone related to specific skin signs) to generate a skin analysis that includes scores for at least some of the skin signs. The value of each score can be generated on the image using a (dedicated) aging sign estimation model (e.g., a type of classifier) based on the ResNet

[27] architecture as described above for analyzing the training set data.

[0096] In this embodiment, skin signs (e.g., their scores) are used to generate personalized recommendations. For example, each product (or service) is associated with one or more skin signs and a specific score (or range of scores) for such signs. In this embodiment, information is stored in a database (e.g., 1460) for use by the e-commerce service 1452 via a suitable lookup that matches the user's data to product and / or service data. In one embodiment, rule-based matching can be utilized to select one or more products and / or rank products / services related to a specific score (or range of scores) of such signs. In one embodiment, additional user data used by the recommendation component 1454 includes any of gender, ethnicity, and location data, etc. For example, location data is related to selecting any of a product / brand, formulation, regulatory requirements, format (e.g., size, etc.), labeling, SKU (stock keeping unit), and can be available or otherwise associated with the user's location. In one embodiment, any of such gender, ethnicity, and / or location data can also assist in selecting and / or ranking the selected products / services or filtering the products / services (e.g., removing products / services not sold at or for a location). In one embodiment, location data is used to determine retail merchants / service providers for which it is available (e.g., with or without a physical business location (e.g., store, salon, office, etc.)), such that the user can purchase products / services locally.

[0097] In this embodiment, the skin sign score of the user's captured image is provided by the e-commerce service for display via an AR application 1434 such as an AR application interface. In this embodiment, the skin sign score is used to generate the generator G ITDefine an aging map for providing to a cloud service 1432 used to define the converted image of 1430. For example, in an embodiment, the skin signature score generated by model 1456 is used as the one initially generated from the image to define the aging map values for some skin signatures. Other initially generated skin signature scores are modified to define the aging map values for some skin signatures. In this embodiment, for example, the user can modify some scores (e.g., only the skin signatures around the eyes) as generated via the interface. For example, in one embodiment, other means are used to modify the scores, such as by applying rules or other codes. In the embodiment, the modification is performed to represent the rejuvenation or aging or any combination of the selected skin signatures. Apparent aging values can be used for some skin signatures as described above instead of the skin signature scores.

[0098] In one embodiment, which is not limiting, the user receives personalized product recommendations as recommended by the e-commerce service 1452. The user selects a particular product or service. That selection triggers a modification of the user's skin signature score for the relevant skin signatures linked to the product or service. The modification adjusts the score to simulate the use of the product or service. The initially generated or modified skin signature scores are used in the aging map, provided to the cloud service 1432, and an aging image is received. As described above herein, the skin signature scores of different signatures may be combined within the map, and the generator G IT can transmit different signatures differently. Thus, in this embodiment, an aging map is defined, some skin signature scores are the ones initially generated for some signatures, and other signatures have modified scores.

[0099] In the embodiment of FIG. 14, an e-commerce service 1452 is configured using a purchase component 1458 to facilitate the purchase of products or services. The products or services include cosmetics or other services. Although not shown, the e-commerce service 1452 and / or the AR application 1434 provide image processing of the captured image to simulate cosmetics or services such as the application of makeup to the captured image to generate an image to which an effect is applied.

[0100] The captured image is used as a source image for processing in the above embodiment, but in one embodiment, other source images (e.g., from sources other than the camera of the device 1410) are used. Embodiments can use the captured image or other source images. In certain embodiments, whether a captured image or another image, a high-resolution image for improving the user experience as a model of the generator G IT 1430 is trained therefor. Although not shown, in this embodiment, the image used by the skin signature analyzer model is reduced when analyzed. For such analysis, other image preprocessing is performed.

[0101] In one embodiment, the AR application 1434 can instruct the user regarding quality features (i.e., lighting, centering, background, hair occlusion, etc.) to improve performance. In one embodiment, if the AR application 1434 does not meet certain minimum requirements and is inappropriate, the image is rejected.

[0102] Although shown as a mobile device in FIG. 14, in one embodiment, the computing device 1410 can have different form factors as described above. The generator G IT 1430 may be hosted and executed locally on a particular computing device having sufficient memory and processing resources, rather than (or in addition to) being provided as a cloud service.

[0103] Thus, in one embodiment, a computing device (e.g., device 1402, 1408, or 1410) is provided, the computing device comprising a processing unit configured to use an age simulation generator to receive an original image of a face and generate a transformed image for presentation, the generator being configured to continuously control a plurality of age-related skin signs between the original image and the transformed image of the face to simulate age, and the generator being configured to transform the original image using respective age targets for the skin signs. Such a computing device (e.g., device 1402, 1408, or 1410) is understood to be configured to execute related method aspects according to one embodiment, as described with reference to, for example, FIG. 15. It will be understood that embodiments of such computing device aspects have corresponding embodiments of the method aspects. Similarly, aspects of the computing device and method have corresponding aspects of a computer program product. The computer program aspects comprise a storage device (e.g., non-transitory) storing instructions that, when executed by a processor of the computing device, configure the computing device to execute a method as according to any respective embodiment herein.

[0104] In one embodiment, the generator is conditional GAN-based. In one embodiment, the target is provided to the generator as an aging map that identifies facial zones associated with respective skin signs, and each zone in the aging map is filled with a respective aging target corresponding to the associated skin sign. In one embodiment, the aging map represents a specific aging target of the associated skin sign by a score value of the associated skin sign. In one embodiment, the aging map represents a specific aging target of the associated skin sign by an apparent aging value of the associated skin sign. In one embodiment, the aging map represents a specific aging target for the associated skin sign by the score value of the associated skin sign when available and by the apparent aging value when the score value is not available. In one embodiment, the aging map is defined to use pixel intensity to represent the aging target.

[0105] In one embodiment, the aging map masks the background of the original image.

[0106] In one embodiment, the generator is configured through training using respective training images and associated aging maps, and the associated aging maps provide weak spatial supervision for inducing aging transformations of respective skin signs. In one embodiment, the skin signs represent ethnic-specific dimensions of aging. In one embodiment, the skin signs represent one or more of "age", "forehead wrinkles", "laugh lines", "under-eye wrinkles", "inter-brow wrinkles", "inter-ocular wrinkles", "lip corner wrinkles", "upper lip", and "lower face sagging".

[0107] In one embodiment, it is a fully convolutional encoder-decoder with residual blocks in the decoder so that the generator incorporates the aging target in the form of an aging map. In one embodiment, it is configured using patch-based training where the generator uses a portion of a specific training image and the corresponding patch of the associated aging map. In one embodiment, the residual blocks further incorporate position information to indicate the respective positions of portions of a specific training image and the corresponding patches of the associated aging map. In one embodiment, the position information is provided using respective X and Y coordinate maps defined from a horizontal gradient map and a vertical gradient map related to the height and width (H×W) size of the original image. In one embodiment, the specific training image is a high-resolution image and the patch size is a portion thereof. In one embodiment, the patch size is less than or equal to 1 / 2 of the high-resolution image.

[0108] In one embodiment, it is configured via an attention mechanism to limit the generator to transform age-related skin signs while minimizing additional transformations applied by the generator.

[0109] In one embodiment, a processing unit (e.g., of device 1410) is configured to communicate with a second computing device (e.g., 1408) that provides the generator for use, and the processing unit communicates the original image and receives the transformed image.

[0110] In one embodiment, the original image is a high-resolution image of 1024×1024 pixels or more.

[0111] In one embodiment, a processing unit (e.g., of computing device 1410) is further configured to provide an augmented reality application for simulating aging using the transformed image. In one embodiment, the computing device includes a camera and the processing unit receives the original image from the camera.

[0112] In one embodiment, the processing unit is configured to provide at least one of a recommendation function for recommending at least one of products and services and an e-commerce function for purchasing at least one of products and services. The act of "providing" in this context includes, in one embodiment, communicating with a web-based or other network-based service provided by another computing device (e.g., 1450) to facilitate recommendation and / or purchase.

[0113] In one embodiment, the product includes one of a rejuvenating product, an anti-aging product, and a cosmetic makeup product.

[0114] In one embodiment, the service includes one of a rejuvenating service, an anti-aging service, and a cosmetic service.

[0115] FIG. 15 is a flowchart of operation 1500 of a method aspect according to one embodiment, as may be executed by, for example, computing device 1402 or 1408. At step 1502, the operation receives an original image of a face, and at step 1504, uses an age simulation generator to generate a transformed image to be presented; the generator is configured to continuously control a plurality of age-related skin signs between the original image and the transformed image of the face to simulate age, and the generator is configured to transform the original image using respective age targets for the skin signs. As described above, embodiments of aspects of the related computing device have corresponding method embodiments.

[0116] In one embodiment, a computing device is provided that is configured to execute a method such as a method of configuring a network model training environment by adjusting an age simulation generator (based on a GAN). In one embodiment, the method defines an age simulation generator having continuous control over a plurality of age-related skin signs between an original image and a face transformation image, and training the generator using individual unpaired training images, each of which is associated with at least some of the age targets of the skin signs, and providing a generator for transforming the image.

[0117] In one embodiment, the generator is conditional GAN-based.

[0118] In one embodiment, the method includes defining an aging target as an aging map that identifies a face zone associated with each of the skin signs, and each zone in the aging map is satisfied with the respective aging target corresponding to the associated skin sign.

[0119] In one embodiment, a computing device is provided that includes a face effect unit including a processing circuit configured to apply at least one face effect to a source image and generate a virtual instance of the applied effect source image on an interface, the face effect unit utilizing a generator for continuously controlling a plurality of age-related skin signs between an original image and a face transformation image to simulate aging, the generator being configured to transform the original image using the respective aging target of the skin signs. In one embodiment, the interface is, for example, an e-commerce interface for enabling a purchase or product / service.

[0120] In one embodiment, a recommendation unit includes a processing circuit configured to present recommendations for products and / or services and receive a selection of a product and / or service, where the product and / or service is associated with at least one aging target modifier of a skin signature. In one embodiment, an interface is an e-commerce interface that enables, for example, the purchase of a recommended product / service. The face effect unit is configured to generate, in response to the selection, an aging target for each skin signature using the aging target modifier, thereby simulating the effect of the product and / or service on the source image. In one embodiment, the recommendation unit is configured to obtain a recommendation by calling a skin signature analyzer to determine a current skin signature score using the source image and using the current skin signature score to determine a product and / or service. In one embodiment, the skin signature analyzer is configured to analyze the source image using a deep learning model. In one embodiment, the aging target is defined from a current skin signature score and an aging target modifier. Conclusion

[0121] The present disclosure presents the use of clinical signs for creating an aging map for facial aging. Latest results regarding high-resolution images that fully control the aging process have been demonstrated. In one embodiment, a patch-based approach enables conditional GANs to be trained on large images while maintaining a large batch size.

[0122] Practical implementations can include any or all of the features described herein. These and other aspects, features, and various combinations can be represented as a method for performing functions, an apparatus, a system, a means, and other ways of combining the features described herein. Some embodiments have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the processes and techniques described herein. Additionally, other steps can be provided or steps can be excluded from the described processes, and other components can be added to or removed from the described systems. Accordingly, other aspects are within the scope of the claims.

[0123] Throughout the description and claims of this specification, the words "comprise" and "contain" and their variants mean "including but not limited to" and are not intended to exclude other components, integers, or steps. Throughout this specification, the singular form encompasses the plural form unless the context requires otherwise. In particular, when an indefinite article is used, it should be understood that both the singular and the plural are intended unless the context requires otherwise.

[0124] Features, integers, characteristics or groups described in connection with a particular aspect, embodiment or example of the present invention are to be understood as applicable to any other aspect, embodiment or example, unless they are incompatible therewith. All features (including any appended claims, abstract and drawings) disclosed herein and / or all steps of any method or process so disclosed may be combined in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive. The invention is not limited to the details of any foregoing example or embodiment. The invention extends to any novel one or any novel combination of features disclosed herein (including any appended claims, abstract and drawings) or any novel one or any novel combination of steps of any method or process so disclosed. References 1. Agustsson, E., Timofte, R., Escalera, S., Baro, X., Guyon, I., Rothe, R.: Apparent and real age estimation in still images with deep residual regressors on appareal database. In: 2017 12th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2017). pp. 87-94. IEEE (2017) 2. Antipov, G., Baccouche, M., Dugelay, J.L.: Face aging with conditional generative adversarial networks. In: 2017 IEEE international conference on image processing (ICIP). pp. 2089-2093. IEEE (2017) 3. Arjovsky, M., Chintala, S., Bottou, L. Wasserstein gan. arXiv preprint arXiv:1701.07875 (2017) 4. Bazin, R., Doublet, E.: Skin aging atlas. volume 1. caucasian type. MED'COM publishing (2007) 5. Bazin, R., Flament, F.: Skin aging atlas. volume 2, asian type (2010) 6. Bazin, R., Flament, F., Giron, F.: Skin aging atlas. volume 3. afro-american type. Paris: Med'com (2012) 7. Bazin, R., Flament, F., Rubert, V.: Skin aging atlas. volume 4, indian type (2015) 8. Brock, A., Donahue, J., Simonyan, K.: Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096 (2018) 9. Chen, B.C., Chen, C.S., Hsu, W.H.: Cross-age reference coding for age-invariant face recognition and retrieval. In: European conference on computer vision. pp. 768-783. Springer (2014) 10. Choi, Y., Choi, M., Kim, M., Ha, J.W., Kim, S., Choo, J.: Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 8789-8797 (2018) 11. Choi, Y., Uh, Y., Yoo, J., Ha, J.W.: Stargan v2: Diverse image synthesis for multiple domains. arXiv preprint arXiv:1912.01865 (2019) 12. Dalal, N., Triggs, B.: Histograms of oriented gradients for human detection. In: 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR'05). vol. 1, pp. 886-893. IEEE (2005) 13. Flament, F., Bazin, R., Qiu, H.: Skin aging atlas. volume 5, photo-aging face & body (2017) 14. Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial nets. In: Advances in neural information processing systems. pp. 2672-2680 (2014) 15. Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., Courville, A.C.: Improved training of wasserstein gans. In: Advances in neural information processing systems. pp. 5767-5777 (2017) 16. Heljakka, A., Solin, A., Kannala, J.: Recursive chaining of reversible image-to-image translators for face aging. In: International Conference on Advanced Concepts for Intelligent Vision Systems. pp. 309-320. Springer (2018) 17. Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S.: Gans trained by a two time-scale update rule converge to a local nash equilibrium. In: Advances in neural information processing systems. pp. 6626-6637 (2017) 18. Isola, P., Zhu, J.Y., Zhou, T., Efros, A.A.: Image-to-image translation with conditional adversarial networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 1125-1134 (2017) 19. Karkkainen, K., Joo, J.: Fairface: Face attribute dataset for balanced race, gender, and age. arXiv preprint arXiv:1908.04913 (2019) 20. Karras, T., Laine, S., Aila, T.: A style-based generator architecture for generative adversarial networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 4401-4410 (2019) 21. Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014) 22. Park, T., Liu, M.Y., Wang, T.C., Zhu, J.Y.: Semantic image synthesis with spatially-adaptive normalization. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 2337-2346 (2019) 23. Pumarola, A., Agudo, A., Martinez, A.M., Sanfeliu, A., Moreno-Noguer, F.: Ganimation: Anatomically-aware facial animation from a single image. In: Proceedings of the European Conference on Computer Vision (ECCV). pp. 818-833 (2018) 24. Ricanek, K., Tesafaye, T.: Morph: A longitudinal image database of normal adult age-progression. In: 7th International Conference on Automatic Face and Gesture Recognition (FGR06). pp. 341-345. IEEE (2006) 25. Rothe, R., Timofte, R., Van Gool, L.: Dex: Deep expectation of apparent age from a single image. In: Proceedings of the IEEE international conference on computer vision workshops. pp. 10-15 (2015) 26. Song, J., Zhang, J., Gao, L., Liu, X., Shen, H.T.: Dual conditional gans for face aging and rejuvenation. In: IJCAI. pp. 899-905 (2018) 27. Szegedy, C., Ioffe, S., Vanhoucke, V., Alemi, A.A.: Inception-v4, inception-resnet and the impact of residual connections on learning. In: Thirty-first AAAI conference on artificial intelligence (2017) 28. Wang, Z., Tang, X., Luo, W., Gao, S.: Face aging with identity-preserved conditional generative adversarial networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 7939-7947 (2018) 29. Yazici, Y., Foo, C.S., Winkler, S., Yap, K.H., Piliouras, G., Chandrasekhar, V.: The unusual effectiveness of averaging in gan training. arXiv preprint arXiv:1806.04498 (2018) 30. Zeng, H., Lai, H., Yin, J.: Controllable face aging. arXiv preprint arXiv:1912.09694 (2019) 31. Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The unreasonable effectiveness of deep features as a perceptual metric. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 586-595 (2018) 32. Zhu, H., Huang, Z., Shan, H., Zhang, J.: Look globally, age locally: Face aging with an attention mechanism. arXiv preprint arXiv:1910.12771 (2019) 33. Zhu, H., Zhou, Q., Zhang, J., Wang, J.Z.: Facial aging and rejuvenation by conditional multi-adversarial autoencoder with ordinal regression. arXiv preprint arXiv:1804.02740 (2018) 34. Zhu, J.Y., Park, T., Isola, P., Efros, A.A.: Unpaired image-to-image translation using cycle-consistent adversarial networks. In: Proceedings of the IEEE international conference on computer vision. pp. 2223-2232 (2017) <Others> <Means> The computing device of Technical Idea 1 includes a processing unit configured to receive an original image of a face and generate a converted image for presentation using an age simulation generator, where the generator continuously controls a plurality of age-related skin signs between the original image and the converted image to simulate aging, and further the generator is configured to transform the original image using each aging target of the skin signs. The computing device of Technical Idea 2 is the computing device described in Technical Idea 1, where the generator is based on conditional GAN. The computing device of Technical Idea 3 is the computing device described in Technical Idea 1 or 2, where the target is provided to the generator as the aging map identification zone of the face associated with each of the skin signs, and each zone in the aging map is filled with each aging target corresponding to the associated skin sign. The computing device of Technical Idea 4 is the computing device described in Technical Idea 3, where the aging map represents a specific aging target of the associated skin sign by the score value of the associated skin sign. The computing device of Technical Idea 5 is the computing device described in Technical Idea 3, where the aging map represents a specific aging target for the associated skin sign by the apparent aging value for the associated skin sign. The computing device of Technical Idea 6 is the computing device described in Technical Idea 3, where the aging map represents a specific aging target for the associated skin sign by the score value of the associated skin sign if available, and by the apparent aging value if the score value is not available. The computing device of Technical Idea 7 is a computing device described in any of Technical Ideas 3 to 6, wherein the aging map is defined to use pixel intensity to represent the aging target. The computing device of Technical Idea 8 is a computing device described in any of Technical Ideas 3 to 7, wherein the aging map masks the background of the original image. The computing device of Technical Idea 9 is a computing device described in any of Technical Ideas 1 to 8, wherein the generator is configured through training using respective training images and associated aging maps, and the associated aging maps provide weak spatial supervision for inducing the aging transformation of the respective skin signs. The computing device of Technical Idea 10 is a computing device described in any of Technical Ideas 1 to 9, wherein the skin signs represent ethnicity-specific dimensions of aging. The computing device of Technical Idea 11 is a computing device described in any of Technical Ideas 1 to 10, wherein the skin signs represent one or more of "age", "forehead wrinkles", "laugh lines", "under-eye wrinkles", "inter-brow wrinkles", "inter-ocular wrinkles", "corners of the mouth wrinkles", "upper lip", and "lower face sagging". The computing device of Technical Idea 12 is a computing device described in any of Technical Ideas 1 to 11, wherein the generator is a fully convolutional encoder-decoder with residual blocks in the decoder to incorporate the aging target in the form of an aging map. The computing device of Technical Idea 13 is a computing device described in Technical Idea 12, wherein the generator is configured using patch-based training that uses a part of a specific training image and the corresponding patch of the associated aging map. The computing device of Technical Idea 14 is a computing device described in Technical Idea 13, wherein the residual block further incorporates position information to indicate the respective positions of the parts of the specific training image and the corresponding patches of the associated aging map. The computing device of Technical Idea 15 is the computing device described in Technical Idea 14, wherein the position information is provided using respective X and Y coordinate maps defined from a horizontal gradient map and a vertical gradient map related to the height and width (H×W) size of the original image. The computing device of Technical Idea 16 is the computing device described in any one of Technical Ideas 13 to 15, wherein the specific training image is a high-resolution image and the patch size is a part of it. The computing device of Technical Idea 17 is the computing device described in Technical Idea 16, wherein the patch size is 1 / 2 or less of the high-resolution image. The computing device of Technical Idea 18 is the computing device described in any one of Technical Ideas 1 to 17, wherein the generator is configured via an attention mechanism to limit the generator to transform the skin signs related to the age while minimizing additional transformations applied. The computing device of Technical Idea 19 is the computing device described in any one of Technical Ideas 1 to 18, wherein the processing unit is configured to communicate with a second computing device that provides the generator for use, the processing unit communicates the original image and receives the transformed image. The computing device of Technical Idea 20 is the computing device described in any one of Technical Ideas 1 to 19, wherein the original image is a high-resolution image of 1024×1024 pixels or more. The computing device of Technical Idea 21 is the computing device described in any one of Technical Ideas 1 to 20, wherein the processing unit is further configured to provide an augmented reality application for simulating aging using the transformed image. The computing device of Technical Idea 22 is the computing device described in Technical Idea 21, comprising a camera, and the processing unit receives the original image from the camera. The computing device of Technical Idea 23 is the computing device described in any of Technical Ideas 1 to 22, wherein the processing unit is configured to provide at least one of a recommendation function for recommending at least one of products and services and an e-commerce transaction function for purchasing at least one of the products and the services. The computing device of Technical Idea 24 is the computing device described in Technical Idea 23, wherein the product includes one of a rejuvenating product, an anti-aging product, and a cosmetic makeup product. The computing device of Technical Idea 25 is the computing device described in Technical Idea 23, wherein the service includes one of a rejuvenating service, an anti-aging service, and a cosmetic service. The method of Technical Idea 26 is to define an age simulation generator having continuous control over a plurality of age-related skin signs between an original image and a transformed image of a face, training the generator using individual unpaired training images each associated with an age target for at least some of the skin signs, and providing a generator for transforming an image, thereby defining the age simulation generator, and providing the transformed image to the generator. The method of Technical Idea 27 is the method described in Technical Idea 26, wherein the generator is conditional GAN-based. The method of Technical Idea 28 is the method described in Technical Idea 26 or 27, including defining the aging target as an aging map that identifies the zones of the face associated with each of the skin signs, and each zone within the aging map being filled with a respective aging target corresponding to the associated skin sign. The computing device of Technical Idea 29 is a face effect unit including a processing circuit configured to apply at least one face effect to a source image and generate a virtual instance of the applied effect source image on an interface, and includes a generator for continuously controlling a plurality of age-related skin signs between an original image and a transformed image of a face to simulate aging, the generator being configured to transform the original image using respective aging targets for the skin signs. The computing device of Technical Idea 30 is, in the computing device of Technical Idea 29, a recommendation unit including a processing circuit configured to present recommendations for the product and / or service and receive a selection of the product and / or service, wherein the product and / or service is associated with an aging target modifier for at least one of the skin signs, and the face effect unit is configured to generate respective aging targets of the skin signs using the aging target modifier in response to the selection, thereby simulating the effect of the product and / or service on the source image. The computing device of Technical Idea 31 is, in the computing device of Technical Idea 30, configured such that the recommendation unit calls a skin sign analyzer to determine a current skin sign score using the source image, and obtains the recommendation by using the current skin sign score to determine the product and / or service. The computing device of Technical Idea 32 is, in the computing device of Technical Idea 31, configured such that the skin sign analyzer analyzes the source image using a deep learning model. The computing device of Technical Idea 33 is, in the computing device of Technical Idea 31 or 32, configured such that the aging target is defined from the current skin sign score and an aging target analyzer. The computing device of Technical Idea 34 is, in the computing device of any one of Technical Ideas 29 to 33, configured such that the generator is conditional GAN-based. The computing device of Technical Idea 35 is the computing device described in any of Technical Ideas 29 to 34, wherein the aging target is provided to the generator as the facial aging map identification zone associated with each of the skin signs, and each zone in the aging map is filled with the respective aging target corresponding to the associated skin sign. The computing device of Technical Idea 36 is the computing device described in Technical Idea 35, wherein the aging map represents a specific aging target of the associated skin sign by a score value of the associated skin sign. The computing device of Technical Idea 37 is the computing device described in Technical Idea 35, wherein the aging map represents a specific aging target for the associated skin sign by an apparent aging value for the associated skin sign. The computing device of Technical Idea 38 is the computing device described in Technical Idea 39, wherein the aging map represents a specific aging target for the associated skin sign, if available, by a score value of the associated skin sign, and if the score value is not available, by an apparent aging value. The computing device of Technical Idea 39 is the computing device described in any of Technical Ideas 35 to 38, wherein the aging map is defined to use pixel intensity to represent the aging target. The computing device of Technical Idea 40 is the computing device described in any of Technical Ideas 35 to 39, wherein the aging map masks the background of the source image. The computing device of Technical Idea 41 is the computing device described in any of Technical Ideas 29 to 44, wherein the skin signs represent one or more of "age", "forehead wrinkles", "laugh lines", "under-eye wrinkles", "inter-brow wrinkles", "inter-ocular wrinkles", "lip corner wrinkles", "upper lip", and "lower face sagging". The computing device of Technical Idea 42 is the computing device described in any one of Technical Ideas 29 to 41, wherein the generator is a fully convolutional encoder-decoder including a residual block in the decoder to incorporate the aging target in the form of an aging map. The computing device of Technical Idea 43 is the computing device described in any one of Technical Ideas 29 to 42, wherein the original image is a high-resolution image of 1024×1024 pixels or more. The computing device of Technical Idea 44 is the computing device described in any one of Technical Ideas 29 to 43, including a camera, and the computing device is configured to generate the original image from the camera. The computing device of Technical Idea 45 is the computing device described in any one of Technical Ideas 29 to 44, wherein the product includes one of a rejuvenating product, an anti-aging product, and a cosmetic makeup product. The computing device of Technical Idea 46 is the computing device described in any one of Technical Ideas 29 to 45, wherein the service includes one of a rejuvenating service, an anti-aging service, and a cosmetic service. The computing device of Technical Idea 47 is the computing device described in any one of Technical Ideas 29 to 46, wherein the interface includes an e-commerce interface for enabling the purchase of either a product or a service.

Claims

1. A computing device, comprising: a processing unit configured to receive an original image of a face and generate a converted image for presentation using a generator; the generator is a fully convolutional encoder-decoder that continuously controls a plurality of age-related skin signs between the original image and the converted image to simulate aging, uses each aging target of the skin signs to convert the original image, and includes a residual block in a decoder for incorporating the aging target in the form of an aging map, and is configured using patch-based training that uses a part of a specific training image and a corresponding patch of the associated aging map; the residual block further incorporates position information to indicate a corresponding patch of the aging map associated with each position of a part of a specific training image; the position information is provided using respective X and Y coordinate maps defined from a horizontal gradient map and a vertical gradient map related to the height and width (H×W) size of the original image, characterized computing device.

2. the generator is based on a conditional GAN; the aging target is provided to the generator as an identification zone of the aging map of the face associated with each of the skin signs, and each zone in the aging map is filled with a respective aging target corresponding to the associated skin sign, the computing device according to claim 1.

3. the aging map is one representing a specific aging target of the associated skin sign by a score value of the associated skin sign, one representing a specific aging target for the associated skin sign by an apparent aging value for the associated skin sign, or one representing a specific aging target for the associated skin sign by the score value of the associated skin sign when available and by the apparent aging value when the score value is not available, the computing device according to claim 2.

4. the aging map is A computing device according to claim 3, characterized in that it is defined to use pixel intensity to represent the aging target, or to mask the background of the original image.

5. The generator is configured through training using the associated aging map that provides weak spatial supervision for inducing the aging transformation of each of the training images and each of the skin signs. The computing device according to any one of claims 1 to 4, characterized in that the skin signs represent one or more of the ethnic-specific dimensions of aging, "age", "forehead wrinkles", "crow's feet", "under-eye wrinkles", "inter-brow wrinkles", "inter-ocular wrinkles", "lip corner wrinkles", "upper lip", and "lower face sagging".

6. A specific one of the training images is a high-resolution image, and the patch size is a part thereof. The computing device according to any one of claims 1 to 5, characterized in that the patch size is 1 / 2 or less of the high-resolution image.

7. The computing device according to any one of claims 1 to 6, characterized in that the generator is configured via an attention mechanism to limit the generator and transform the skin signs related to the age while minimizing additional transformations to be applied.

8. The processing unit is configured to communicate with a second computing device that provides the generator for use. The processing unit communicates the original image and receives the transformed image. The processing unit is further configured to provide an augmented reality application for simulating aging using the transformed image, or to receive the original image from a camera. The computing device according to any one of claims 1 to 7, characterized in that the original image is a high-resolution image of 1024×1024 pixels or more.

9. The computing device according to any one of claims 1 to 8, characterized in that the processing unit is configured to provide at least one of a recommendation function for recommending at least one of products and services and an e-commerce function for purchasing at least one of the products and the services.

10. The product includes one of a rejuvenating product, an anti-aging product, and a cosmetic makeup product, The computing device according to claim 9, wherein the service includes one of a rejuvenating service, an anti-aging service, and a cosmetic service.

11. A method comprising: defining a generator having continuous control over a plurality of age-related skin signs between an original image and a transformed image of a face, wherein each of the training images is an individual unpaired training image associated with an aging target for at least some of the skin signs, training the generator using the unpaired training images, and providing the generator for transforming images, thereby defining the generator; providing the transformed image to the generator; The generator is a fully convolutional encoder-decoder comprising a residual block in a decoder for incorporating the aging target in the form of an aging map, and is configured using patch-based training using a part of a specific training image and a corresponding patch of the associated aging map; The residual block further incorporates position information for indicating a corresponding patch of the aging map associated with each position of a part of a specific training image; The method is characterized in that the position information is provided using respective X and Y coordinate maps defined from a horizontal gradient map and a vertical gradient map related to the height and width (H×W) size of the original image.

12. The generator is conditional GAN-based; The method according to claim 11, wherein defining the aging target includes defining the aging target as the aging map that identifies a zone of the face associated with each of the skin signs, and each zone in the aging map is filled with a respective aging target corresponding to the associated skin sign.

13. A computing device, A face effect unit including a processing circuit configured to apply at least one face effect to a source image and generate a virtual instance of the applied effect source image on an interface, the face effect unit using a generator for continuously controlling a plurality of age-related skin signs between an original image and a transformed image of a face to simulate aging, the generator being configured to transform the original image using respective aging targets for the skin signs. Furthermore, the generator is a fully convolutional encoder-decoder including a residual block in a decoder for incorporating the aging targets in the form of an aging map, and is configured using patch-based training that uses a part of a specific training image and a corresponding patch of the associated aging map. The residual block further incorporates position information to indicate a corresponding patch of the aging map associated with each position of a part of the specific training image. The position information is provided using respective X and Y coordinate maps defined from a horizontal gradient map and a vertical gradient map related to the height and width (H×W) size of the original image, characterized by a computing device.

14. A recommendation unit including a processing circuit configured to present recommendations for products and / or services and receive a selection of the products and / or services, wherein the products and / or services are associated with an aging target modifier for at least one of the skin signs, and the face effect unit generates respective aging targets for the skin signs using the aging target modifier in response to the selection, thereby simulating the effect of the products and / or services on the source image. The recommendation unit is configured to call a skin sign analyzer to determine a score value of the current skin signs using the source image, and obtain the recommendation by using the score value of the current skin signs to determine the products and / or services. The skin sign analyzer is configured to analyze the source image using a deep learning model. The aging target is defined from the current score value of the skin sign and an aging target analyzer, or is provided to the generator as an identification zone of the facial aging map associated with each of the skin signs, and each zone in the aging map is satisfied with the respective aging target corresponding to the associated skin sign. The generator is the computing device according to claim 13, characterized in that it is based on conditional GAN.

15. The aging map represents a specific aging target of the associated skin sign by the score value of the associated skin sign, represents a specific aging target for the associated skin sign by an apparent aging value for the associated skin sign, or represents a specific aging target for the associated skin sign by the score value of the associated skin sign when available, and by an apparent aging value when the score value is unavailable. The computing device according to claim 14, characterized in that it is as described above.

16. The aging map is defined to use pixel intensity to represent the aging target, or masks the background of the source image. The computing device according to claim 15, characterized in that it is as described above.

17. The skin sign represents one or more of "age", "forehead wrinkles", "laugh lines", "under-eye wrinkles", "inter-brow wrinkles", "inter-ocular wrinkles", "lip corner wrinkles", "upper lip", and "lower face sagging". The original image is a high-resolution image of 1024×1024 pixels or more. The computing device is configured to generate the original image from a camera. The product includes one of a rejuvenating product, an anti-aging product, and a cosmetic makeup product. The service includes one of a rejuvenating service, an anti-aging service, and a cosmetic service. The interface includes an e-commerce interface for enabling the purchase of any of the product and the service. The computing device according to any one of claims 13 to 16, characterized in that it is as described above.

18. The aging map is separate and different from the original image, and represents a specific aging target of the relevant skin sign related by the score value of the relevant skin sign, characterized in that the computing device according to claim 1 or 13.

Citation Information

Patent Citations

  • Method for aging appearance simulation

    JP2020515952A