DEEP LEARNING SYSTEMS, DEVICES AND METHODS FOR PREDICTING HIGH DYNAMIC RANGE ENVIRONMENTAL PANORAMAS

A synthetic dataset-based deep learning system addresses the high cost and bias issues of existing lighting estimation models by predicting HDR panoramas from LDR images, ensuring diverse and realistic lighting conditions for augmented reality.

FR3160257B3Active Publication Date: 2026-04-10LOREAL SA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Utility models
Current Assignee / Owner
LOREAL SA
Filing Date
2024-03-14
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing deep learning models for lighting estimation in augmented reality are costly and time-consuming to train due to the need for matched datasets of real portrait images and HDR environment maps, leading to model biases towards certain demographic groups and limited environmental and subject diversity.

Method used

A deep learning system using a synthetic dataset trained on diverse subjects and environments to predict HDR environment panoramas from LDR limited field-of-view portrait images, employing a generator and discriminator model with specific loss functions to improve lighting estimation accuracy.

Benefits of technology

The system reduces training costs and biases by generating realistic lighting conditions across various demographics and environments, enabling efficient and unbiased lighting estimation for augmented reality applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000025_0000
    Figure 00000025_0000
  • Figure 00000025_0001
    Figure 00000025_0001
  • Figure 00000026_0000
    Figure 00000026_0000
Patent Text Reader

Abstract

DEEP LEARNING SYSTEMS, DEVICES, AND METHODS FOR PREDICTING HIGH DYNAMIC RANGE ENVIRONMENTAL PANORAMAS. Aspects of lighting estimation and related models are proposed, including aspects for training such models. A lighting estimation model pre-trained using synthetic data is proposed to alleviate the cost and difficulty of obtaining matched datasets of real portrait images and HDR environment maps. To improve model performance, the model is trained using a discriminator configured to predict one or more average color values ​​for a defined percentage of the highest intensity pixels in a predicted environment map and to determine a color loss associated with the predicted environment map and the one or more average color values.The trained model can be used for a wide range of downstream tasks, including generating hair renderings with realistic lighting effects for virtual try-on experiences. Figure for the abstract: none.
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: DEEP LEARNING SYSTEMS, DEVICES AND METHODS FOR PANORAMIC PREDICTION HIGH DYNAMIC BEACH ENVIRONMENT FIELD OF INVENTION

[0001] This disclosure relates to computer image processing and artificial intelligence, in particular lighting estimation systems and methods and more particularly to the prediction of high dynamic range (HDR) environmental panoramas from single low dynamic range (LDR) limited field of view portrait images. CONTEXT

[0002] The ambient light color and direction in a real-world scene have a significant impact on the content of an image that captures a view of that scene. Scene images can be processed, and new images can be produced as output to provide augmented reality experiences, including virtual try-on (VTO) experiences where one or more effects are applied to an object in the images. For example, processing can apply hair effects, makeup effects, and similar effects to a person's head in the image. VTO and other augmented reality experiences are often presented via a consumer-oriented computing device such as a smartphone, tablet, laptop, or desktop computer, where an input image of a scene includes a single, low dynamic range (LDR) limited field of view (FOV) portrait image, such as a user's "selfie."

[0003] A deep learning lighting estimation model can estimate the lighting scene of an input image, for example, to provide an environment map. In one example, the environment map can provide information on realistic rendering results in an augmented reality experience. Compiling portrait images with matched lighting data (e.g., an environment map) to train a lighting estimation model is expensive and time-consuming.

[0004] It is desired to make improvements to deep learning models for lighting estimation and their training. Summary

[0005] In at least some aspects, it is desirable to have a deep learning computer system and a method that predicts the HDR environment panorama(s) from LDR limited field-of-view portrait image(s) for estimate lighting conditions and furthermore generate realistic image data for use in augmented reality experiments.

[0006] In at least some aspects, the ability to estimate lighting conditions and predict HDR environment panorama(s) using a deep learning computer system and a method employing an estimation model trained on a synthetic image dataset would help overcome the costly and time-consuming process associated with collecting matched data of portrait images and HDR environment maps to train lighting estimation models. In at least some aspects, such a model trained on a synthetic image dataset would help address the problem of model biases with respect to certain demographic groups, which can result from the typical need to train such systems on a small real-world dataset due to time or cost constraints.

[0007] According to one aspect of the present invention, a computer system and a method using deep learning lighting estimation model(s), including aspects for training such model(s), are proposed. Furthermore, according to one embodiment, a synthetic dataset is proposed for training lighting estimation model(s) in order to reduce the cost and difficulty of obtaining matched datasets of real portrait images and HDR environment maps.According to one embodiment, a generated synthetic dataset is proposed in which a large synthetic dataset of subjects and environments is sampled to ensure that the dataset covers a diversity of lighting conditions and subject demographics, including subjects of diverse sexes, ethnic origins, and age groups, in order to ensure that lighting estimation models are not biased towards certain demographic groups.

[0008] In order to further improve model performance, according to one embodiment, an evaluation pipeline is proposed to assess the model's ability to infer lighting conditions, in the form of an environment map, from real portrait images. According to one embodiment, a video evaluation pipeline is proposed to qualitatively measure the predictive consistency of the model's environment map.

[0009] The following statements describe various aspects and features disclosed in the embodiments herein. These aspects, as well as others, will be obvious to those skilled in the art, such as aspects related to the computer program product. It is also understood that aspects of the computer device may exhibit corresponding process aspects, and vice versa.

[0010] Declaration 1: Computing device comprising a processor coupled to a storage device storing instructions executable by the processor to cause the computing device to: predict a predicted environment map for lighting conditions in an input image using a generator of a deep learning model, the predicted environment map encoding one or more light sources and an estimate of the light source color in the input image, the deep learning model having been defined by training using a discriminator that predicts one or more average color values ​​of a defined percentage of the highest intensity pixels of an environment map,the discriminator being defined using a color loss determined from: one or more average color values ​​of a defined percentage of the highest intensity pixels of a terrain reality environment map; and a prediction by the discriminator for a predicted environment map in training time generated by the generator; and generating an output image comprising one or more objects from the input image to which one or more respective effects are applied, in which a property of each of the respective effects is adapted to the predicted environment map.

[0011] Declaration 2: Computer device of Declaration 1, wherein the predicted training-time environment map encodes one or more predicted light sources and a light source color estimate of one of a plurality of synthetically generated images supplied to the generator, the plurality of synthetically generated images having lighting conditions comprising the light source direction and the light source color.

[0012] Declaration 3: Computer device of Declaration 2, wherein at least one of the following: the plurality of synthetically generated images comprises a set of portrait image data containing combinations of subjects and environments sampled to cover a diversity of demographics, light source directions and light source colors; the plurality of synthetically generated images comprises synthetic faces sampled on the basis of various sexes, ethnic origins and ages; the plurality of synthetically generated images are generated to include instances of neutral facial expressions and extreme facial expressions and at different head positions and orientations; or the plurality of synthetically generated images comprises images with a field of view established at a viewing angle between 50 and 70 degrees.

[0013] Declaration 4: Computer device of Declaration 1, wherein the deep learning model is defined by training comprising the generation of one or more bounding boxes to align a center of a head in the input image and perform one or both of a rotation or a transposition of an environment from the input image to align a center of the environment with the center of the head.

[0014] Declaration 5: Computer device of Declaration 1, wherein the defined percentage of the highest intensity pixels of the terrain reality environment map is in the range of 5 percent to 10 percent.

[0015] Declaration 6: Computer device of Declaration 5, in which the pixel intensity of the terrain reality environment map is defined on the basis of a perceptual brightness channel (L*) of the CIELAB color space.

[0016] Declaration 7: Computer device of Declaration 1, wherein the generator comprises one or more convolution output layers and one or more oversampling blocks, each of the oversampling blocks comprising a convolution layer, a ReLu activation layer and a bilinear oversampling layer.

[0017] Declaration 8: Computer device of Declaration 1, wherein the input image comprises a low dynamic range portrait image with a limited field of view and the output image comprises a high dynamic range image.

[0018] Declaration 9: Computer device of Declaration 1, wherein the effect applied to one or more objects from the input image is a virtual try-on effect.

[0019] Declaration 10: Computer device of Declaration 1, wherein the generator is defined by training using a reconstruction loss, the reconstruction loss being defined from an L2 loss applied to the predicted environment map in training time and to the field reality environment map.

[0020] Declaration 11: Computer device of Declaration 10, wherein the generator is defined by driving using a generator loss, the generator loss being determined from a weighted sum of the reconstruction loss and the color loss, and wherein the weighted sum is set to 10 to 1.

[0021] Declaration 12: Computer device of Declaration 1, wherein the generator is defined by training using a loss of coherence, the loss of coherence being determined from: a first predicted environment map generated by the generator for the lighting conditions of a first consecutive video frame; and the sum of a second predicted environment map generated by the generator for the lighting conditions of a second consecutive video frame and a third predicted environment map generated by the generator for the lighting conditions of a third consecutive video frame.

[0022] Declaration 13: Computer device of Declaration 1, further comprising a camera device configured to perform one or both of the capture or recording of the input image.

[0023] Declaration 14: A method for configuring a deep learning model that predicts lighting conditions in an input image and generates a predicted environment map for the lighting conditions in the input image, the method comprising: defining the deep learning model by training using a discriminator that predicts one or more average color values ​​of a defined percentage of the highest intensity pixels of an environment map, the discriminator being defined using a color loss determined from: one or more average color values ​​of a defined percentage of the highest intensity pixels of a field reality environment map; and a prediction by the discriminator for a predicted environment map in training time generated by a generator of the deep learning model for a lighting condition of a training image;and the generation of the predicted environment map using the generator, the generator having been defined by training using color loss.

[0024] Declaration 15: A method according to Declaration 14, wherein the training image comprises one of a plurality of synthetically generated images supplied to the generator, the plurality of synthetically generated images having lighting conditions comprising a light source direction and a light source color.

[0025] Declaration 16: A method of Declaration 15, wherein at least one of the following: the plurality of synthetically generated images comprises a set of portrait image data containing combinations of subjects and environments sampled to cover a diversity of demographics, light source directions and light source colors; the plurality of subjects of synthetically generated images comprises synthetic faces sampled on the basis of various sexes, ethnicities and ages; the plurality of synthetically generated images are generated to include instances of neutral facial expressions and extreme facial expressions and at different head positions and orientations; or the plurality of synthetically generated images comprises images with a field of view established at a viewing angle between 50 and 70 degrees.

[0026] Declaration 17: Method of Declaration 14, wherein the generator is defined by training using a reconstruction loss, the reconstruction loss being defined from an L2 loss applied to the training-time predicted environment map and the field-reality environment map.

[0027] Declaration 18: A method of Declaration 17, wherein the generator is defined by driving using a generator loss, the generator loss being determined from a weighted sum of the reconstruction loss and the color loss, and wherein the weighted sum is set to 10 to 1.

[0028] Declaration 19: A method of Declaration 14, wherein the generator is defined by training using a loss of coherence, the loss of coherence being determined from a first predicted environment map generated by the generator for the lighting conditions of a first consecutive video frame; and the sum of a second predicted environment map generated by the generator for the lighting conditions of a second consecutive video frame and a third predicted environment map generated by the generator for the lighting conditions of a third consecutive video frame.

[0029] Declaration 20: Computer device comprising: a VTO (virtual try-on) rendering engine configured to produce an output image for display, wherein the output image includes an object from an input image to which a VTO effect is applied, the VTO rendering engine comprising a deep learning model configured by training to provide a predicted environment map for the lighting conditions in the input image and wherein the effect is applied to the object as adapted by the predicted environment map, the deep learning model being defined by training by: the provision of a generator;the provision of a discriminator that predicts one or more average color values ​​of a defined percentage of the highest intensity pixels of an environment map, the discriminator being defined using a color loss determined from: one or more average color values ​​of a defined percentage of the highest intensity pixels of a field reality environment map; and a prediction by the discriminator for a training-time predicted environment map generated by the generator; and wherein the training-time predicted environment map encodes one or more predicted light sources and a light source color estimate from one or more of a plurality of synthetically generated images supplied to the generator, the plurality of synthetically generated images having lighting conditions comprising a light source direction and a light source color;and wherein the generator is defined by training using a reconstruction loss, the reconstruction loss being defined from an L2 loss applied to the training-time predicted environment map and the field reality environment map; and one or both of a VTO product recommendation interface or a product purchase interface, wherein the VTO effect is associated with a product to simulate a fitting. Brief description of the drawings

[0030] These and other features of the invention will become more evident from the following description, in which reference is made to the accompanying drawings on which:

[0031] [Fig-1] [Fig. 1] is a schematic diagram showing an example of a deep learning lighting estimation model, according to one embodiment.

[0032] [Fig.2] Fig.2 is a schematic diagram showing the estimation model lighting of [Fig.1] as adapted with a hair rendering, according to one embodiment.

[0033] [Fig.3] Fig.3 shows a set of evaluation images in accordance with a method of implementation.

[0034] [Fig.4] Fig.4 is a schematic diagram of a computer system, in accordance with an embodiment.

[0035] [Fig.5] Fig.5 is a schematic diagram showing the estimation model lighting of [Fig.1] as adapted with additional training refinement using consecutive video frames, according to one embodiment. DETAILED DESCRIPTION

[0036] The systems, devices, and methods described herein aim to improve existing deep learning applications for lighting estimation by providing a lighting estimation model trained on a synthetic image dataset. Generally, in at least some embodiments, systems and methods are proposed that use an estimation model trained on synthetic datasets to estimate the lighting conditions in an input photograph.

[0037] In order to obtain realistic rendering results in augmented reality experiences, research has been conducted to estimate the lighting scene of images using deep learning models. Gardner et al. [9] first presented an end-to-end deep neural network that directly regresses key lighting locations and intensities from a field-of-view photograph, without strong assumptions about scene geometry, material properties, or lighting. Previous models used to estimate the lighting scene of images can be classified into two main categories: regression models that estimate low-dimensional panoramic lighting parameters and generative models that generate nonparametric illumination maps or light probes.For example, there are models that regress on more lighting parameters such as the color of the light source [9], the angular size of the light in steradians. [9] or the 5th-order spherical harmonic representation of lighting

[10] . With the increasing popularity of generative models, there are also frame-by-frame generative models that directly generate an HDR environment map, such as styleLight [7]. Regression models have less flexibility and loss of detail in describing lighting information and require additional post-processing steps to convert lighting parameters into an HDR environment map for image-based relighting; as such, a generative model for lighting estimation is desirable.

[0038] Previous work has explored ways to infer lighting information in real-time augmented reality applications. LeGender et al. [2] first proposed models that predict low-resolution HDR light probes based on unconstrained, field-of-view LDR images, and then extended the work to focus on predicting light probes based on field-of-view LDR portrait images [5]. These models are small models that can infer in real time on mobile applications. However, because the light probes have low resolution and only capture lighting information behind the camera, this limits realism when using light probes to render images. To address this issue, Somanath et al.[6] presented a model that deduces high-resolution HDR panorama maps based on unconstrained, limited-field-of-view LDR images.

[0039] Developing generative models for lighting estimation presents numerous challenges. For example, collecting paired portrait image and HDR environment map data is a costly task, as it requires that the facial lighting in the portrait images be consistent with the corresponding environment map. This limits the ability to collect a large dataset of portrait and light map data with diverse environments and subjects. Thus, models trained on small real-world datasets cover only a limited variety of subjects and environments, and may introduce model bias toward certain demographic groups.

[0040] Previously, the Laval Face+Lighting HDR dataset [1] was created by inviting 9 subjects into 25 environments with different lighting conditions, and sequentially photographing the environments at different orientations and exposures, and the subjects. LeGender et al. [2] generated such a dataset by first recording the reflectance field and alpha mat of 70 diverse subjects using a bright scene, and separately collecting HDR environment maps, then re-lighting the subjects using image-based re-lighting to create portrait image-environment map pairs. Similarly, Sztrajman et al. [3] also generated the dataset by combining scanned faces re-illuminated images from the ICT 3d Relightable Facial Expression database [4] with environment maps from the HDR Laval Face+Lighting dataset [1]. This can increase the variety of environment maps in the dataset, but the dataset still only covers a small variety of subjects. MODEL ARCHITECTURE

[0041] Figure 1 shows a schematic diagram of a computing device in which one or more technologies or methodologies can be implemented, such as, for example, estimating lighting conditions for synthetic content in an augmented reality environment, a mixed reality environment, a virtual fitting environment, and the like. In one embodiment, the computing device includes processor circuitry and one or more deep learning components configured to predict high dynamic range (HDR) environment panoramas from single low dynamic range (LDR) limited field-of-view portrait images.

[0042] In one embodiment, the computer device is configured to provide an estimation stream 100 according to an embodiment in which an estimation model 102 is configured to receive an input image 104 as input, and provides an output environment map prediction 112. In a training mode, the input image 104 represents a training image associated with a matched terrain reality map 126. In an inference time mode, the input image 104 represents an inference time image such as a user image to be used in an augmented reality experience.

[0043] It should be understood that [Fig. 1] represents the estimation stream 100 as data components and executable components stored within one or more storage devices 103 of a computing device 105, according to one embodiment. The executable components are executable by a processor (not shown) of the computing device. In one embodiment, the input image 104 comprises a limited field-of-view LDR portrait image. In another embodiment, the input image 104 comprises a synthetically generated image. In one embodiment, the estimation model 102 comprises a generator based on an encoder-decoder architecture. In one embodiment, the image feature extraction 106 uses MobileNet v2 to extract image features from the input image 104.In one embodiment, the oversampling blocks 108 comprise 5 oversampling blocks, in which each oversampling block comprises a convolution layer, a ReLu activation layer, and a bilinear oversampling layer. According to the embodiment of [Fig. 1], the estimation model 102 further comprises the convolution output layer 110. The estimation model. 102 outputs a predicted environment map, generated by a generator, as environment map prediction 112. In one embodiment, environment map prediction 112 is predicted on a logarithmic scale to convey long-tail RGB color distribution data from HDR environment maps.

[0044] In one embodiment, the discriminator 116 comprises 6 convolution layers, and each convolution layer is followed by a Leaky ReLu activation layer. In some embodiments, the discriminator 116 further comprises a fully connected output layer (not shown) configured to predict average RGB values. In one embodiment, the estimation stream 100 represents a training stream with one training time. The components 102 are retained during the inference time.

[0045] The input image 104 can be received from another computing device (not shown) over a communication network (not shown) or at least partially entered by a user on a computing device for the estimation model (e.g., a computing device 300 shown in [Fig. 3]). Loss functions

[0046] Previous work [7, 8] typically applies a loss per pixel to the first 5 to 10% of pixels of highest intensity in an environment map. In accordance with experimental results and the techniques and teachings presented herein, the training of the estimation model is improved when such a loss is instead calculated by a discriminator. As such, in one embodiment, the environment map prediction 112 generated by a generator is output by the estimation model 102 and passed to the discriminator 116 (an executable component).The discriminator 116 is trained to predict the average RGB value of defined percentages of the highest intensity pixels in a terrain reality environment map (e.g., matched terrain reality map 126) based on an L1 loss, in which defined percentages of the highest intensity pixels in the terrain reality environment map are calculated based on the perceptual brightness (L*) channel of the CIELAB color space. In one embodiment, the defined percentages are the first 10% of highest intensity pixels and the first 10 to 50% of highest intensity pixels. In one embodiment, the color loss based on a discriminator 118 is determined according to the formula: .

[0047] Discriminator loss = L1(D toplQ(card gt ), card gt. mask topWgt )+ L1(D toplQ_5Q(card gt ), card gt. mask topw>-wgt )

[0048] wherein D is the discriminator 116, map gt is the matched ground reality map 126, and maskgt is a binary mask based on the defined percentage of highest intensity pixels of the matched truth map. In one embodiment, the defined percentage of mask toPiogt is the first 10% of highest intensity pixels. In another embodiment, the defined percentage of mask io_sogt is the first 10 to 50% of highest intensity pixels.

[0049] In one embodiment, the estimation model 102 is trained on the sum of a reconstruction loss 120 and a light source color loss 114, in which the relative weights of the reconstruction loss with respect to the light source color loss are set at 10 to 1. In another embodiment, the color loss based on a discriminator 118 is further used to update the training of the estimation model 102.

[0050] In one embodiment, the light source color loss 114 is calculated by first passing a predicted environment map (for example, environment map prediction 112) to the discriminator 116 to calculate the average RGB of a defined percentage of highest intensity pixels, where intensity is defined on the basis of the perceptual brightness channel (L*) of the CIELAB color space. An L1 loss is applied between the output of the discriminator 116 and the average RGB of the defined percentage of highest intensity pixels of the ground reality environment map (for example, the matched ground reality map 126) according to the formula:

[0051] Light loss wallCe = L1(D lopW(E(x)), map gt* mask lopWgl )+ L1(D lopW_50(E(x)), map gt * toPio-50gt mask )

[0052] wherein D is the discriminator 116, E is the estimation model 102, x is the input image 104, cartegt is the matched terrain reality map 126, and masquegt is a binary mask based on the defined percentage of highest intensity pixels in the matched terrain reality map 126. In one embodiment, the defined percentage of mask topWgt is the first 10% of highest intensity pixels. In another embodiment, the defined percentage of mask topio-5ogt is the first 10 to 50% of highest intensity pixels. In one embodiment, the color loss module 124 includes an executable component configured to execute a light source color loss function and provide a light source color loss 114 to the estimation model 102.

[0053] In one embodiment, the reconstruction loss 120 includes an L2 loss applied to minimize a difference per pixel between the environment map prediction 112 and a ground reality environment map (for example, the matched ground reality map 126) according to the formula:

[0054] Loss Reconstitution = E2(E(x), card gt )

[0055] in which E is the estimation model 102, x is the input image 104, and the map includes a ground reality environment map (e.g., matched ground reality map 126).

[0056] In one embodiment, an overall training loss function is applied by the estimation model 102 comprising the sum of the color loss based on a discriminator 118 and the weighted sum of the reconstruction color loss 120 and a light source color loss 114, wherein the relative weights of the reconstruction color loss with respect to the light source color loss are set at 10 to 1. In one embodiment, the reconstruction loss module 122 includes an executable component configured to execute a reconstruction loss function and provide a reconstruction loss 120 to the estimation model 102. Synthetic training data

[0057] In one embodiment, the estimation model 102 is trained with images (for example, the input image 104) sampled from a synthetic training dataset created using the DataGen platform. In one embodiment, the synthetic training dataset comprises 50,000 portrait images with combinations of more than 5,000 subjects and 250 synthetic environment maps. According to one embodiment, the synthetic subjects and environments are broadly sampled to ensure that the synthetic training dataset covers a diversity of demographics and lighting conditions. In one embodiment, the environment maps include both indoor and outdoor scenes and are at different times of day, for example, morning and evening.

[0058] To provide a synthetic training dataset representing directional lights from different angles, according to one embodiment, the environment maps in the dataset are rotated horizontally at different angles during portrait image rendering. To ensure that an estimation model (e.g., 102) is not biased toward certain demographic groups, synthetic subjects are sampled from the synthetic dataset based on various gender, ethnicity, and age groups. According to another embodiment, synthetic subjects are generated for the synthetic training dataset with neutral and extreme expressions, and at different head positions and orientations, such as frontal and profile faces. According to one embodiment, the field of The camera's field of view in synthetic portrait images is fixed between 50 and 70 degrees to resemble the front camera of mobiles and webcams.

[0059] In one embodiment, bounding boxes are used to center the aligned heads in the portrait images to ensure that the estimation model (e.g., 102) focuses on learning the face light map during training. In another embodiment, the environment maps are rotated and / or shifted so that the centers of the environment maps are aligned with the location of the subjects' faces in the portrait images. In yet another embodiment, random horizontal image flipping and increased color flicker of the images are used during the training of an estimation model (e.g., 102).

[0060] In one embodiment, the environment map prediction 112 delivered as output by the estimation model 102 can be used to locate key directional light sources and estimate the ambient light color from the input image 104. In one embodiment, the estimation model 102 can infer environment panoramas from the input image 104 in real time and can be deployed in mobile or web applications. Evaluation using synthetic data

[0061] In one embodiment, the performance of an estimation model (e.g., 102) is evaluated quantitatively using synthetic data. An estimation model (e.g., 102) is examined on the basis of the residual mean squared error (RMSE), the scaled residual mean squared error (sMRSE), the RGB angular error and the Fréchet inception distance (FID) between the environment map prediction (e.g., 112) and a terrain reality environment map (e.g., the matched terrain reality map 126).

[0062] In one embodiment, the angular error is calculated directly on predicted equirectangular environment maps (for example, 112), while other metrics are calculated on the basis of a cubic map converted from the predicted equirectangular maps. Hair rendering estimation model

[0063] Figure 2 shows a schematic diagram of an example of a computer device configured to provide an estimation stream 100 according to an embodiment of Figure 1. In one embodiment, the computer device 105 is further configured to generate a hair rendering 222 suitable for the environment map prediction 112. In one embodiment, the hair rendering 222 imitates real hair in portrait images (for example, the input image 104). For example, rendering operations overlay hair pixels onto a portrait image (e.g., an input image) to simulate a hairstyle in a generated output image. In one embodiment, the hairstyle is defined according to a 3D hair mesh that defines a plurality of hair strands, etc., for the hairstyle. In another embodiment, the light source direction(s) and light source color(s) of the environment map are used to adapt a property of the pixels in the output image. The color and / or brightness of the hair pixels or other pixels can be adjusted, for example, using shading and lighting techniques to illuminate, for example, the hair to match the ambient lighting of the input image. Shadows can be determined and rendered.In one embodiment, the environment map is converted into a set of discrete light sources by detecting the brightest points on the map and finding the average color for each point to determine the color and intensity of the light. Each bright point is converted into a directional light source. In one embodiment, hair shading is performed using a Marschner hair shading model with, for example, directional light sources. In another embodiment, indirect lighting is approximated using a double diffusion technique. In one embodiment, lookup tables help optimize real-time performance. Evaluation based on real data

[0064] Since the provided estimation models are trained on synthetic data, it is preferable to assess how well the estimation models fit real data. Evaluation with real data is typically done either by directly comparing predicted environment maps (e.g., 112) with real environment maps, or by rendering relit objects based on predicted environment maps (e.g., 112) and maps of the real environment. Such an assessment is difficult due to the lack of ground-reality lighting maps and face or hair albedo maps in portrait images.

[0065] To meet this challenge, an estimation model (e.g., 102) is evaluated by relit hair renderings (e.g., 222) based on predicted environment maps (e.g., 112) that mimic real hair in portrait images, and by comparing the frame changes in the predicted environment maps (e.g., 112) for real videos.

[0066] According to one embodiment, to examine the model's ability to predict ambient light color and locate strong directional lights, light maps are rendered on hair renderings that mimic hair real-world portrait images. Ten hair experts were asked to annotate hair color on 1000 portraits. For each portrait image, the color with the highest agreement among the hair experts' annotations was used as the real-world color. Based on the annotated hair color, the average RGB of the corresponding hair strand images was calculated, and hair renderings were generated based on the average RGB. In one embodiment, the real-world portrait images (e.g., 104) are fed into an estimation model (e.g., 102) to predict light maps (e.g., environment map prediction 112), and hair meshes (e.g., a hair simulation from a 3D mesh definition) are rendered (e.g., 222) based on the predicted light maps.The rendered hair mesh images are compared to the actual portrait images entered into the estimation model to assess how close the overall hair colors are and whether the specular light locations are consistent.

[0067] Figure 3 shows a set of evaluation images 300 according to one embodiment. The image set 300 includes portrait images 302A, 302B (e.g., input images), rendered hair mesh images 304A, 304B based on the predicted environment map (not shown), and annotated hair color images 306A and 306B. Images 306A and 306B are compiled from the average RGB annotated by hair experts. For simplicity, Figure 3 is represented as grayscale representations of the images based on the original RGB colors. The subject in image 302A is originally shown against a predominantly green vegetation background with dark-colored (almost black) hair. The subject of image 302B is shown against a predominantly white and grey background with hair that is generally dark reddish-blond in color.

[0068] In one embodiment, a video evaluation method is also used to examine the ability of an estimation model (e.g., 102) to locate key directional lights and the consistency in generating an environment map prediction (e.g., 112). Videos with stable and moving camera positions are used to assess whether the model is capable of making reasonable and consistent predictions of powerful light sources. According to one embodiment, lighting estimation models are run on each frame of a video input in an estimation model to estimate a light map per frame.Subsequently, qualitative assessments are conducted on the stability of a predicted light map in videos where the background of an input video is stable, and on the regularity of the movements of key light sources in videos when the camera recording the input video is moving.

[0069] Quantitative analysis indicated that an estimation model according to one embodiment was able to locate strong light sources from an input image and to identify whether the ambient light of an input image is cool or warm.

[0070] Computer device using an estimation model in a VTO pipeline

[0071] Figure 4 is a schematic diagram of a computer system according to one embodiment. In this embodiment, an estimation model according to one embodiment of this system is integrated into a virtual try-on (VTO) application to provide hair simulation, such as hair simulation from a 3D hair mesh providing a definition of a hairstyle to be simulated. The rendering can be performed as described with reference to Figure 2. Figure 4 is not exhaustive and is simplified for brevity.

[0072] The computer device 400 includes one or more processors 402, one or more input devices 404, one or more communication units 408, one or more output devices 410, a display screen 448 (for example, providing one or more graphical user interfaces), a camera 446, and memory 406. The computer device 400 also includes one or more storage devices 414 storing one or more executable computer modules including an application and lighting estimation data 416 comprising: input image 418 (for example, 104), environment map prediction 422 (for example, 112), estimation model 420 (for example, estimation model 102), hair rendering 424 (for example, hair rendering 222), estimation model data store 428, VTO pipeline 440, and a user interface. 448, product purchase interface 442, and image with hair rendering 444.The computer device 400 may include additional computer modules or data stores in various embodiments. Additional computer modules and devices that may be included in various embodiments are not shown in [Fig. 4] to avoid undue complexity in the description, such as communication with one or more other computer devices, as appropriate, using a communication unit(s) 408, to obtain the input image 418, including via a communication network (not shown).

[0073] The storage device(s) 414 stores data and / or computer-readable instructions for execution by a processing unit (e.g., processor(s) 402), such that once executed, the instructions cause the computer to perform operations such as one or more processes. The one or more storage devices 414 can take various forms and / or configurations, for example, short-term or long-term memory. The 414 storage device(s) can be configured for short-term storage of information as volatile memory, which does not retain its contents when power is turned off. Examples of volatile memory include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), etc. The 414 storage device(s), in some examples, also include one or more computer-readable storage media, for example, to store larger amounts of information than volatile memory and / or to store this information long-term, retaining the information when power is turned off.Examples of non-volatile memory include magnetic hard drives, solid-state drives, optical discs, floppy disks, flash memory, or forms of electrically programmable read-only memory (EPROM) or electrically programmable eraseable read-only memory (EEPROM). The computing device 400 can store data / information (e.g., the input image 418, the environment map prediction 422, or the image with hair rendering 444) in the storage device(s) 414.

[0074] One or more processors 402 can implement functionality and / or execute instructions within the computing device 400. For example, the processor(s) 402 can be configured to receive instructions and / or data from the storage device(s) 414 to execute the functionality of the estimation models (e.g., 102) shown in Figures 1 and 2, among other modules (e.g., operating system 430, browser 432, email and / or messaging application 434, social networking application 436, etc.). The processor(s) 402 comprise one or more central processing units (e.g., CPUs) and / or graphics processing units (e.g., GPUs) having one or more processors / microprocessors, controllers / microcontrollers, etc. Other types of processors may be used.GPUs can be particularly useful for accelerating graphics processing tasks and / or deep learning processing tasks (e.g., training and / or inference).

[0075] The input device(s) 404 and output device(s) 410 may include one or more buttons, switches, pointing devices, a keyboard, a microphone, one or more sensors, a speaker, a buzzer, one or more lights, etc. One or more of these may be connected via a wired connection (e.g., Ethernet, USB-A, USB-C, Thunderbolt (Intel Corporation™), or other communication channel). In one embodiment, the camera 446 is a camera capable of capturing and / or recording a field-of-view LDR portrait image limited vision. In one embodiment, the 446 camera can be used to take selfies.

[0076] One or more communication units 408 can communicate with external computing devices via one or more networks by transmitting and / or receiving network signals on one or more networks. The communication units 408 can include various antennas and / or network interface cards, etc., for wireless and / or wired communications.

[0077] The display screen 448 presents images such as components of a graphical user interface, camera images, etc. In one embodiment, the display screen is a touch screen device, a type of I / O device, configured to receive gesture inputs (e.g., swipes, taps, etc.) which interact with the screen region(s) and in association with the user interface components (e.g., commands) presented by an application executed by the processor(s) 402.

[0078] The computing device 400 may include additional computing modules or data stores in various embodiments. Additional modules, data stores, and devices that may be included in various embodiments may not be shown in [Fig. 4] to avoid undue complexity in the description. Other examples of computing device 400 may be an electronic tablet, a personal digital assistant (PDA), a laptop computer, a desktop computer, a portable media player, an e-reader, a watch, a client device, a user device, or another type of computing device.

[0079] The storage device(s) 414 store(s) components of a light estimation application and related data (e.g., 416). Representative components are shown. The light estimation application and data 416 include the user interface component 438 (e.g., screens, instructions, icons, commands, etc.). The user interface provides output to a user and receives input such as input for the application workflow. An estimation model 420 is provided and includes an estimation model (e.g., 102) as previously described for predicting image lighting conditions (e.g., environment map prediction 422). A VTO pipeline 440 is provided for simulating hair rendering with environment-sensitive lighting in association with (e.g., on or in) an input image.

[0080] In one embodiment, a user can provide an input image (e.g., 418) similar to the input image 104, for lighting estimation. In one embodiment, the input image (e.g., 418) is captured and / or recorded by a camera (e.g., camera 446). The environment map prediction 422 is generated by the estimation model 420 using the input image 418 according to an embodiment proposed herein. The environment map prediction 422 can be used to generate a hair rendering 424 simulating hair adapted to the lighting conditions present in the input image 418. The hair rendering 424 can be displayed on the display screen 448. The user can invoke the light estimation application 416 (via an input to a command) so that the application simulates hair in response to the lighting conditions present in an input image 418, using the image 418 and the environment map prediction 422 as input to produce an output image with hair rendering 444.

[0081] The output image with hair rendering 444 can be presented via the user interface component 438 on the display screen 448. The products or services or both can be purchased via the interface 442. Such an interface 442 can direct the user (e.g., the computer device) to an e-commerce service on the Web (e.g., a website (not shown)) to make a purchase, a reservation or the like.

[0082] The other components stored in the storage device(s) 414 include an operating system 430, a browser 432 (for example, for browsing web pages), an email and / or messaging application 434 (for example, SMS or other type) and a social networking application 436. The output image with hair rendering 444 can be shared (for example, communicated) via the applications 434 and / or 436, for example.

[0083] Figure 5 shows a schematic diagram of an example computer device configured to provide an estimation stream 100 according to an embodiment of Figure 1. In one embodiment, the computer device 105 is further configured to train the estimation model 102 by refinement associated with video frames. Consecutive video frames 528, comprising two or more consecutive frames from an input video, are provided to the estimation model 102. In one example embodiment, three consecutive frames from an input video are provided to the estimation model 102. The estimation model 102 is configured to provide a video frame prediction map 530, comprising a light environment map prediction for each of two or more consecutive video frames 528, to the coherence loss module 532.

[0084] In one embodiment, the coherence loss 534 is calculated by applying an L1 loss per pixel between the video frame prediction map 530 of one of the consecutive video frames 428, and a sum of the video frame prediction map 530 for two additional consecutive video frames 528 according to the formula:

[0085] Loss Consistency = Ll((E(v l+] )), 0.5 * (E(vt) + E(v t+2 )))

[0086] in which E is the estimation model 102, vt indicates 528 consecutive video frames at time t, vt+1 indicates 528 consecutive video frames at time t + 1, and vt+2 indicates 528 consecutive video frames at time t + 2.

[0087] In one embodiment, the computer device 105 is configured to refine the training of the estimation model 102 using a loss function comprising the sum of the reconstruction loss 120, the light source color loss 114, the discriminator-based color loss 118, and the coherence loss 534.

[0088] A practical implementation may include all or part of the features described herein. These features, characteristics, and various combinations thereof, as well as others, may be expressed in the form of processes, devices, systems, means of performing functions, program products, and other means, combining the features described herein. A number of embodiments have been described. Nevertheless, it is understood that various modifications may be made without departing from the spirit and scope of the processes and techniques described herein. Furthermore, other steps may be provided, or steps may be eliminated, from the described process, and other components may be added to or removed from the described systems. Accordingly, other embodiments fall within the scope of the following claims.

[0089] Throughout the description and claims of this document, the words "include," "contain," and their variations mean "including but not limited to," and are not intended to exclude (and do not exclude) other components, integers, or steps. Throughout this document, the singular encompasses the plural unless the context requires otherwise. In particular, where the indefinite article is used, this document is to be understood as considering plurality as well as singularity, unless the context requires otherwise.

[0090] The features, integers, characteristics, or groups described in conjunction with a particular aspect, embodiment, or example of the invention shall be understood as applicable to any other aspect, embodiment, or example, unless inconsistent with them. All features disclosed herein (including the claims, abstract, and accompanying drawings), and / or all steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of these features and / or steps are mutually exclusive. The invention is not limited to the details of the preceding examples or embodiments. The invention extends to any new feature, or any new combination, of the features disclosed herein (including any claim, abstract and any attached drawings) or to any new feature, or any new combination, of the steps of any disclosed method or process. REFERENCES

[0091] The following documents are cited:

[0092] D. Calian, J.-F. Lalonde, PFU Gotardo, T. Simon, I. Matthews and K. Mitchell, “From Faces to Outdoor Light Probes”, Computer Graphics Forum, 37.2 / 2018

[0093] LeGendre, C., Ma, WC, Pandey, R., Fanello, S., Rhemann, C., Dourgarian, J., Busch, J. and Debevec, P., 2020. Leaming illumination from diverse portraits. In SIGGRAPH Asia 2020 Technical Communications (pp. 1-4).

[0094] Sztrajman, A., Neophytou, A., Weyrich, T. and Sommerlade, E., 2020, November. High-Dynamic-Range Lighting Estimation From Face Portraits. In 2020 International Conference on 3D Vision (3DV) (pp. 355–363). IEEE.

[0095] Stratou, G., Ghosh, A., Debevec, P. et Morency, L.P., 2011, mars. Effect of illumination on automatic expression récognition: a novel 3D relightable facial database. Dans 2011 IEEE International Conférence on Automatic Face & Gesture Récognition (FG) (pp. 611-618). IEEE.

[0096] LeGendre, C., Ma, W.C., Fyffe, G., Flynn, J., Charbonnel, L., Busch, J. et Debevec, P., 2019. Deeplight: Leaming illumination for unconstrained mobile mixed reality. Dans Proceedings of the IEEE / CVF Conférence on Computer Vision and Pattern Récognition (pp. 5918-5928).

[0097] Somanath, G. et Kurz, D., 2021. HDR environment map estimation for real-time augmented reality. Dans Proceedings of the IEEE / CVF Conférence on Computer Vision and Pattern Récognition (pp. 11298—>11306).

[0098] Wang, G., Yang, Y., Loy, CC and Liu, Z., 2022, October. Stylelight: HDR panorama generation for lighting estimation and editing. In European Conference on Computer Vision (pp. 477-492). Cham: Springer Nature Switzerland.

[0099] Gardner, MA, Hold-Geoffroy, Y., Sunkavalli, K., Gagné, C. and Lalonde, JF, 2019. Deep parametric indoor lighting estimation. In Proceedings of the IEEE / CVF International Conference on Computer Vision (pp. 7175-7183).

[0100] Gardner, MA, Sunkavalli, K., Yumer, E., Shen, X., Gambaretto, E., Gagné, C. and Lalonde, JF, 2017. Leaming to predict indoor illumination from a single image. arXiv preprint arXiv: 1704.00090.

[0101] Garon, M., Sunkavalli, K., Hadap, S., Carr, N. and Lalonde, JF, 2019. Fast spatially-varying indoor lighting estimation. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 6908-6917).

Claims

Demands

1. A computing device comprising a processor coupled to a storage device storing instructions executable by the processor to enable the computing device to: (a) predict a predicted environment map for lighting conditions in an input image using a generator of a deep learning model, the predicted environment map encoding one or more light sources and a light source color estimate in the input image, the deep learning model having been defined by training using a discriminator that predicts one or more average color values ​​of a defined percentage of the highest intensity pixels of an environment map,the discriminator being defined using a color loss determined from: (i) one or more average color values ​​of a defined percentage of the highest intensity pixels of a terrain reality environment map; and (ii) a prediction by the discriminator of a predicted training-time environment map generated by the generator; and (b) generating an output image comprising one or more objects from the input image to which one or more respective effects are applied, wherein a property of each of the respective effects is adapted to the predicted environment map.

2. A computer device according to claim 1, wherein the training-time predicted environment map encodes one or more predicted light sources and a light source color estimate from one of a plurality of synthetically generated images supplied to the generator, the plurality of synthetically generated images having lighting conditions comprising a light source direction and a light source color.

3. A computer device according to claim 2, wherein at least one of: (a) the plurality of synthetically generated images comprises a portrait image dataset containing combinations of subjects and environments sampled to cover a diversity of demographic data, light source directions and light source colors; (b) the plurality of synthetically generated image subjects includes synthetic faces sampled on the basis of various sexes, ethnicities and ages; (c) the plurality of synthetically generated images are generated to include instances of neutral facial expressions and extreme facial expressions and at different head positions and orientations; or (d) the plurality of synthetically generated images includes images with a field of view established at a viewing angle between 50 and 70 degrees.

4. A computer device according to claim 1, wherein the deep learning model is defined by training comprising generating one or more bounding boxes to align a center of a head in the input image and performing one or both of a rotation or a transposition of an environment of the input image to align a center of the environment with the center of the head.

5. A computer device according to claim 1, wherein the defined percentage of the highest intensity pixels of the terrain reality environment map is in the range of 5 percent to 10 percent; and wherein the pixel intensity of the terrain reality environment map is defined on the basis of a perceptual brightness channel (L*) of the CIELAB color space.

6. Computer device according to claim 1, wherein the generator comprises one or more convolution output layers and one or more oversampling blocks, each of the oversampling blocks comprising a convolution layer, a ReLu activation layer and a bilinear oversampling layer.

7. A computer device according to claim 1, wherein the input image comprises a low dynamic range portrait image with a limited field of view and the output image comprises a high dynamic range image.

8. Computer device according to claim 1, wherein the effect applied to one or more objects in the input image is a virtual try-on effect.

9. Computer device according to claim 1, wherein the generator is defined by training using a reconstruction loss, the reconstruction loss being defined from an L2 loss applied to the training-time predicted environment map and the field-reality environment map.

10. A computer device according to claim 1, wherein the generator is defined by training using a loss of coherence, the loss of coherence being determined from: (a) a first predicted environment map generated by the generator for the lighting conditions of a first consecutive video frame; and (b) the sum of a second predicted environment map generated by the generator for the lighting conditions of a second consecutive video frame and a third predicted environment map generated by the generator for the lighting conditions of a third consecutive video frame.