Learn lighting from various portraits
By using a machine learning-based method, multiple bidirectional reflectance distribution functions and a large-scale lighting environment database, the difficulty of lighting estimation in augmented reality applications is solved, and accurate lighting restoration and realistic rendering of virtual objects under different skin colors are achieved.
Patent Information
- Application Number
- CN202080089261.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-20
- Filing Date
- 2020-09-21
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2040-09-21
AI Technical Summary
Existing technologies struggle to effectively match the lighting of real-world scenes in augmented reality applications, especially in portraits taken with front-facing cameras. Lighting estimation methods are limited by the complexity of skin reflections and the inherent ambiguity between light source intensity and surface albedo, resulting in inaccurate lighting matching between virtual content and real scenes.
A machine learning-based approach is adopted, using multiple bidirectional reflectance distribution functions as loss functions. The model is trained to estimate high dynamic range omnidirectional lighting from portrait photos. A large database of indoor and outdoor lighting environments is used for re-lighting. The multi-scale adversarial loss function and the rendering-based loss function are combined to generate accurate lighting estimates.
Under different natural skin colors, it can accurately restore light intensity and surface reflectivity, achieve realistic rendering of virtual objects in portrait photos, and is suitable for real-time lighting estimation in augmented reality applications.
Smart Images

Figure CN114846521B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is a non-provisional application based on and claims priority to U.S. Provisional Patent Application No. 62 / 704,657, filed on May 20, 2020, entitled “LEARNING ILLUMINATION FROMPORTRAITS,” the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] The present disclosure relates to determining lighting from a portrait for use in, for example, augmented reality applications. Background Art
[0004] A common problem in both still photography and video applications is matching the lighting of a real-world scene so that rendered virtual content plausibly matches the appearance of the scene. For example, lighting schemes can be designed for augmented reality (AR) use cases with a world-facing camera, such as in the rear camera of a mobile device, where one might want to render a synthetic object (such as a piece of furniture) into a live camera feed of a real-world scene. Summary of the Invention
[0005] Embodiments disclosed herein provide a learning-based technique for estimating high dynamic range (HDR) omnidirectional lighting from a single low dynamic range (LDR) portrait image captured under arbitrary indoor or outdoor lighting conditions. This technique involves training a model using portrait photos paired with the portrait photos' ground truth ambient lighting. This training involves generating a rich collection of such photos by using a light stage to record reflectance fields and alpha masks for 70 diverse subjects with various expressions. The subjects are then relighted using image-based relighting from a database of one million HDR lighting environments, composited onto paired high-resolution background images recorded during the light acquisition. The lighting estimation model is trained using a rendering-based loss function and, in some cases, a multi-scale adversarial loss to estimate plausible high-frequency lighting details. This learning-based technique robustly handles the inherent ambiguity between overall lighting intensity and surface albedo, thereby recovering similar-scale lighting for subjects with a variety of natural skin colors. This technique further allows virtual objects and digital characters to be added to portrait photos with consistent lighting. This lighting estimation can run in real time on a smartphone, enabling realistic rendering and compositing of virtual objects into live video for augmented reality (AR) applications.
[0006] In one general aspect, a method can include receiving image training data representing a plurality of images, each of the plurality of images including at least one of a plurality of human faces, each of the plurality of human faces having been formed by combining images of one or more faces illuminated by at least one of a plurality of illumination sources in a physical or virtual environment, each of the plurality of illumination sources having been positioned in a respective orientation of a plurality of orientations within the physical or virtual environment. The method can also include generating a prediction engine based on the plurality of images, the prediction engine configured to generate a predicted illumination profile from input image data representing an input human face.
[0007] In another general aspect, a computer program product includes a non-transitory storage medium, the computer program product including code that, when executed by processing circuitry of a computing device, causes the processing circuitry to perform a method. The method can include receiving image training data representing a plurality of images, each of the plurality of images including at least one of a plurality of human faces, each of the plurality of human faces formed by combining images of one or more faces illuminated by at least one of a plurality of illumination sources in a physical or virtual environment, each of the plurality of illumination sources being in a respective orientation of a plurality of orientations within the physical or virtual environment. The method can also include generating a prediction engine based on the plurality of images, the prediction engine configured to generate a predicted illumination profile from input image data, the input image data representing an input human face.
[0008] In another general aspect, an electronic device includes a memory and a control circuit coupled to the memory. The control circuit can be configured to receive image training data representing a plurality of images, each of the plurality of images including at least one of a plurality of human faces, each of the plurality of human faces formed by combining images of one or more faces illuminated by at least one of a plurality of illumination sources in a physical or virtual environment, each of the plurality of illumination sources being positioned in a respective orientation of a plurality of orientations within the physical or virtual environment. The control circuit can also be configured to generate a prediction engine based on the plurality of images, the prediction engine configured to generate a predicted illumination profile from input image data representing an input human face.
[0009] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a schematic diagram illustrating an example electronic environment in which the improved techniques described herein may be implemented.
[0011] Figure 2 is a flow chart illustrating an example method of estimating lighting from a portrait in accordance with disclosed embodiments.
[0012] Figure 3 is a schematic diagram illustrating an example system configured to estimate lighting from a portrait in accordance with disclosed embodiments.
[0013] Figure 4 It is an icon Figure 3 Schematic diagram of an example convolutional neural network (CNN) within the example system illustrated in .
[0014] Figure 5 It is an icon Figure 3 Schematic diagram of an example discriminator within the example system illustrated in .
[0015] Figure 6 is a schematic diagram illustrating examples of computer devices and mobile computer devices that can be used to implement the described techniques. DETAILED DESCRIPTION
[0016] One challenge in video applications such as augmented reality (AR) involves rendering synthetic objects into real scenes so that the object appears to be in the scene. One issue is matching the lighting of the real-world scene so that the rendered virtual content plausibly matches the appearance of the scene. For example, a lighting scheme can be designed for AR use cases with a world-facing camera, such as the rear camera of a mobile device, where one might want to render a synthetic object (such as a piece of furniture) into a live camera feed of a real-world scene.
[0017] However, such lighting schemes designed for world-facing cameras may differ from lighting schemes designed for front-facing cameras (e.g., for selfie images). For example, in portrait photos, lighting affects the look and feel of a given shot. Photographers light their subjects to convey a specific aesthetic and emotional tone. One method used by film visual effects practitioners to capture real-world lighting schemes involves recording the color and intensity of the omnidirectional lighting by shooting a mirror ball using multiple exposures. The result of this conventional approach is an HDR "image-based lighting" (IBL) environment used to realistically render virtual content into real-world photos.
[0018] AR shares the goal of realistically blending visual content with real-world imagery with film visual effects. However, in real-time AR, light measurements from specialized capture hardware are unavailable because acquisition is impractical for the average mobile phone or headset user. Similarly, for post-production visual effects in film, on-set light measurements are not always available, yet lighting designers must still use cues from the scene to infer lighting.
[0019] The challenge is therefore to determine the lighting solution for the front camera given an image of a face within a lighting environment. Some concepts have exploited strong geometric structure and reflectance priors from the face to solve for the lighting from the portrait. In the years since some researchers have introduced portrait backlighting, most such techniques have attempted to recover both the face geometry and a low-frequency approximation of the distant scene lighting, typically using up to a second-order spherical harmonic (SH) basis representation. The rationale for this approximation is that skin reflection is primarily diffuse (Lambertian) and therefore acts as a low-pass filter for the incident illumination. For diffusely reflecting materials, the irradiance is indeed very close to a nine-dimensional subspace that is well represented by this basis.
[0020] However, the lighting at capture time can reveal itself not only through the diffuse reflectance of the skin, but also through the direction and extent of cast shadows and the intensity and position of specular highlights. Inspired by these cues, some methods train neural networks to perform inverse lighting from portraits, thereby estimating the full range of HDR lighting without assuming any specific skin reflectance model. Such methods can produce higher-frequency lighting that can be used to convincingly render novel objects into real-world portraits, with applications in both visual effects and AR when offline lighting measurements are unavailable.
[0021] Conventional methods for estimating illumination given an LDR image of a face include generating such an illumination estimate based on a modeled bidirectional reflectance distribution function (BRDF) that defines the relationship between the incident light irradiance and the reflected light radiance on the face. The BDRF can be expressed as the ratio of the differential of the light irradiance, or the power per unit solid angle, about the incident light ray direction per unit projected area normal to the light ray to the differential of the outgoing light irradiance, or the power per unit surface area.
[0022] A technical problem with the conventional methods for estimating illumination from images of faces described above is that they base illumination estimation on a single reflectance function, such as the Lambertian or Phong models, which can limit the robustness of illumination estimation in the presence of extremely complex skin reflectances involving subsurface scattering, as well as roughness and Fresnel reflections, such as in the presence of varying natural skin color. Furthermore, the inherent ambiguity between light source intensity and surface albedo prevents straightforward recovery of the correct scale for illumination of subjects of various skin tones, even though a simple Lambertian model can accurately predict skin reflectance.
[0023] According to embodiments described herein, a solution to the aforementioned technical problem involves generating a lighting estimate from a single image of a face using a machine learning (ML) system using multiple bidirectional reflectance distribution functions (BRDFs) as a loss function. In some embodiments, the ML system is trained using images of faces formed using HDR lighting captured using an LDR lighting acquisition method. The solution involves training a lighting estimation model in a supervised manner using a dataset of portraits and their corresponding ground truth lighting. In the example dataset, 70 various subjects were captured in a lighting stage system illuminated by 331 directional light sources forming a sphere-based basis, allowing the captured subjects to be re-lit to appear as they would in any scene with image-based re-lighting. While several databases of real-world lighting environments captured using traditional HDR panoramic capture techniques are publicly available, the LDR lighting collection technique employed in some embodiments has been extended to instead capture approximately 1 million indoor and outdoor lighting environments, which are then converted to HDR using a novel non-negative least squares solver formulation before being used for re-lighting.
[0024] A technical advantage of the disclosed embodiments is that the ML system produces substantially the same lighting estimate at the correct scale or exposure value, regardless of the natural skin color of the face in the input image. Any attempt at lighting estimation is complicated by the inherent ambiguity between surface reflectance (albedo) and light source intensity. In other words, the shadow of a pixel is rendered unchanged if its albedo is halved when the light source intensity is doubled. The improved technique described above explicitly evaluates the performance of the model on a variety of subjects with different natural skin colors. For a given lighting condition, the improved technique is able to restore lighting at similar scales for a variety of subjects.
[0025] Furthermore, ML systems can estimate HDR lighting even when trained on LDR portrait images generated using HDR lighting. Several recent works have attempted to recover lighting from portraits without relying on a low-frequency illumination basis or BRDF models, including deep learning methods for arbitrary scenes and for outdoor scenes containing only the sun. The technical problem described in this paper outperforms both of these approaches and generalizes to arbitrary indoor or outdoor scenes. These models rely on computer-generated humanoid models as training data and therefore do not generalize to portraits of real, natural environments when inferring.
[0026] Figure 1 is a schematic diagram of an example electronic environment 100 in which the above-described technical solutions can be implemented. A computer 120 is configured to train and operate a prediction engine configured to estimate lighting from a portrait.
[0027] The computer 120 includes a network interface 122, one or more processing units 124, and a memory 126. The network interface 122 includes, for example, an Ethernet adapter, a token ring adapter, etc., for converting electronic and / or optical signals received from the network 150 into an electronic form for use by the computer 120. The set of processing units 124 includes one or more processing chips and / or components. The memory 126 includes volatile memory (e.g., RAM) and non-volatile memory, such as one or more ROMs, disk drives, solid-state drives, etc. The set of processing units 124 and the memory 126 together form a control circuit that is configured and arranged to perform the various methods and functions described herein.
[0028] In some implementations, one or more of the components of computer 120 can be or can include a processor (e.g., processing unit 124) configured to process instructions stored in memory 126. Figure 1 Examples of such instructions depicted in include image acquisition manager 130 and prediction engine training manager 140. Figure 1 As shown in , the memory 126 is configured to store various data, and the various data are described with respect to corresponding management using such data.
[0029] Image acquisition manager 130 is configured to receive image training data 131 and reference object data 136. In some embodiments, image acquisition manager 130 receives image training data 131 and reference object data 136 from display device 170 via network interface 122, i.e., via a network, such as network 190. In some embodiments, image acquisition manager 130 receives image training data 131 and reference object data 136 from local storage (e.g., a disk drive, a flash drive, an SSD, etc.).
[0030] In some embodiments, the image acquisition manager 130 is further configured to crop and resize facial images from the image training data 131 to produce standard-sized portraits. By cropping and resizing the images to a standard size, the training of the ML system is made more robust.
[0031] Image training data 131 represents a collection of portraits of faces captured using various lighting arrangements. In some embodiments, image training data 131 includes images or portraits of faces formed using HDR lighting recovered from low dynamic range (LDR) lighting environments. Figure 1As shown in , image training data 131 includes a plurality of images 132(1), ... 132(M), where M is the number of images in image training data 131. Each image, for example, image 132(1), includes light direction data 134(1) and pose data 135(1).
[0032] Light direction data 134(1...M) represents one of a specified number of directions (e.g., 331) from which a face is illuminated for a portrait used in image training data 131. In some embodiments, light direction data 134(1) includes polar and azimuthal angles, i.e., coordinates on a unit sphere. In some embodiments, light direction data 134(1) includes a triplet of direction cosines. In some embodiments, light direction data 134(1) includes a set of Euler angles. In the examples described above and in some embodiments, the angular configuration represented by light direction data 134(1) is one of the 331 configurations used to train the ML system.
[0033] The pose data 135 (1...M) represents one of a plurality (e.g., 9) of specified poses in which an image of a face is captured. In some embodiments, a pose includes a facial expression. In some embodiments, there is a fixed number of facial expressions (e.g., 3, 6, 9, 12, or more).
[0034] A four-dimensional reflectance field R(θ, φ, x, y) can represent a subject illuminated from any lighting direction (θ, φ) for each image pixel (x, y) according to the light direction data 134 (1…M). It has been demonstrated that taking the dot product of this reflectance field with an HDR lighting environment parameterized similarly by (θ, φ) relights the subject to appear as if they were in the scene. To capture the reflectance field of the subject, a computer-controlled sphere of white LED light sources is used with lights spaced 12° apart at the equator. In such embodiments, the reflectance field is formed from a collection of reflectance basis images, capturing the subject as each of the directional LED light sources is individually turned on within the sphere setup, one at a time. In some embodiments, these one-light-at-a-time (OLAT) images are captured for multiple camera viewpoints. In some embodiments, 331 OLAT images are captured for each subject using six color machine vision cameras with 12-megapixel resolution, placed 1.7 meters from the subject, although in some embodiments, these values, the number of OLAT images, and the type of camera used may vary. In some embodiments, the cameras are positioned generally in front of the subject, with five cameras with 35mm lenses capturing the subject's upper body from different angles, and an additional camera with a 50mm lens capturing close-up images of the face using a more rigid configuration.
[0035] In some embodiments, for 70 different subjects' reflection fields, each subject performing nine different facial expressions based on pose data 135 (1 ... M) and wearing different accessories, approximately 630 sets of OLAT sequences from six different camera viewpoints were obtained, totaling 3780 unique OLAT sequences. Other sets of OLAT sequences can be used. Subjects with a wide range of natural skin colors were captured.
[0036] Because acquiring a complete OLAT sequence for a subject takes some time, for example, about six seconds, there may be some slight subject motion between frames. In some embodiments, optical flow techniques are used to align the images, occasionally (for example, every 11 OLAT frames) interspersed with an additional "tracking" frame with uniform illumination to ensure that the brightness constrain for optical flow is met. This step can preserve the sharpness of image features when performing a relighting operation that linearly combines the aligned OLAT images.
[0037] To re-illuminate the subject using the captured reflective field, in some embodiments, a large database of HDR lighting environments in which no light sources are clipped is used. Although several such databases exist, containing approximately thousands of indoor panoramas or the upper hemispheres of outdoor panoramas, deep learning models are typically enhanced with larger amounts of training data. Therefore, approximately 1 million indoor and outdoor lighting environments were collected. In some embodiments, a mobile phone capture device was used to simultaneously capture an automatically exposed and white-balanced LDR video of a high-resolution background image and the corresponding LDR appearances of three spheres of different reflectivity (diffuse, specular, and matte silver with a rough mirror reflection). These three spheres reveal different clues about the scene lighting. The specular sphere reflects omnidirectional, high-frequency light, but since bright light sources are often clipped in a single exposure image, their intensity and color will be incorrect. In contrast, the diffuse sphere's near-Lambertian BRDF acts as a low-pass filter for the incident illumination, capturing a blurred but relatively complete record of the total scene radiance.
[0038] The embodiments herein enable having a true HDR record of the scene lighting for relighting the subject after explicitly lifting the three ball appearance into a near HDR lighting environment.
[0039] Reference object data 136 represents reference objects, such as balls of different reflectivity. Such reference objects are used to provide ground truth lighting in ML systems. Figure 1As shown in , the reference object data 136 includes a plurality of reference sets 137(1), ..., 137(N), where N is the number of HDR lighting environments considered. Each reference set in the reference sets 137(1 ...N), for example, reference set 137(1), includes BRDF data for specular 138(1), matte silver 139(1), and diffuse gray 141(1). In some embodiments, the BRDF data 138(1), 139(1), and 141(1) include arrays of BRDF values. In some embodiments, the BRDF data 138(1), 139(1), and 141(1) include sets of coefficients for SH expansion.
[0040] To train a model for estimating lighting from image training data 131 in a supervised manner, in some embodiments, portraits represented by image training data 131 are labeled with ground truth lighting, such as reference object data 136. In some embodiments, portraits using data-driven techniques for image-based relighting are synthesized, in some cases shown to produce photo-realistic relighting results for human faces, thereby appropriately capturing complex light transport phenomena for human skin and hair, such as subsurface and roughness scattering and Fresnel reflections. Such synthesis is in contrast to renderings of 3D models of faces, which often fail to represent these complex phenomena.
[0041] The prediction engine training manager 140 is configured to generate prediction engine data 150 representing the above-described ML system for estimating lighting from a portrait. Figure 1 As shown in , the prediction engine training manager 140 includes an encoder 142 , a decoder 143 , and a discriminator 144 .
[0042] The encoder 142 is configured to take as input a cropped portrait (i.e., light direction data 134 (1...L) from the image 132 (1...M) and from the image training data 131 to produce parameter values to be input into the fully connected layers in the decoder 143. The decoder 143 is configured to take as input the parameter values produced by the encoder 142 and produce lighting profile data 153 representing a predicted HDR lighting estimate. The discriminator 144 is configured to take as input the lighting profile data 153 and the reference object data 136 and produce cost function data 154 that is fed back into the decoder 143 to produce convolutional layer data 151 and blur pooling data 152. It should be noted that a cost function as used in an ML system is a function to be minimized by the ML system. The cost function in this case reflects, for example, the difference between a ground truth sphere image for multiple BRDFs and a corresponding network rendered sphere illuminated with the predicted lighting. Reference Figure 3 Describe further details about the ML system.
[0043] Returning to the reference object data 136, given the captured images of the three reflective spheres, possibly with clipped pixels, some embodiments solve for HDR lighting that can have plausibly produced the appearance of the three spheres. In some embodiments, the reflection fields for the diffuse and matte silver spheres can first be captured using the lighting stage system again. Some embodiments convert the reflection basis images into the same correlated radiometric space normalized based on the incident light source color. Some embodiments then project the reflection basis images into a mirror sphere mapping (Lambert azimuthal equal-area projection), thereby accumulating the energy from the input image for each new lighting direction (θ, φ) over a 32×32 image of the mirror sphere, as in some embodiments, or sliced into individual pixels R x,y (θ, φ).
[0044] For the lighting direction (θ, φ) in the captured mirror ball image without cropping for the color channel c, some embodiments recover the scene lighting L by simply scaling the mirror ball image pixel values by the inverse of the measured mirror ball reflectivity (82.7%) c (θ, φ). For lighting directions (θ, φ) with clipped pixels in the original mirror sphere image, some embodiments set the pixel value to 1.0, scale this by the inverse of the measured reflectivity, and form the scene lighting L c (θ, φ), and then use the non-negative least squares solver formula to solve for the residual lost light intensity U c (θ, φ). Given a BRDF index k (e.g., diffuse or matte silver), a color channel c, and a measured reflectance field R x,y,c,k The original image pixel value p of (θ, φ) x,y,c,k , due to the superposition principle of light, the following equation is satisfied:
[0045]
[0046] Equation (1) represents a set of m equations for each BRDF k and color channel c, equal to the number of spherical pixels in the reflected base image, where n is the unknown residual light intensity. For the unclipping illumination direction, U c (θ, φ) = 0. For each color channel, where km>n, the unknown U can be solved using non-negative least squares method c (θ, φ) values, thus ensuring that light is only added and not removed. In fact, some embodiments exclude the pruning pixel p from the solution x,y,c,k Some methods have recovered the clipped light intensity by comparing pixel values from a photographed diffuse sphere with the diffuse convolution of the cropped panorama, but these implementations are the first to use a photographic reflection basis and multiple BRDFs.
[0047] In some embodiments, it is observed that when solving for U c When each color channel is treated independently in (θ, φ), brightly hued red, green, and blue light sources are generally produced in geometrically nearby lighting directions, rather than a single light source with greater intensity in all three color channels. To recover results with more realistic neutrally colored light sources, some embodiments add cross-color channel regularization based on the following insight: the color of a photographed diffuse gray sphere reveals the average color balance of bright light sources in the scene (R avg , G avg , B avg ). Some embodiments add to the system of equations a new set of linear equations with weight λ=0.5:
[0048]
[0049] These regularization terms penalize the recovery of strongly hued light sources that have a different color balance than the target diffuse sphere. Some embodiments add regularization terms to encourage similar intensities for geometrically nearby lighting directions, but this will not necessarily prevent the recovery of strongly hued lights. Some embodiments use the Ceres solver to recover U c (θ, φ), thereby promoting the appearance of one million captured balls to HDR lighting. Since the LDR images from this video rate data collection method are 8-bit and encoded as sRGB, they may have a local natural color mapping, so some embodiments first linearize the ball images, assuming a gamma value γ = 2.2, as included in the linear system formula.
[0050] Using the captured reflectance fields and HDR-enhanced lighting for each subject, some embodiments generate re-illuminated portraits with ground truth lighting to use as training data. Some embodiments again transform the reflectance basis images into the same correlated radiometric space, calibrated based on the incident light source color. Since the lighting environment is represented as, for example, a 32×32 mirror sphere image, some embodiments project the reflectance fields onto this basis, again accumulating, as in some embodiments, the energy from the input image for each new lighting direction (θ, φ). Each new basis image is a linear combination of the original 331 OLAT images.
[0051] The lighting capture technique also produces a high-resolution background image corresponding to the appearance of the three balls. Since even arbitrary images contain useful clues for extracting lighting estimates, some embodiments composite the re-lit subject onto the background rather than onto a black frame as in some embodiments. Since the background image can be 8-bit sRGB, some embodiments crop and apply the transfer function to the re-lit subject before compositing. Since natural environment portraits may contain cropped pixels (especially for 8-bit live video for mobile AR), some embodiments discard HDR data for the re-lit subject to match the expected inference-time input.
[0052] Although background images can provide contextual cues to aid in lighting estimation, some embodiments compute a face bounding box for each input, and during training and inference, some embodiments crop each image to expand the bounding box by 25%. During training, some embodiments add slight crop region variations to randomly change their position and extent.
[0053] The components of the user device 120 (e.g., modules, processing unit 124) can be configured to operate based on one or more platforms (e.g., one or more similar or different platforms) that can include one or more types of hardware, software, firmware, operating systems, runtime libraries, etc. In some embodiments, the components of the computer 120 can be configured to operate within a cluster of devices (e.g., a server farm). In such embodiments, the functions and processing of the components of the computer 120 can be distributed to several devices of the cluster of devices.
[0054] The components of computer 120 can be or can include any type of hardware and / or software configured to process attributes. In some embodiments, Figure 1 One or more portions of the components shown in the components of the computer 120 in FIG. 1 can be or can include hardware-based modules (e.g., digital signal processors (DSPs), field programmable gate arrays (FPGAs), memory), firmware modules, and / or software-based modules (e.g., modules of computer code, sets of computer-readable instructions that can be executed at a computer). For example, in some embodiments, one or more portions of the components of the computer 120 can be or can include software modules configured for execution by at least one processor (not shown). In some embodiments, the functionality of the components can be included in Figure 1 Those different modules and / or different components shown in the drawings include combining the functions of two components illustrated in the drawings into a single component.
[0055] Although not shown, in some embodiments, the components (or portions thereof) of the computer 120 can be configured to operate, for example, within a data center (e.g., a cloud computing environment), a computer system, one or more server / host devices, and the like. In some embodiments, the components (or portions thereof) of the computer 120 can be configured to operate within a network. Thus, the components (or portions thereof) of the computer 120 can be configured to work within various types of network environments that can include one or more devices and / or one or more server devices. For example, a network can be or can include a local area network (LAN), a wide area network (WAN), and the like. A network can be or can include a wireless network and / or a wireless network implemented using, for example, a gateway device, a bridge, a switch, and the like. A network can include one or more segments and / or can have portions based on, for example, an Internet Protocol (IP) and / or a proprietary protocol. A network can include at least a portion of the Internet.
[0056] In some embodiments, one or more components of the computer 120 can be or include a processor configured to process instructions stored in a memory. For example, the image acquisition manager 130 (and / or portions thereof) and the predictive image training manager 140 (and / or portions thereof) can be a combination of a processor and memory configured to execute instructions associated with performing one or more functions.
[0057] In some embodiments, the memory 126 can be any type of memory, such as random access memory, disk drive memory, flash memory, and the like. In some embodiments, the memory 126 can be implemented as more than one memory component associated with a component of the VR server computer 120 (e.g., more than one RAM component or disk drive memory). In some embodiments, the memory 126 can be a database memory. In some embodiments, the memory 126 can be or can include non-local memory. For example, the memory 126 can be or can include memory shared by multiple devices (not shown). In some embodiments, the memory 126 can be associated with a server device (not shown) within a network and configured as a component of the service computer 120. Figure 1 As shown in , the memory 126 is configured to store various data, including image training data 131 , reference object data 136 , and prediction engine data 150 .
[0058] Figure 2 is a flow chart depicting an example method 200 for performing visual search according to the improved techniques described above. The method 200 may be executed by a combination of a computer 120 residing in the memory 126 and executed by a set of processing units 124. Figure 1 Describes the software structure implementation.
[0059] At 202, image acquisition manager 130 receives a plurality of images of a plurality of human faces in a physical environment (e.g., image training data 131). Each of the plurality of human faces is illuminated by at least one of a plurality of illumination sources oriented within the physical environment according to at least one of a plurality of orientations (e.g., light direction data 134(1...M)).
[0060] At 204, prediction engine training manager 140 generates a prediction engine (e.g., prediction engine data 150) configured to generate a predicted lighting profile based on a plurality of images of a plurality of human faces. The prediction engine is configured to generate the predicted lighting profile based on input image data. The input image data represents at least one human face. The prediction engine includes a cost function (e.g., discriminator 144 and cost function data 154) based on a plurality of bidirectional reflectance distribution functions (BRDFs) corresponding to each of reference objects (e.g., reference object data 136). The predicted lighting profile represents the spatial distribution of illumination incident on the subject of the portrait. An example of the predicted lighting represents coefficients of a spherical harmonic expansion of an illumination function comprising an angle. Another example of the predicted lighting represents a grid comprising pixels, each pixel having a value of the illumination function for a solid angle.
[0061] Figure 3 is a schematic diagram illustrating an example ML system 300 configured to estimate lighting from a portrait. Figure 3 As shown in FIG, the ML system 300 includes a generator network 314 and an auxiliary adversarial discriminator 312. The input to the generator network 314 is an sRGB encoded LDR image, e.g., an LDR portrait 302, with a crop 306 of the face region of each image detected by the face detector 304, which is resized to the input resolution of 256×256 and normalized to the range of [-0.5, 0.5]. Figure 3 As shown in , the generator network 314 has an encoder / decoder architecture including an encoder 142 and a decoder 143, with a latent vector representation of log-space HDR lighting of size 1024 at the bottleneck. In some embodiments, the encoder 142 and the decoder 143 are implemented as convolutional neural networks (CNNs). The final output of the generator network 314 includes a 32×32 HDR image representing a mirrored sphere with omnidirectional lighting in log space. Figure 4 Further details about the encoder 142 and decoder 143 are shown; see Figure 5 Further details regarding the auxiliary adversarial discriminator 312 are shown.
[0062] Figure 4 is a schematic diagram illustrating example details for encoder 142 and decoder 143. Figure 4As shown in , the encoder 142 includes five 3×3 convolutions, each followed by a blur pooling operation with successive filter depths of 16, 32, 64, 128, and 256, followed by a final convolution with a filter size of 8×8 and a depth of 256, and finally a fully connected layer. The decoder 143 includes three sets of 3×3 convolutions with filter depths of 64, 32, and 16, each followed by a bilinear upsampling operation.
[0063] Figure 5 is a schematic diagram illustrating an example auxiliary adversarial discriminator 312. The auxiliary adversarial discriminator 312 is configured to provide an adversarial loss term, thereby enforcing the estimation of plausible high-frequency lighting. Figure 5 As shown in [ ], the auxiliary adversarial discriminator 312 takes as input the ground truth pruned image and the predicted illumination from the main model and attempts to discriminate between real examples and generated ones. The discriminator has an encoder consisting of three 3×3 convolutions, each followed by a max pooling operation, with successive filter depths of 64, 128, and 256, followed by a fully connected layer of size 1024 before the final output layer. Because the main network's decoder includes several upsampling operations, the network implicitly learns information at multiple scales. Some embodiments leverage this multi-scale output to provide input to the discriminator, not only the full-resolution 32×32 pruned illumination image, but also illumination images at each scale: 4×4, 8×8, and 16×16, using the multi-scale gradient technique of the MSG-GAN. Because the lower-resolution feature maps produced by the generator network have more than three channels, some embodiments add convolution operations at each scale as additional branches of the network, producing multiple scales of the 3-channel illumination image to feed to the discriminator.
[0064] return Figure 3 , the generator network 314 and the auxiliary adversarial discriminator 312 use various cost functions to build a prediction engine and discriminate between real and generated lighting estimates. Some embodiments describe methods for training a network to estimate HDR lighting from unconstrained images. The method minimizes the loss between a ground truth sphere image I and a corresponding network-rendered sphere image I using the predicted illumination for multiple BRDFs. Some embodiments use this technique to train the model for backlighting from portraits, thereby relying on these sphere renderings to learn useful lighting for rendering virtual objects of various BRDFs. Some embodiments use image-based relighting and the captured reflectance field and color channel c of each sphere for BRDF index k (specular, matte silver, or diffuse) to generate sphere renderings in the network in As the intensity of light for the direction (θ, φ):
[0065]
[0066] Since in some embodiments, the network similarly outputs pixels with values Q c The logarithmic space image Q of the HDR lighting of (θ, φ), so the sphere image is rendered as
[0067]
[0068] Using a binary mask that masks the corners of each ball γ = 2.2 for gamma encoding, λ as an optional weight for each BRDF k and a soft pruning function Λ as in some embodiments, which transforms the ground truth image I k with network rendered images The final LDR image reconstruction loss L for comparison rec yes
[0069]
[0070] The binary operator ⊙ represents element-wise multiplication.
[0071] Instead of using the LDR sphere image captured in video rate data collection as the reference image I k Instead, some embodiments render the sphere using HDR lighting recovered from a linear solver (e.g., Equation (1)), gamma encoding the rendering with γ = 2.2. This ensures that the same lighting is used to render the "ground truth" sphere as the input portrait, preventing residual errors from HDR lighting recovery from being propagated into the model training phase.
[0072] Some embodiments finally add additional convolution branches to convert the decoder's multi-scale feature maps into 3-channel images representing log-space HDR lighting at continuous scales. Some embodiments then extend the rendering loss function of some embodiments (Equation (6)) to the multi-scale domain, rendering specular, matte silver, and diffuse spheres during training with sizes 4×4, 8×8, 16×16, and 32×32. Using the scale index denoted by s and as λ s For each optional weight, the multi-scale image reconstruction loss is written as
[0073]
[0074] Recent work in unconstrained illumination estimation has shown that the adversarial loss term improves the recovery of high-frequency information compared to using only the image reconstruction loss. Therefore, some embodiments utilize weights λ as in some embodiments advAdding an adversarial loss term. However, in contrast to this technique, some embodiments use a multi-scale GAN architecture that flows gradients from the discriminator to the generator network at multiple scales, thereby providing the discriminator with real and generated cropped mirror ball images of different sizes.
[0075] Some embodiments use Tensorflow and the ADAM optimizer with β1=0.9, β2=0.999, a learning rate of 0.00015 for the generator network, and a 100x lower learning rate for the discriminator network as is common, alternating between training the generator and the discriminator. Some embodiments set λ for specular, diffuse, and matte silver BRDFs separately. k =0.2, 0.6, 0.2, set λ s = 1 to weight all image scales equally, set λ adv =0.004, and a batch size of 32 is used. Since the number of lighting environments can be orders of magnitude larger than the number of subjects, for some embodiments, stopping early at 1.2 rounds prevents overfitting to the subjects in the training set. Some embodiments use the ReLU activation function for the generator network and the ELU activation function for the discriminator. To augment the dataset, some embodiments flip both the input image and the lighting environment across the vertical axis. Some embodiments augment the dataset with a slight image rotation (+ / - 15 degrees) of the input image in the image plane.
[0076] Some embodiments split the 70 subjects into two groups: 63 for training and 7 for evaluation, ensuring that all expressions and camera views for a given subject belong to the same subset. Some embodiments include manually selecting 7 subjects to include a variety of natural skin colors. Overall, for each of the 1 million lighting environments, some embodiments include randomly selecting 8 OLAT sequences from the training set to relight (across subjects, facial expressions, and camera views), thereby generating a training dataset of 8 million portraits with ground truth lighting. Using the same approach, some embodiments capture lighting environments in both indoor and outdoor locations not seen in training for evaluation, pairing only these with the evaluation subjects.
[0077] An accurately estimated lighting should correctly render objects with any reflectance properties, so the performance of the model is better than that of the model using L rec This metric compares the appearance of three spheres (diffuse, matte silver, and specular) as rendered with ground truth and estimated lighting.
[0078] For the LDR image reconstruction loss, the model outperforms some implementations for diffuse and matte silver spheres. However, some implementations can outperform this implementation for mirror spheres. The second-order SH approximation of the ground truth illumination can outperform the LDR implementation for diffuse spheres. rec This model is not suitable for rendering Lambertian materials because the low frequency representation of the lighting is sufficient. However, this implementation is able to outperform the L for both matte silver and mirror sphere with non-Lambertian BRDFs. rec This shows that the lighting produced by this implementation is better suited for rendering various materials.
[0079] Some embodiments add a loss function based on cross-subject consistency based on the difference between a first predicted lighting profile from a first person's face and a second predicted lighting profile from a second person's face. Such a loss function can provide a measure of lighting consistency for various natural skin colors and head poses.
[0080] Figure 6 1 illustrates an example of a general purpose computer device 600 and a general purpose mobile computer device 650 that can be used with the techniques described herein. Figure 1 and Figure 2 An example configuration of computer 120 is shown.
[0081] like Figure 6 As shown, computing device 600 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Computing device 650 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are meant to be exemplary only and are not meant to limit the embodiments of the invention described and / or claimed in this document.
[0082] Computing device 600 includes a processor 602, a memory 604, a storage device 606, a high-speed controller 608 connected to memory 604 and a high-speed expansion port 610, and a low-speed controller 612 connected to a low-speed expansion port 614 and storage device 606. Each of components 602, 604, 606, 608, 610, and 612 is interconnected using various buses and can be mounted on a common motherboard or otherwise installed where appropriate. Processor 602 is capable of processing instructions for execution within computing device 600, including instructions stored in memory 604 or on storage device 606 to display graphical information such as a GUI on an internal input / output device coupled to display 616 of high-speed controller 608. In other embodiments, multiple processors and / or multiple buses can be used, as appropriate, along with multiple memories and multiple types of memory. In addition, multiple computing devices 600 can be connected, with each device providing a portion of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).
[0083] Memory 604 stores information within computing device 600. In one embodiment, memory 604 is one or more volatile memory units. In another embodiment, memory 604 is one or more non-volatile memory units. Memory 604 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.
[0084] The storage device 606 can provide mass storage for the computing device 600. In one embodiment, the storage device 606 can be or include a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device or a tape device, a flash memory or other similar solid-state memory device, or an array of devices, including devices in a storage area network or other configuration. A computer program product can be tangibly embodied in an information carrier. The computer program product can also include instructions that, when executed, perform one or more methods such as those described above. The information carrier is a computer or machine-readable medium, such as the memory 604, the storage device 606, or a memory on the processor 602.
[0085] The high-speed controller 608 manages bandwidth-intensive operations of the computing device 600, while the low-speed controller 612 manages less bandwidth-intensive operations. This allocation of functions is merely exemplary. In one embodiment, the high-speed controller 608 is coupled to the memory 604, the display 616 (e.g., via a graphics processor or accelerator), and to the high-speed expansion port 610, which can accept various expansion cards (not shown). In an embodiment, the low-speed controller 612 is coupled to the storage device 506 and the low-speed expansion port 614. The low-speed expansion port (which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet)) can be coupled to one or more input / output devices such as a keyboard, pointing device, scanner, or a network device such as a switch or router, for example, via a network adapter.
[0086] Computing device 600 can be implemented in many different forms, as shown. For example, it can be implemented as a standard server 620 or implemented multiple times in a group of such servers. It can also be implemented as part of a rack server system 624. In addition, it can be implemented in a personal computer such as laptop computer 622. Alternatively, components from computing device 600 can be combined with other components in a mobile device (not shown) such as device 650. Each of such devices can contain one or more of computing devices 600, 650, and the entire system can be composed of multiple computing devices 600, 650 communicating with each other.
[0087] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations of one or more computer programs executable and / or interpretable on a programmable system that includes at least one programmable processor (which can be special purpose or general purpose) coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to the storage system, at least one input device, and at least one output device.
[0088] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the terms "machine-readable medium," "computer-readable medium," and "machine-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0089] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input from the user can be received in any form, including sound, voice, or tactile input.
[0090] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface or a web browser through which a user can interact with the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.
[0091] A computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0092] A number of implementations have been described, however, it will be understood that various modifications can be made without departing from the spirit and scope of the disclosure.
[0093] It will also be understood that when an element is referred to as being on another element, being connected to another element, being electrically connected to another element, being coupled to another element, or being electrically coupled to another element, it can be directly on another element, being connected or being coupled to another element, or one or more intermediary elements can be present. On the contrary, when an element is referred to as being directly on another element, being directly connected to or being directly coupled to another element, there is no intermediary element. Although the term directly on, being directly connected to or being directly coupled to can not be used in the entire specific embodiment, the element that is shown as being directly on, being directly connected to or being directly coupled can be referred to as such. The claims of the present application can be modified to narrate the exemplary relationship described in this specification or shown in each figure.
[0094] Although certain features of the described embodiments have been described as described herein, many modifications, substitutions, variations and equivalents will now occur to those skilled in the art. Therefore, it should be understood that the appended claims are intended to encompass all such modifications and variations as fall within the scope of the embodiments. It should be understood that they have been presented by way of example only and not limitation, and various changes in form and detail may be made. Any part of the apparatus and / or method described herein may be combined in any combination except for mutually exclusive combinations. The factual manner described herein can include various combinations and / or sub-combinations of the functions, components and / or features of the different embodiments described.
[0095] Additionally, the logic flows depicted in the figures do not require the specific order shown or sequential order to achieve the desired results. Additionally, other steps may be provided, or steps may be eliminated from the described flows, and other components may be added to or removed from the described systems. Accordingly, other implementations are within the scope of the following claims.
[0096] In the following, some examples are described.
[0097] Example 1: A method comprising:
[0098] receiving image training data representing a plurality of images, each image of the plurality of images including at least one human face of a plurality of human faces, each human face of the plurality of faces having been formed by combining images of one or more faces as illuminated by at least one of a plurality of illumination sources in a physical or virtual environment, each of the plurality of illumination sources having been in a respective orientation of a plurality of orientations within the physical or virtual environment; and
[0099] A prediction engine is generated based on the plurality of images, the prediction engine being configured to generate a predicted lighting profile from input image data, the input image data representing an input human face.
[0100] Example 2: The method according to Example 1, further comprising:
[0101] The images of the one or more human faces illuminated by the at least one of a plurality of illumination sources are combined to synthetically render each of the plurality of human faces to appear illuminated by a high dynamic range (HDR) lighting environment.
[0102] Example 3: The method of Example 2, wherein combining the images comprises:
[0103] The HDR lighting environment is generated based on a low dynamic range (LDR) image of a set of reference objects, each reference object in the set of reference objects having a corresponding bidirectional reflectance distribution function (BRDF).
[0104] Example 4: The method according to Example 3, wherein the reference object set includes a mirror ball, a matte silver ball, and a gray diffuse reflection ball.
[0105] Example 5: The method of Example 1, wherein generating the prediction engine comprises:
[0106] performing differentiable rendering on a set of reference objects using the predicted lighting profile to generate a rendered image of the set of reference objects, each reference object in the set of reference objects having a corresponding bidirectional reflectance distribution function (BRDF); and
[0107] A difference between the rendered image of the set of reference objects and a ground truth image of the set of reference objects is generated as a cost function of the prediction engine.
[0108] Example 6: The method of Example 5, wherein the cost function comprises a BRDF-weighted L1 loss on the rendered image of the set of reference objects.
[0109] Example 7: The method of Example 5, wherein the cost function is a first cost function, and
[0110] Wherein, the prediction engine includes a second cost function, which is an adversarial loss function based on high-frequency mirror reflections from the mirror ball.
[0111] Example 8: The method of Example 5, wherein the differentiable rendering is performed using image-based relighting (IBRL) to produce a high dynamic range (HDR) illuminated image.
[0112] Example 9: The method of Example 5, wherein the cost function is a first cost function, and
[0113] Wherein, the prediction engine includes a second cost function, which is a cross-subject consistency based loss function based on the difference between a first predicted lighting profile from a first human face and a second predicted lighting profile from a second human face.
[0114] Example 10: The method of Example 1, wherein generating the prediction engine comprises:
[0115] During the generation of the prediction engine, a facial landmark detection operation is performed on image training data to produce facial landmark identifiers that identify facial landmarks.
[0116] Example 11: The method of Example 1, wherein generating the prediction engine comprises:
[0117] Each pixel of the image of a face of the plurality of faces is projected into a common UV space.
[0118] Example 12: The method of Example 1, wherein each image of the plurality of images is gamma encoded.
[0119] Example 13: A computer program product comprising a non-transitory storage medium, the computer program product comprising code that, when executed by a processing circuit of a computer, causes the processing circuit to perform a method comprising:
[0120] receiving image training data representing a plurality of images, each image of the plurality of images including at least one human face of a plurality of human faces, each human face of the plurality of faces having been formed by combining images of one or more faces illuminated by at least one of a plurality of illumination sources in a physical or virtual environment, each illumination source of the plurality of illumination sources having been in a respective orientation of a plurality of orientations within the physical or virtual environment; and
[0121] A prediction engine is generated based on the plurality of images, the prediction engine being configured to generate a predicted lighting profile from input image data, the input image data representing an input human face.
[0122] Example 14: The computer program product of Example 13, wherein generating the prediction engine comprises:
[0123] The images of the one or more human faces as illuminated by the at least one of a plurality of illumination sources are combined to synthetically render each of the plurality of human faces to appear illuminated by a high dynamic range (HDR) lighting environment.
[0124] Example 15: The computer program product of Example 14, wherein combining the images comprises:
[0125] The HDR lighting environment is generated based on a low dynamic range (LDR) image in a reference object set, where each reference object in the reference object set has a corresponding bidirectional reflectance distribution function (BRDF).
[0126] Example 16: The computer program product of Example 15, wherein the set of reference objects includes a mirrored sphere, a matte silver sphere, and a gray diffuse sphere.
[0127] Example 17: The computer program product of Example 13, wherein generating the prediction engine comprises:
[0128] performing differentiable rendering of a set of reference objects using the predicted lighting profile to produce a rendered image of the set of reference objects, each reference object in the set of reference objects having a corresponding bidirectional reflectance distribution function (BRDF); and
[0129] A difference between the rendered image in the set of reference objects and a ground truth image in the set of reference objects is generated as a cost function for the prediction engine.
[0130] Example 18: The computer program product of Example 14, wherein the cost function comprises a BRDF-weighted L1 loss on the rendered images in the set of reference objects.
[0131] Example 19: The computer program product of Example 18, wherein the differentiable rendering is performed using Image-Based Relighting (IBRL) to produce a High Dynamic Range (HDR) illuminated image.
[0132] Example 20: An electronic device comprising:
[0133] Memory; and
[0134] a processing circuit coupled to the memory, the processing circuit being configured to:
[0135] image training data representing a plurality of images, each image in the plurality of images including at least one human face from a plurality of human faces, each human face in the plurality of faces having been formed by combining images of one or more faces illuminated by at least one illumination source from a plurality of illumination sources in a physical or virtual environment, each illumination source from the plurality of illumination sources having been in a respective orientation from a plurality of orientations within the physical or virtual environment; and
[0136] A prediction engine is generated based on the plurality of images, the prediction engine being configured to generate a predicted lighting profile from input image data, the input image data representing an input human face.
Claims
1. A method for determining lighting from a portrait, comprising: receiving image training data representing a plurality of images, each image in the plurality of images including at least one human face from a plurality of human faces, each human face in the plurality of faces having been formed by combining images of one or more faces illuminated by at least one illumination source from a plurality of illumination sources in a physical or virtual environment, each illumination source from the plurality of illumination sources having been oriented in a respective one of a plurality of orientations within the physical or virtual environment; as well as generating a prediction engine based on the plurality of images, the prediction engine configured to generate a predicted lighting profile from input image data, the input image data representing an input human face, Wherein the prediction engine comprises a first cost function based on a difference between a first predicted lighting profile from a representation of a first human face and a second predicted lighting profile from a representation of a second human face.
2. The method according to claim 1, further comprising: The images of the one or more faces illuminated by the at least one of a plurality of illumination sources are combined to synthetically render each of the plurality of faces to appear illuminated by a high dynamic range (HDR) lighting environment.
3. The method according to claim 2, wherein: Combining the images comprises: The HDR lighting environment is generated based on a low dynamic range (LDR) image in a reference object set, where each reference object in the reference object set has a corresponding bidirectional reflectance distribution function (BRDF).
4. The method according to claim 3, wherein: The reference object set includes a mirror ball, a matte silver ball, and a gray diffuse reflection ball.
5. The method according to claim 1, wherein Generating the prediction engine includes: performing differentiable rendering of a set of reference objects using the predicted lighting profile to produce a rendered image of the set of reference objects, each reference object in the set of reference objects having a corresponding bidirectional reflectance distribution function (BRDF); and Wherein, the prediction engine further includes a second cost function, the second cost function being based on a difference between the rendered image in the set of reference objects and a ground truth image in the set of reference objects.
6. The method according to claim 5, wherein: The second cost function comprises a BRDF-weighted L1 loss on the rendered images in the set of reference objects.
7. The method according to claim 5, wherein: The prediction engine includes another cost function that is an adversarial loss function based on high-frequency mirror reflections from a mirror ball.
8. The method according to claim 5, wherein The differentiable rendering is performed using image-based relighting (IBRL) to produce high dynamic range (HDR) illuminated images.
9. The method according to claim 1, wherein: Generating the prediction engine includes: During generation of the prediction engine, a facial landmark detection operation is performed on image training data to produce facial landmark identifiers that identify facial landmarks.
10. The method according to claim 1, wherein Generating the prediction engine includes: Each pixel of the image of a face of the plurality of faces is projected into a common UV space.
11. The method according to claim 1, wherein Each image of the plurality of images is gamma encoded.
12. A computer program product comprising a non-transitory storage medium, the computer program product comprising code that, when executed by a processing circuit of a server computing device, causes the processing circuit to perform a method comprising: receiving image training data representing a plurality of images, each image in the plurality of images including at least one human face from a plurality of human faces, each human face in the plurality of faces having been formed by combining images of one or more faces illuminated by at least one illumination source from a plurality of illumination sources in a physical or virtual environment, each illumination source from the plurality of illumination sources having been oriented in a respective one of a plurality of orientations within the physical or virtual environment; as well as generating a prediction engine based on the plurality of images, the prediction engine configured to generate a predicted lighting profile from input image data, the input image data representing an input human face, Wherein the prediction engine comprises a first cost function based on a difference between a first predicted lighting profile from a representation of a first human face and a second predicted lighting profile from a representation of a second human face.
13. The computer program product of claim 12, wherein: Generating the prediction engine includes: The images of the one or more faces illuminated by the at least one of a plurality of illumination sources are combined to synthetically render each of the plurality of faces to appear illuminated by a high dynamic range (HDR) lighting environment.
14. The computer program product of claim 13, wherein: Combining the images comprises: The HDR lighting environment is generated based on a low dynamic range (LDR) image in a reference object set, where each reference object in the reference object set has a corresponding bidirectional reflectance distribution function (BRDF).
15. The computer program product of claim 14, wherein: The reference object set includes a mirror ball, a matte silver ball, and a gray diffuse reflection ball.
16. The computer program product of claim 12, wherein: Generating the prediction engine includes: performing differentiable rendering of a set of reference objects using the predicted lighting profile to produce a rendered image of the set of reference objects, each reference object in the set of reference objects having a corresponding bidirectional reflectance distribution function (BRDF); and Wherein, the prediction engine further includes a second cost function, the second cost function being based on a difference between the rendered image in the set of reference objects and a ground truth image in the set of reference objects.
17. The computer program product of claim 16, wherein: The second cost function comprises a BRDF-weighted L1 loss on the rendered images in the set of reference objects.
18. The computer program product of claim 16, wherein: The differentiable rendering is performed using image-based relighting (IBRL) to produce high dynamic range (HDR) illuminated images.
19. An electronic device, comprising: Memory; as well as a control circuit coupled to the memory, the control circuit being configured to: receiving image training data representing a plurality of images, each image in the plurality of images including at least one human face from a plurality of human faces, each human face in the plurality of faces having been formed by combining images of one or more faces illuminated by at least one illumination source from a plurality of illumination sources in a physical or virtual environment, each illumination source from the plurality of illumination sources having been oriented in a respective one of a plurality of orientations within the physical or virtual environment; as well as generating a prediction engine based on the plurality of images, the prediction engine configured to generate a predicted lighting profile from input image data, the input image data representing an input human face, Wherein the prediction engine comprises a first cost function based on a difference between a first predicted lighting profile from a representation of a first human face and a second predicted lighting profile from a representation of a second human face.