Enhanced photo relighting based on machine learning models
By applying machine learning models to adjust image lighting on mobile devices, the challenge of taking professional-grade portrait photos has been solved, achieving high-quality lighting effects in real time or post-processing and enhancing the visual expressiveness of images.
Patent Information
- Application Number
- CN202180067003.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-30
- Filing Date
- 2021-05-17
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-05-17
AI Technical Summary
Mobile device users struggle to take professional-grade portrait photos due to a lack of professional lighting knowledge and resources. Existing technologies also make it difficult to adjust the lighting effects of images in real time or in post-processing to achieve high-quality visual effects.
Employing a machine learning-based convolutional neural network model, the system receives input images and lighting model data, predicts and generates adjusted output images, and automatically or interactively enhances image lighting by combining surface geometry and ambient light estimation.
It enables real-time or post-processing adjustment of image lighting on mobile devices, improving image quality, enhancing user creative flexibility and the visual effects of photos, and is suitable for various lighting conditions and scenarios.
Smart Images

Figure CN116324899B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 085,529, filed on September 30, 2020, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to methods and apparatus for relighting photographs to achieve enhancement, and more particularly to methods and apparatus for relighting photographs based on machine learning models. Background Technology
[0004] Many modern computing devices, including mobile phones, personal computers, and tablets, include image capture devices such as still and / or video cameras. Image capture devices can capture images, such as images of people, animals, landscapes, and / or objects.
[0005] Some image capture devices and / or computing devices can correct or otherwise modify captured images. For example, some image capture devices can provide "red-eye" correction, which removes artifacts that may exist in images captured using strong light, such as flash illumination, such as the appearance of red eyes in people and animals. After correction, the corrected image can be saved, displayed, transmitted, printed on paper, and / or used in other ways.
[0006] Professional photographers (such as portrait photographers) utilize the properties of light on their subjects to create striking photographs. Such photographers often use specialized equipment (such as external camera flashes and reflectors) to position the lighting and illuminate their subjects to achieve a professional look. In some instances, such activities take place in controlled studio environments and involve expertise in equipment, lighting, and so on.
[0007] Mobile phone users typically lack access to specialized portrait studio resources or are unaware of how to use them. However, users may prefer the professional, high-quality results from experienced portrait photographers. Summary of the Invention
[0008] In one aspect, image capture devices can be configured to translate a professional photographer's understanding of light and use of external camera lighting into computer-implemented methods. Powered by machine learning components, image capture devices can be configured to enable users to create attractive lighting for portraits or other types of images.
[0009] In some aspects, the mobile device can be configured with these features so that images can be enhanced in real-time. In some cases, the mobile device can automatically enhance images. In other aspects, the mobile phone user can non-destructively enhance images to match their preferences. Furthermore, pre-existing images in a user's image library can be enhanced, for example, based on the techniques described herein.
[0010] In one aspect, a computer-implemented method is provided. A computing device applies a geometry model to an input image to determine, based on surface geometry of an object, a surface orientation map indicative of a distribution of illumination on the object in the input image. The computing device applies an ambient light estimation model to the input image to determine a direction of synthetic illumination to be applied to the input image to enhance at least a portion of the input image. The computing device applies, based on the surface orientation map and the direction of synthetic illumination, a light energy model to determine a quotient image indicative of an amount of light energy to be applied to each pixel of the input image. The computing device enhances the portion of the input image based on the quotient image.
[0011] In another aspect, a computing device is provided. The computing device includes one or more processors and a data storage device. The data storage device has stored thereon computer executable instructions that, when executed by the one or more processors, cause the computing device to perform functions. The functions include: applying a geometry model to an input image to determine, based on surface geometry of an object, a surface orientation map indicative of a distribution of illumination on the object in the input image; applying an ambient light estimation model to the input image to determine a direction of synthetic illumination to be applied to the input image to enhance at least a portion of the input image; applying, based on the surface orientation map and the direction of synthetic illumination, a light energy model to determine a quotient image indicative of an amount of light energy to be applied to each pixel of the input image; and enhancing the portion of the input image based on the quotient image.
[0012] In another aspect, an article of manufacture is provided. The article of manufacture includes one or more computer readable media having stored thereon computer readable instructions that, when executed by one or more processors of a computing device, cause the computing device to perform functions. The functions include: applying a geometry model to an input image to determine, based on surface geometry of an object, a surface orientation map indicative of a distribution of illumination on the object in the input image; applying an ambient light estimation model to the input image to determine a direction of synthetic illumination to be applied to the input image to enhance at least a portion of the input image; applying, based on the surface orientation map and the direction of synthetic illumination, a light energy model to determine a quotient image indicative of an amount of light energy to be applied to each pixel of the input image; and enhancing the portion of the input image based on the quotient image.
[0013] In another aspect, a system is provided. The system includes: means for applying a geometry model to an input image to determine a surface orientation map indicative of a distribution of illumination on an object in the input image based on surface geometry of the object; means for applying an ambient light estimation model to the input image to determine a direction of synthetic illumination to be applied to the input image to enhance at least a portion of the input image; means for applying a light energy model based on the surface orientation map and the direction of synthetic illumination to determine a quotient image indicative of an amount of light energy to be applied to each pixel of the input image; and means for enhancing the portion of the input image based on the quotient image.
[0014] The foregoing overview is illustrative only and is not intended to serve as a limitation. Further aspects, embodiments, and features will become apparent from the following detailed description and from the drawings. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 An image 100 with different lighting is shown according to example embodiments.
[0016] Figure 2 is a diagram depicting a relighting network for enhancing lighting of an image according to example embodiments.
[0017] Figure 3 is a diagram depicting an example network for predicting a surface orientation map according to example embodiments.
[0018] Figure 4 is an example architecture of a neural network for predicting a surface orientation map according to example embodiments.
[0019] Figure 5 An example image from a photographed reflective field is shown according to example embodiments.
[0020] Figure 6 An example method of recommending light directions is shown according to example embodiments.
[0021] Figure 7 An example image with estimated lighting is shown according to example embodiments.
[0022] Figure 8A An example network architecture for determining a light visibility map is shown according to example embodiments.
[0023] Figure 8B An example image showing the effect of highlighting surface geometry in a photo lighting is shown according to example embodiments.
[0024] Figure 9 is a diagram depicting an example network for predicting a quotient image according to example embodiments.
[0025] Figure 10 is an example architecture of a neural network that predicts merchant images according to example embodiments.
[0026] Figure 11 is an example network illustrating post-processing techniques according to example embodiments.
[0027] Figure 12 is an example interactive graphical user interface according to example embodiments.
[0028] Figure 13 illustrates example input-output image pairs based on surface orientation maps according to example embodiments.
[0029] Figure 14 illustrates an example relighting pipeline according to example embodiments.
[0030] Figure 15 illustrates an example input image with predicted facial light estimates according to example embodiments.
[0031] Figure 16 is a diagram illustrating training and inference stages of a machine learning model according to example embodiments.
[0032] Figure 17 depicts a distributed computing architecture according to example embodiments.
[0033] Figure 18 is a block diagram of a computing device according to example embodiments.
[0034] Figure 19 depicts a network of computing clusters arranged as a cloud-based server system according to example embodiments.
[0035] Figure 20 is a flowchart of a method according to example embodiments. DETAILED DESCRIPTION
[0036] This application relates to enhancing images of objects (e.g., objects depicting human faces) using machine learning techniques, such as but not limited to neural network techniques. When a mobile computing device user takes an image of an object, such as a person, the resulting image can not always have ideal lighting. For example, the image can be too bright or too dark, the light can be coming from an undesirable direction, or the lighting can include different colors that give the image an undesirable color cast. Furthermore, even if the image does have desirable lighting at one time, the user can want to change the lighting at a later time. As such, there arises a technical problem related to image processing that involves adjusting the lighting of an obtained image.
[0037] To allow a user to control the lighting of an image, particularly an image of a human face, the techniques described herein apply a model based on a convolutional neural network to adjust the lighting of the image. The techniques described herein include receiving an input image and data regarding a particular lighting model to be applied to the input image, using a convolutional neural network to predict an output image that applies the data regarding the particular lighting model to be applied to the input image, and generating an output based on the output image. The input image and the output image can be high resolution images, such as images of millions of pixels in size captured by a camera of a mobile computing device. The convolutional neural network is well suited to handle input images captured under a variety of natural and artificial lighting conditions. In some examples, a trained model of the convolutional neural network can work on a variety of computing devices, including but not limited to mobile computing devices (e.g., smartphones, tablet computers, cell phones, laptop computers), stationary computing devices (e.g., desktop computers), and server computing devices. The convolutional neural network can apply a particular lighting model to an input image, thereby adjusting the lighting of the input image and addressing the technical problem of adjusting the lighting of an obtained image.
[0038] A neural network, such as a convolutional neural network, can be trained using a training dataset of images to perform one or more aspects described herein. In some examples, the neural network can be arranged as an encoder / decoder neural network.
[0039] While the examples described herein relate to determining and applying a lighting model for images of objects having human faces, a neural network can be trained to determine and apply a lighting model for images of other objects that similarly reflect light as a human face. Human faces generally diffuse reflect light, but can also include some specular highlights due to direct reflection of light. For example, specular highlights can result from direct light reflection from the surface of an eye, glasses, jewelry, etc. In many images of human faces, the area of such specular highlights is relatively small compared to the area of the facial surfaces that diffuse reflect light. Thus, a neural network can be trained to apply a lighting model to images of other objects that diffuse reflect light, where these diffusely reflecting objects can have some relatively small specular highlights (e.g., a tomato or a wall painted with a matte finish). The images in the training dataset can illustrate one or more particular objects using lighting provided under a variety of different conditions (e.g., lighting provided from different directions, lighting provided at different intensities (e.g., brighter and dimmer lighting), lighting provided with different colors of light sources, lighting provided with different numbers of light sources, etc.).
[0040] The trained neural network can process an input image to predict ambient lighting. An optimal light direction to compensate for existing portrait lighting of the input image can be recommended based on the predicted ambient lighting. The trained relighting network can take the optimal light direction, a prediction of surface geometry of a subject in the input image, and a prediction of a merchandise image indicating an amount of light energy to apply to each pixel of the input image. The trained neural network can also process the image to apply a desired lighting to the original image and predict an output image in which the desired lighting has been applied to the input image from the recommended light direction. The trained neural network can then provide an output including the predicted output image.
[0041] In one example, a copy of the trained neural network can reside on a mobile computing device. The mobile computing device can include a camera that can capture an input image of a subject, such as a portrait of a human face. A user of the mobile computing device can view the input image and determine that the input image should be relit. The user can then provide the input image and information about how the input image should be relit to the trained neural network residing on the mobile computing device. In response, the trained neural network can generate a predicted output image showing the input image relit as indicated by the user and subsequently output the output image (e.g., provide the output image for display by the mobile computing device). In other examples, the trained neural network does not reside on the mobile computing device; rather, the mobile computing device provides the input image and information about how the input image should be relit to a remotely located trained neural network (e.g., via the Internet or another data network). The remotely located convolutional neural network can process the input image and the information about how the input image should be relit as indicated above and provide an output image to the mobile computing device showing the input image relit as indicated by the user. In other examples, non-mobile computing devices can also use the trained neural network to relight images, including images that are not captured by a camera of a computing device.
[0042] In some examples, the trained neural network can work in conjunction with other neural networks (or other software) and / or be trained to identify whether an input image of a subject is poorly lit. Then, upon determining that the input image is poorly lit, the trained neural network described herein can apply a correction lighting model to the poorly lit input image, thereby correcting the poor lighting of the input image. The correction lighting model can be selected based on user input and / or predetermined. For example, a user input lighting model or a predetermined lighting model can be used to provide a "flat light" or light that looks like a "fill flash" technique, or to light a subject with a lighting model that produces undesirable shadows on a face (e.g., from a backlit scene), thereby flattening the image.
[0043] In some examples, the trained neural network can take as input an input image and one or more lighting models and provide one or more resulting output images. The trained neural network can then determine the one or more resulting output images by applying each of the plurality of lighting models to the input image. For example, the one or more lighting models can include a plurality of lighting models representing one (or more) light source changing position, lighting color, and / or other characteristics in each of the plurality of lighting models. More specifically, the plurality of lighting models can represent one or more light sources, where at least one light source changes position (e.g., by a predetermined amount) between the provided models. In this approach, the resulting output images represent the input image shown with the changed light source appearing to rotate or otherwise move around the object(s) depicted in the input image. Similarly, the changed light source can change color between the provided models (e.g., by a predetermined distance in a color space), such that the resulting output images represent the input image shown with light of various colors. The plurality of output images can be provided as still images and / or video images. Other effects can be produced by having the trained neural network apply a plurality of lighting models to one image (or, relatedly, having the trained neural network apply one lighting model to a plurality of input images).
[0044] Accordingly, the techniques described herein can improve images by applying more desirable and / or selectable lighting models to the images, thereby enhancing their actual and / or perceived quality. Enhancing the actual and / or perceived quality of images, including portraits of people, can provide emotional benefits to those who believe their photos look better. The techniques are flexible, and thus a wide variety of lighting models can be applied to images of human faces and other objects (particularly other objects having similar lighting characteristics). Moreover, by changing the lighting models, different aspects of the images can be highlighted, which can lead to a better understanding of the objects depicted in the images.
[0045] Techniques for image relighting using neural networks
[0046] Figure 1An image 100 with different illuminations according to example embodiments is shown. The image 100 includes an image 110 and an image 120 depicting the human face of the same person under two different types of illumination. According to example embodiments, the image 110 is relit to obtain an image 130, where the right side of the face in the image 110 is illuminated. The image 130 is an image including a human face where the illumination is darker than the illumination in the image 140. Other examples of images with different illuminations and / or other types of imperfect illumination are also possible. The image 100 shows the effect of illumination on an image. For example, the illumination can affect the perceived expression, the perceived emotion, and / or other aspects of the subject’s features and / or personality.
[0047] Relighting network
[0048] A relighting network can be designed to add additional, repositionable light sources to an image by computation, with initial illumination direction and intensity automatically chosen to compensate for the original lighting conditions of the image. For example, under less than ideal original lighting conditions, such as in a backlit scene, the additional light sources improve the exposure on the eyes and face. If the image already has compelling lighting (as is often the case with scenes with some directional lighting), the image can be elegantly enhanced for dramatic effect, for example, by emphasizing the contours and shapes of the face or person in the image.
[0049] Figure 2 is a diagram depicting a relighting network 200 for enhancing the illumination of an image according to example embodiments. A first convolutional neural network 215 can utilize an input image 210 to generate a surface orientation map 220 indicating the surface geometry of the input image 200. In general, when a photographer adds an additional light source to a scene, the orientation of the light source relative to the face geometry of the subject determines how bright each part of the face presented in the input image 210 is. Optically, according to Lambert’s cosine law, the amount of light reflected by an object composed of a relatively subluminous material is proportional to the cosine of the angle between its surface orientation and the direction of the incident light. To model this behavior, the first convolutional neural network 215 is trained to estimate the surface orientation from the input image 210. The output of the first convolutional neural network 215 is a color representation of a set of three-dimensional vectors in camera space.
[0050] For example, the spatial vector can be mapped to a colorized red-green-blue (RGB) surface orientation map 220. As shown, the original light source illuminates the left side of the face in the input image 210, there is less illumination of the right side of the face in the input image 210, and the background and hair color are substantially darker colors. In the surface orientation map 220, the illuminated left side of the face corresponds to blue pixels, the background and hair color correspond to green pixels, and the less illuminated right side of the face corresponds to red pixels.
[0051] In some aspects, the second convolutional neural network 225 is trained to estimate the ambient lighting corresponding to the input image 210. For example, to recommend optimal light directions, the second convolutional neural network 225 is trained to estimate a high dynamic range, omnidirectional lighting profile of a scene based on an input portrait. This lighting estimation model can infer the direction, relative intensity, and color of all light sources from all directions in the scene of the input image 210 by treating the face as a light probe. The optimal light directions 230 for the relighting network 200 are then automatically determined with the ambient light estimation. In some aspects, the pose of the subject of the portrait can be estimated to determine the optimal light directions 230. Further, for example, labeled image data can be generated that correlates different positions of light sources with optimal light placements for images, and a machine learning model can be trained on the labeled data to automatically determine the light directions 230. For example, in studio portrait photography, the main light source or "key light" is often placed about 30° above the subject's line of sight, and between 30° and 60° from the camera axis when looking at the subject from above. The relighting network 200 can be configured to follow similar guidelines for classic portrait appearance, thereby enhancing the pre-existing lighting directionality in the input image 210 while targeting a balanced, subtle key fill lighting ratio of about 2: 1.
[0052] As previously described, the input image 210 indicates that the right side of the face is less illuminated than the left side of the face. Accordingly, in some embodiments, the relighting network 200 can automatically determine the light directions 230 to illuminate this less illuminated right side of the face. In some embodiments, the relighting network 200 can receive user preferences for the light directions 230 via an interactive graphical user interface.
[0053] Based on Lambert's cosine law, a dot product 235 between the three-dimensional vector representation of the surface orientation map 220 and the three-dimensional vector representation of the light direction 230 can be computed to generate a light visibility map 240. Generally, the light visibility map 240 indicates areas in the portrait that are to be illuminated by the synthetic lighting. For example, the light visibility map 240 indicates, based on the surface orientation, which portions of the input image 200 can see light and which portions cannot see light. As previously described, the input image 210 indicates that the right side of the face is less lit than the left side of the face, and the relighting network 200 can automatically determine the light direction 230 to illuminate this less lit right side of the face. Further, for example, the surface orientation map 220 takes into account the surface geometry of the face in the input image 200 to highlight, with red pixels, portions that the ambient lighting does not provide enough light for. Thus, based on the light direction 230 and the surface orientation map 220, the light visibility map 240 indicates the right side of the face as a portion that needs lighting.
[0054] In enhancing the input image 210 in near real-time, it is advantageous to reduce the utilization of computational resources (e.g., including processing time, processing speed, and memory allocation). Specifically, it can not be desirable to directly utilize the light visibility map 240 to enhance the input image 210. The light visibility map 240 indicates areas in the input image 210 that need to be illuminated and how much light to illuminate these areas, but the light visibility map 240 does not capture the material properties of the object being relit. Depending on the material of the object being relit, the same light can result in very different outcomes, e.g., shiny / specular materials versus dull / diffuse materials reflect light very differently.
[0055] Conversely, the light visibility map 240 and the input image 210 can be input to a third convolutional neural network 245 to predict a quotient image 250. The quotient prediction network 245 must learn the properties of the materials on the skin, eyes, and even clothing from the training data and produce a quotient image 250 that takes into account both the material properties learned from the input image 210 and the optimized light direction and geometry information provided by the light visibility map 240. Typically, the quotient image 250 is a real-valued multiplicative factor for each pixel that indicates the amount of illumination to be applied to each pixel of the input image 210. This can also be supervised with ground truth images. Because the quotient image 250 is a multiplicative factor, it does not attenuate the details from the original input image 210. For example, the quotient image 250 can be predicted even at much lower resolutions, e.g., for a blurry image. On the other hand, processing a high-resolution input image 210 can involve more computational complexity, resulting in a delay for real-time processing. Typically, the details in the high-resolution input image 210 can be preserved while the input image 210 is enhanced to bring out less visible aspects, resulting in a realistic image. The multiplier in the quotient image 250 makes a given pixel lighter or darker. The range of values can vary, sometimes by a factor of 10, depending on the intensity of the synthetic illumination. The quotient image 250 enhances low-frequency lighting changes without affecting high-frequency image details that are transferred directly from the input image 210 to maintain image quality. In some aspects, post-processing can be applied to the quotient image 250 to adjust highlights, exposure adjustments, and matting, resulting in a photo-realistic image enhancement.
[0056] Multiplication 255 of the quotient image 250 and the input image 210 generates a relit image 260. This approach is also computationally efficient because the third convolutional neural network 245 predicts a lower resolution quotient image that is upsampled before multiplication 255 with the high-resolution input. Thus, the relit network 200 combines automatic estimation of surface geometry and light direction to generate a relit image 260 in which high-frequency details of the input image 210 are preserved while low-frequency details of the input image 210 are enhanced. In some aspects, the relit network 200 can be optimized to run at interactive frame rates on a mobile device with a total model size of less than 10 MB. The result can be produced with a combination of utilization of the UNet model and a version of separable, depthwise convolutions and cascaded for skip connections. This, together with float 16 quantization, can result in a model size of 2.4 MB per UNet, thus a total model size of 4.8 MB.
[0057] An interactive user experience can be provided in which users can adjust the position and brightness of the lights. This provides additional creative flexibility to users to find their own balance between light and shadow. Moreover, aspects of the relighting network 200 can be applied to existing images from a user’s photo library, for example. For example, for an existing image in which a face can be slightly underexposed, the relighting network 200 can be applied to light and shape the face. This can be particularly advantageous for images in which a single person posed directly in front of a camera.
[0058] Several aspects of the relighting network 200 will be described in more detail below. For example, the relighting network 200 can take an input image 210 and, without any prior knowledge of the camera, exposure, composition, and without any additional photography hardware devices, the relighting network 200 can derive the geometry, lighting, and find the optimal exposure to enhance the input image 210. Based on the estimation of the original lighting in the input image 210, the relighting network 200 can automatically derive the optimal light directions for the synthetic lighting. Combining the optimal light directions with the knowledge of the surface geometry customizes the application of the synthetic lighting to the geometric features of the objects in the input image 210. Moreover, for example, using a multiplicative factor for each pixel to preserve the high-frequency details of the input image while enhancing the low-frequency lighting details is an important factor in reducing the computational complexity of the image enhancement technique. Further post-processing techniques of adjusting highlights, exposure adjustment, and matting enable the presentation of photo-realistic image enhancements. The intermediate interpretable outputs, such as, for example, the surface orientation map 220, the light directions 230, the light visibility map 240, and the merchandise image 250, provide opportunities for intermediate adjustments that optimize the loss function of the neural network, improve the image quality, and otherwise inform the overall quality of the enhanced image. For example, the surface orientation map 220 can be used to identify sources of error, and the ground truth data can be updated with additional input from photography techniques (e.g., adjustments made by a professional photographer) to correct the errors. An interactive user experience is another important feature of the techniques described herein in which users can adjust the position and brightness of the lights for any image.
[0059] Surface Orientation Maps - Supervising Relighting with Geometric Estimation
[0060] Figure 3is a representation of an example network to predict surface orientation maps according to example embodiments. For example, the convolutional neural network 215 can take an input image 210, infer a reflectance field 310 that indicates the surface geometry of an object in the input image 210 (e.g., a face in a portrait and / or an entire body in the input image 210). Based on ground truth data that indicates the correlation between the lighting and the surface geometry, the convolutional neural network 215 can predict a surface orientation map 220. As described herein, the surface orientation map 220 can be combined with one or more aspects of the relighting network 200 to generate a relit image 260. Further, for example, additional or alternative methods of ground truth geometry / surface orientation can be applied, such as, for example, data from a depth sensor. In some example embodiments, depth from a mobile phone depth sensor can be used to determine the surface orientation map 220, rather than using a network (e.g., the convolutional neural network 215).
[0061] From the light data, a four-dimensional reflectance field 310, i.e., R(u, v, 0, f), can represent the subject illuminated from any lighting direction (0, f) for each image pixel (u, v). In general, the reflectance field 310 describes how the volume of space enclosed by a surface A transforms a directional light (0 i , f i ) into a radiance field of illumination R(0 r , f r , u r , v r ) at a point (u r , v r ). The light data represents one of a specified number of directions from which the face in the portrait in the image training data was lit to train the convolutional neural network 215.
[0062] While the examples described herein involve determining and applying an illumination model for images of objects with human faces, the convolutional neural network 215 can be trained to determine and apply an illumination model for images of other objects that reflect light similarly to human faces. Human faces generally diffuse light, but can include some specular highlights due to directly reflected light. For example, specular highlights can result from direct light reflection from the surface of the eyes, glasses, jewelry, etc. In many human face images, the area of such specular highlights is relatively small compared to the area of the face surface that diffuses light. Thus, the convolutional neural network 215 can be trained to apply an illumination model to images of other objects that diffuse light, where these diffusing objects can have some relatively small specular highlights (e.g., a tomato or a wall painted with a matte finish). The images in the training dataset can show one or more particular objects using illumination provided in a variety of different conditions (e.g., illumination provided from different directions, illumination provided at different intensities (e.g., brighter and dimmer illumination), illumination provided with different colors of light sources, illumination provided with different numbers of light sources, etc.). Once trained, the convolutional neural network 215 can receive an input image 210 and information about the original illumination. The trained convolutional neural network 215 can process the input image 210 to determine a prediction of the illumination based on the surface geometry, generating a surface orientation map 220.
[0063] Figure 4is an example architecture of a neural network that predicts a surface orientation map according to example embodiments. For example, the convolutional neural network 215 can be modeled around a UNet with skip connections. In some embodiments, the input image 210 can be of size (256 x 192 x 3), and can go through 6 encoder blocks (Enc LI 410, Enc L2 412, Enc L3 414, Enc L4 416, Enc L5 418, and Enc L6 420) and 6 decoder blocks (Dec LI 422, Dec L2 424, Dec L3 426, Dec L4 428, Dec L5 430, and Dec L6 432). Each encoder block of the encoder 401 includes a single convolution followed by a down-sampling by a factor of 2 followed by a blur pooling operation. Each decoder block of the decoder 403 includes a single convolution followed by an up-sampling by a factor of 2 followed by a bilinear up-sampling operation. Depending on the network version, the skip connections from the encoder 401 are concatenated or added to the up-sampled features of the decoder 403. In some example embodiments, the filters of the encoder 401 can include (16, 32, 64, 128, 256, 256) filters per encoder block, which results in a final bottleneck size of (4 x 3 x 256). Similarly, the filters of the decoder 403 can include (256, 256, 128, 64, 32, 16) features to return to a (256 x 192) image resolution. A final convolution with 3 filters can be used to produce the surface orientation map 220 from the decoder output. Relu activations can be used after all convolutions.
[0064] The image training data 405 represents a set of facial portraits taken with various lighting arrangements. In some implementations, the image training data 405 includes images of faces or portraits formed with high dynamic range (HDR) lighting captured from low dynamic range (LDR) lighting environments recovered. As shown, the image training data 405 includes a plurality of images 406(1),..., 406(M), where M is the number of images in the image training data 405. Each image, such as image 406(1), includes light data 408(1) and can also include pose data. Figure 4
[0065] The set of light data 408 (1...M) based on the ground truth data represents one of a specified number (e.g., 331) of directions from which the face is illuminated for the portrait used in the image training data 405. In some implementations, the light data 408 (1) includes a polar and azimuthal angle, i.e., coordinates on the unit sphere. In some implementations, the light data 408 (1) includes a triple of direction cosines. In some implementations, the light data 408 (1) includes a set of Euler angles. In some implementations, the angular configuration represented by the light data 408 (1) is one of 331 configurations used to train the convolutional neural network 215.
[0066] In some implementations, to capture the reflectance field of a subject, a computer- controllable sphere of white LED light sources can be used, with lights spaced at 12° intervals at the equator. In such implementations, the reflectance field is formed from a set of reflectance basis images, taken of the subject as each directional LED light source is individually turned on one at a time within the spherical mount. Such one-light-at-a-time (OLAT) images are captured for multiple camera viewpoints. In some implementations, 331 OLAT images are captured for each subject using six color machine vision cameras with 12 megapixel resolution placed 1.7 meters from the subject, although these values and the number of OLAT images and the type of cameras used can differ in some implementations. In some implementations, the cameras are positioned approximately in front of the subject, with five cameras having 35mm lenses capturing the upper body of the subject from different angles, and one additional camera having a 50mm lens capturing close-up images of the face.
[0067] In some implementations, using a reflectance field of 70 distinct subjects, each exhibiting ten different facial expressions and wearing different accessories, approximately 700 sets of OLAT sequences are produced from six different camera viewpoints, for a total of 4200 unique OLAT sequences. Other numbers of sets of OLAT sequences can be used. Subjects are captured that span a wide range of skin pigmentation. In addition, for example, 32 custom high-resolution (e.g., 12MP) depth sensors can be used. As another example, 62 high-resolution (e.g., 12MP) RGB cameras can be used. In total, over 15 million images can be generated as the image training data 405.
[0068] Since it takes some time, e.g., approximately six seconds, to acquire a complete OLAT sequence of a subject, there can be some slight subject motion between frames. In some implementations, optical flow techniques are used to align the images, occasionally (e.g., every 11th OLAT frame) interspersing one additional “tracking” frame with uniformly consistent lighting to ensure that the optical flow’s brightness constancy constraint is met. This step can preserve the sharpness of image features when performing the relighting operation that linearly combines the aligned OLAT images.
[0069] Figure 5 Example images of the captured reflective field are shown according to an exemplary embodiment. For example, a spherical illumination apparatus may include 64 cameras with different viewpoints and 331 individually programmable LED light sources. Figure 5 As shown, each individual can be photographed by each light source according to the illumination of the OLAT, forming a reflected field (e.g., Figure 3 Image 510 depicts the appearance of an individual illuminated by tiny cones of light in a 360° environment. The reflection field 310 encodes the unique color and reflective properties of each individual's skin, hair, and clothing. For example, the reflection field 310 encodes information about how bright or dark each material appears. Due to the principle of light superposition, these OLAT images can then be linearly added together to present a realistic image 520 of the individual, as if it were in any image-based lighting environment, correctly representing complex light transport phenomena such as subsurface scattering. Synthetic portraits of each individual can be generated in many different lighting environments with and without added directional light, presenting millions of image pairs for image training data 405. The dataset can include model performance across disparate lighting environments and individuals.
[0070] The convolutional neural network 215 can be a fully convolutional neural network as described herein. During training, the convolutional neural network 215 can receive one or more input training images as input. The convolutional neural network 215 can include layers of nodes for processing the input image 210. Example layers can include, but are not limited to, an input layer, a convolutional layer, an activation layer, a pooling layer, and an output layer. The input layer can store input data, such as pixel data of the input image 210 and inputs from other layers of the convolutional neural network 215. The convolutional layer can compute outputs of neurons connected to local regions in the input. In some examples, the predicted outputs can be fed back into the convolutional neural network 215 as input to perform iterative refinements. The activation layer can determine whether outputs of a previous layer are “activated” or actually provided (e.g., to a successor layer). The pooling layer can downsample the input. For example, the convolutional neural network 215 can include one or more pooling layers that downsample the input in the horizontal and / or vertical dimensions by a predetermined factor (e.g., a factor of 2). The output layer can provide outputs of the convolutional neural network 215 to software and / or hardware interfacing with the convolutional neural network 215, such as to hardware and / or software for displaying, printing, communicating, and / or otherwise providing the surface orientation map 220 (e.g., to one or more components of the relighting network 200). The layers 410, 412, 414, 416, 418, 420, 422, 424, 426, 428, 430, 432 can include one or more input layers, output layers, convolutional layers, activation layers, pooling layers, and / or other layers described herein.
[0071] In some implementations, the convolutional neural network 215 can include encoding layers 410, 412, 414, 416, 418 arranged in order of layers LI, L2, L3, L4, L5, respectively, each layer successively convolving its input and providing its output to a successor layer until reaching the encoding layer L6 420. In some implementations, the convolutional neural network 215 can include decoding layers 422, 424, 426, 428, 430 arranged in order of layers L6, L5, L4, L3, LI, respectively, each layer successively deconvolving its input and providing its output to a predecessor layer until reaching the decoding layer LI 410. Figure 4 In some implementations, the depicted layers can include one or more actual layers. For example, the encoding layer LI 410 can have one or more input layers, one or more activation layers, and / or one or more additional layers. As another example, the encoding layer L2 412, the encoding layer L3 414, the encoding layer L4 416, the encoding layer L5 418, and / or the encoding layer L6 420 can include one or more convolutional layers, one or more activation layers (e.g., in one-to-one relationship with the one or more convolutional layers), one or more pooling layers, and / or one or more additional layers.
[0072] In some examples, some or all of the pooling layers in the convolutional neural network 215 can downsample the input by a common factor in both the horizontal and vertical dimensions without downsampling the depth dimension associated with the input. The depth dimension can store data for pixel colors (red, green, blue) and / or data representing scores. One or more of the layers of the convolutional neural network 215 can also use other common factors for downsampling other than 2.
[0073] The encoding layer LI 410 can receive and process the input image 210 and provide an output to the encoding layer L2 412. The encoding layer L2 412 can process the output of the encoding layer LI 410 and provide an output to the encoding layer L3 414. The encoding layer L3 414 can process the output of the encoding layer L2 412 and provide an output to the encoding layer L4 416. The encoding layer L4 416 can process the output of the encoding layer L3 414 and provide an output to the encoding layer L5 418. The encoding layer L5 418 can process the output of the encoding layer L4 416 and provide an output to the encoding layer L6 420.
[0074] The encoding layer L6 420 can provide output to the decoding layer L1 422 to begin predicting the surface orientation map 220. The decoding layer L2 424 can receive and process input from both the decoding layer L1 422 and the encoding layer L5 418 (e.g., using a skip connection between the encoding layer L5 418 and the decoding layer L2 424) to provide output to the decoding layer L3 426. The decoding layer L3 426 can receive and process input from both the decoding layer L2 424 and the encoding layer L4 416 (e.g., using a skip connection between the encoding layer L4 416 and the decoding layer L3 426) to provide output to the decoding layer L4 428. The decoding layer L4 428 can receive and process input from both the decoding layer L3 426 and the encoding layer L3 414 (e.g., using a skip connection between the encoding layer L3 414 and the decoding layer L4 428) to provide output to the decoding layer L5 430. The decoding layer L5 430 can receive and process input from both the decoding layer L4 428 and the encoding layer L2 412 (e.g., using a skip connection between the encoding layer L2 412 and the decoding layer L5 430) to provide output to the decoding layer L6 432. The decoding layer L6 432 can receive and process input from both the decoding layer L5 430 and the encoding layer L1 410 (e.g., using a skip connection between the encoding layer L1 410 and the decoding layer L6 432) to provide a prediction of the surface orientation map 220, which can then be output from the decoding layer L6 432. The data provided by the skip connections between the encoding layers 418, 416, 414, 412, 410 and the corresponding decoding layers 424, 426, 428, 430, 432 can be used by each respective decoding layer to provide additional detail for generating a contribution to the prediction of the surface orientation map 220 by the decoding layer. In some examples, each of the decoding layers 422, 424, 426, 428, 430, 432 used to predict the surface orientation map 220 can include one or more convolutional layers, one or more activation layers, and possibly one or more input and / or output layers. In some examples, some or all of the layers 410, 412, 414, 416, 418, 420, 422, 424, 426, 428, 430, 432 can function as a convolutional encoder / decoder network.
[0075] In some implementations, the convolutional neural network 215 is trained end-to-end with a loss on the predicted surface orientation map 220. For example, the convolutional neural network 215 can be trained using a combination of LI loss and adversarial loss. Generally, empirical data suggests that adversarial loss is an important factor for good generalization to in-the-wild images when the training data is limited (e.g., 15 subjects). However, LI loss achieves similar results with larger datasets (e.g., 70 subjects). Moreover, adversarial loss can become more difficult to train as the variation in viewpoint and subject clothing grows. In some implementations, adversarial loss can be selectively applied to facial portions of the image. Other loss metrics can likewise or alternatively be used. For example, during training of the convolutional neural network 215 for predicting the surface orientation map 220, an L2 loss metric between the surface orientation map prediction and the training image can be minimized.
[0076] As described herein, the convolutional neural network 215 can include perceptual loss processing. For example, the convolutional neural network 215 can use a generative adversarial network (GAN) loss function to determine whether a portion or all of the image is to be predicted to generate a surface orientation map, and thus satisfy one or more perceptually-related conditions regarding the lighting of that portion of the image. In some examples, a cycle loss can be used to feed back the predicted surface orientation map into the convolutional neural network 215 to generate and / or refine a further predicted surface orientation map. In some examples, the convolutional neural network 215 can utilize a deep supervision technique to provide constraints on intermediate layers. In some examples, the convolutional neural network 215 can have more, fewer, and / or different layers than those shown. Figure 4
[0077] Light Direction - Automatic Light Placement
[0078] Figure 6 An example method for recommending light directions is shown in accordance with example embodiments. An input image 210 can be processed by a face detection component 610 to detect a human face in the input image 210. The image 620 shows the face of a human subject within a bounding box. This portion of the image 620 can be cropped and resized (e.g., enlarged or reduced) to obtain a face 630. While the background image can provide contextual cues that aid in lighting estimation, some implementations compute a face bounding box for each input and, during training and inference, some implementations crop each image, extending the bounding box by 25%. During training, some implementations add slight cropping region variations, changing their position and extent randomly.
[0079] The face 630 and / or the image 620 can be input to a convolutional neural network 225. In some embodiments, the architecture components of the convolutional neural network 225 can be similar to those referenced in the description of the convolutional neural network 215. Figure 4 The described architecture components of the convolutional neural network 215. The convolutional neural network 225 can take an input face 630 and predict an ambient lighting profile for the input image 210. The convolutional neural network 225 is a lighting estimation model that infers the direction, relative intensity, and color of all light sources in the face 630 from all directions, treating the face 630 as a light probe.
[0080] For example, an image of a person can be taken while being illuminated by each individual light source in the image 620 and / or face 630. In some aspects, the training data for the convolutional neural network 225 can include Figure 4 image training data 405. During training, the convolutional neural network 225 can receive one or more input training images as input and input a target lighting profile. For example, the convolutional neural network 225 can be trained on input images 410. During training, the convolutional neural network 225 is directed to generate a prediction of the original lighting profile as the estimated lighting 640.
[0081] The original lighting model is a light estimation model that predicts the actual lighting conditions used to illuminate an input image 210. The lighting model can include a grid or other arrangement of lighting profile data related to the lighting of some or all of one or more images. The lighting profile data can include, but is not limited to, data representing one or more colors, intensities, directions of the lighting of some or all of one or more images. The convolutional neural network 225 can be configured to produce a predicted lighting profile or estimated lighting 640 based on a plurality of images of a plurality of human faces. The convolutional neural network 225 is configured to produce a predicted lighting profile or estimated lighting 640 based on an input image 210. The input image 210 represents at least one human face. In some aspects, the convolutional neural network 225 includes a cost function (e.g., a discriminator and cost function data) based on a plurality of bidirectional reflectance distribution functions (BRDFs) corresponding to each reference object. The estimated lighting 640 represents a spatial distribution of lighting incident on the portrait subject. An example representation of the estimated lighting 640 includes coefficients of a spherical harmonic expansion of an illumination function of an angle. Another example representation of the estimated lighting 640 includes a grid of pixels, each pixel having a value of an illumination function of a solid angle.
[0082] As described herein, the input to the convolutional neural network 225 is an sRGB encoded LDR image, e.g., LDR portrait 620, with a crop 630 of the face region of each image detected by the face detector 610, resized to an input resolution of 256x256, and normalized to the range [-0.5, 0.5]. The convolutional neural network 225 has an encoder / decoder architecture that includes a latent vector representation of the log-space HDR illumination with a size of 1024 at the bottleneck. In some implementations, the encoder and decoder are implemented as convolutional neural networks. In some implementations, the encoder includes five 3x3 convolutions, each followed by a blur-pooling operation, with successive filter depths of 16, 32, 64, 128, and 256, followed by a final convolution with a filter size of 8x8 and a depth of 256, followed by a fully connected layer. The decoder includes three sets of 3x3 convolutions with filter depths of 64, 32, and 16, each followed by a bilinear up-sampling operation. The final output of the convolutional neural network 225 includes a 32x32 HDR image of the estimated illumination 640, which can be visualized by rendering three different spheres (diffuse, sub- specular silver, and specular) and represents a log-space omnidirectional illumination.
[0083] Figure 7 Example images with estimated illumination are shown in accordance with example embodiments. For example, images 710, 720, and 730 are shown along with respective estimated illuminations 640A, 640B, and 640C. For example, the estimated illuminations 640A, 640B, and 640C estimate a high dynamic range, omnidirectional illumination profile from the images 710, 720, and 730, respectively. The three spheres (640A, 640B, and 640C) to the right of each image are rendered using the estimated illumination, showing the color, intensity, and directionality of the ambient lighting in the images 710, 720, and 730, respectively.
[0084] Referring back to Figure 6 The input image 210 can be processed to estimate a pose 650 of a person in the input image 210. In some embodiments, this operation can be performed by a facial geometry solution (e.g., a facial mesh). A machine learning model can infer 3D surface geometry from a single camera input without a dedicated depth sensor. The facial geometry data can include common 3D geometry primitives, including a facial pose transformation matrix and a triangular facial mesh. Further, for example, a machine learning solution for high-fidelity upper body pose tracking can be configured to infer 25 2D upper body landmarks from RGB video frames. Using a detector, the estimated pose detector first localizes a pose region of interest (ROI) within an image frame. The detector then uses the ROI-cropped frame as input to predict the pose landmarks within the ROI.
[0085] Photographers often rely on perceptual cues when deciding how to augment ambient lighting with camera external light sources. They assess the intensity and direction of light falling on the face and also adjust their subject's head pose to compensate for it. The computational equivalent of these operations is the estimated lighting 640 and the estimated pose 650. Furthermore, in studio portrait photography, for example, the main camera external light source or "key light" is typically placed about 30° above the subject's line of sight, between 30° and 60° from the camera axis when looking at the subject from above. These guidelines can be configured to be performed by computation based on the estimated lighting 640 and the estimated pose 650 to automatically suggest optimal light directions 230. For example, the light directions 230 indicate directions in which synthetic lighting can be added to the input image 210, similar to how a professional photographer would place a camera external light source to light a subject. Thus, pre-existing lighting directionality in the scene can be augmented while targeting a balanced, subtle key fill lighting ratio of about 2: 1. As used herein, fill lighting generally refers to additional lighting that can be added to light the low frequency parts of an image.
[0086] Given the desired lighting directions 230 and the input image 210, a machine learning model can be trained to add synthetic light from directional light sources to the original photograph. For supervised training, millions of pairs of portraits with and without extra light are used. Capturing such a dataset in real life would be challenging, requiring near perfect registration of portraits taken under different lighting conditions. Instead, many individuals with different face shapes, genders, skin tones, hairstyles, and clothing / accessories can be photographed in a computational lighting system.
[0087] Surface orientation map * optimal light directions = light visibility map
[0088] Figure 8A An example network architecture that determines a light visibility map is shown according to example embodiments. An input image 210 can be processed to output a surface orientation map 220 and light directions 230. According to Lambert's law, the amount of light reflected by an object composed of a relatively matte material is proportional to the cosine of the angle between its surface orientation (e.g., surface orientation map 220) and the direction of incident light (e.g., light directions 230). Thus, Lambert's law can be applied to take the dot product 235 to obtain a light visibility map 240.
[0089] The light visibility map 240 is a representation of the light directions 230 tailored to the surface geometry. Such a representation can serve two purposes. First, the light directions 230 are more tightly linked to the prediction of the surface geometry, which can greatly simplify the image quality assessment of the relighting network 200. Many errors that arise in relighting can be traced back to errors in the predicted surface orientation map 220. Second, the light visibility map 240 serves as a better per-pixel representation of the light directions 230 for the commercial image predicting UNet, e.g., the third convolutional network 245. The relighting network 200 does not need to infer which pixels will be lit, but only needs to determine how they are lit (e.g., material properties), and how subtle shadows are cast (e.g., occlusions). Figure 2
[0090] Figure 8B Example images showing the effect of highlighting surface geometry in photo lighting according to example embodiments are shown. For example, an image 810 is obtained in a lighting network that does not use the surface orientation map 220, and an image 820 is obtained in a lighting network that uses the surface orientation map 220. As shown, the geometry of the face in the image 820 is more detailed, and in the lower left half of the face, from the nose to the chin, there is less flattening.
[0091] Commercial image - learning details - preserving relighting
[0092] Figure 9 is a diagram depicting an example network that predicts a commercial image according to example embodiments. For example, the convolutional neural network 245 can take the input image 210 and the light visibility map 240 to predict a commercial image 250 that indicates the per-pixel light energy needed to enhance the input image 210. In general, the commercial image 250 is the quotient of the relit image 260 and the input image 210. Predicting such a multiplicative factor enables the relighting network 200 to preserve high frequency details in the input image 210, such as skin pores and individual strands of hair, rather than attempting to reproduce these high frequency details through a neural network. The resolution of the predicted commercial image 250 is significantly lower than the input image 210. To apply the low resolution commercial image 250 to the high resolution input image 210, the predicted commercial image 250 can be upsampled using bilinear or other upsampling methods.
[0093] Figure 10 is an example architecture of a neural network that predicts quotient images according to example embodiments. For example, the convolutional neural network 245 can be modeled around a UNet with skip connections. In some embodiments, the input image 210 and the light visibility map 240 can go through 6 encoder blocks (Enc LI 1010, Enc L2 1012, Enc L3 1014, Enc L4 1016, Enc L5 1018, and Enc L6 1020) and 6 decoder blocks (Dec LI 1022, Dec L2 1024, Dec L3 1026, Dec L4 1028, Dec L5 1030, and Dec L6 1032). Each encoder block of the encoder 1001 includes a single convolution followed by a down-sampling by a factor of 2 followed by a blur pooling operation. Each decoder block of the decoder 1003 includes a single convolution followed by an up-sampling by a factor of 2 followed by a bilinear up-sampling operation. Depending on the network version, the skip connections from the encoder 1001 are concatenated or added to the up-sampled features of the decoder 1003. In some example embodiments, the filters of the encoder 1001 can include (16, 32, 64, 128, 256, 256) filters per encoder block resulting in a final bottleneck size of (4 x 3 x 256). Similarly, the filters of the decoder 1003 can include (256, 256, 128, 64, 32, 16) features to return to a (256 x 192) image resolution. A final convolution with 3 filters can be used to produce the quotient image 250 from the decoder output. Relu activations can be used after all convolutions. This can also be supervised with ground truth image data 1005.
[0094] The convolutional neural network 245 can be a fully convolutional neural network as described herein. During training, the convolutional neural network 245 can receive as input one or more input training images 1006(1...M) and light data 1008(1...M) from ground truth image data 1005. The convolutional neural network 245 can include layers of nodes for processing input images 210 and light visibility maps 240. Example layers can include, but are not limited to, input layers, convolutional layers, activation layers, pooling layers, and output layers. Input layers can store input data, such as pixel data for input images 210 and light visibility maps 240, and inputs from other layers of the convolutional neural network 245. Convolutional layers can compute outputs of neurons connected to local regions in the input. In some cases, bilinear upsampling is performed followed by convolution to apply a filter to a relatively small input, expanding / up-sampling the relatively small input to a larger output. Activation layers can determine whether outputs of a previous layer are “activated” or actually provided (e.g., to a successor layer). Pooling layers can down-sample inputs. For example, the convolutional neural network 245 can use one or more pooling layers to down-sample inputs in the horizontal and / or vertical dimensions by a predetermined factor (e.g., a factor of 2). Output layers can provide outputs of the convolutional neural network 245 to software and / or hardware interfacing with the convolutional neural network 245, such as hardware and / or software for displaying, printing, transmitting, and / or otherwise providing merchant images 250 (e.g., to one or more components of the relighting network 200). Layers 1010, 1012, 1014, 1016, 1018, 1020, 1022, 1024, 1026, 1028, 1030, 1032 can include one or more input layers, output layers, convolutional layers, activation layers, pooling layers, and / or other layers described herein.
[0095] In some implementations, the convolutional neural network 245 can include encoding layers 1010, 1012, 1014, 1016, 1018 arranged in order of layers LI, L2, L3, L4, L5, respectively, each layer successively convolving its input and providing its output to a successor layer, until reaching encoding layer L6 1020. In some implementations, the convolutional neural network 245 can include decoding layers 1022, 1024, 1026, 1028, 1030 arranged in order of layers L6, L5, L4, L3, LI, respectively, each layer successively deconvolving its input and providing its output to a predecessor layer, until reaching the output layer 1032. Figure 10 In some implementations, the depicted layers can include one or more actual layers. For example, the encoding layer LI 1010 can have one or more input layers, one or more activation layers, and / or one or more additional layers. As another example, the encoding layer L2 1012, the encoding layer L3 1014, the encoding layer L4 1016, the encoding layer L5 1018, and / or the encoding layer L6 1020 can include one or more convolutional layers, one or more activation layers (e.g., in one-to-one relationship with the one or more convolutional layers), one or more pooling layers, and / or one or more additional layers.
[0096] In some examples, some or all of the pooling layers in the convolutional neural network 245 can downsample the input by a common factor in both the horizontal and vertical dimensions without down sampling the depth dimension associated with the input. The depth dimension can store data for pixel color (red, green, blue) and / or data representing scores. One or more of the layers of the convolutional neural network 245 can also use other common factors for downsampling other than 2.
[0097] The encoding layer LI 1010 can receive and process the input image 210 and the light visibility map 240 and provide an output to the encoding layer L2 1012. The encoding layer L2 1012 can process the output of the encoding layer LI 1010 and provide an output to the encoding layer L3 1014. The encoding layer L3 1014 can process the output of the encoding layer L2 1012 and provide an output to the encoding layer L4 1016. The encoding layer L4 1016 can process the output of the encoding layer L3 1014 and provide an output to the encoding layer L5 1018. The encoding layer L5 1018 can process the output of the encoding layer L4 1016 and provide an output to the encoding layer L6 1020.
[0098] The encoding layer L6 1020 can provide output to the decoding layer Ll 1022 to begin predicting the merchant image 250. The decoding layer L2 1024 can receive and process input from both the decoding layer Ll 1022 and the encoding layer L5 1018 (e.g., using a skip connection between the encoding layer L5 1018 and the decoding layer L2 1024) to provide output to the decoding layer L3 1026. The decoding layer L3 1026 can receive and process input from both the decoding layer L2 1024 and the encoding layer L4 1016 (e.g., using a skip connection between the encoding layer L4 1016 and the decoding layer L3 1026) to provide output to the decoding layer L4 1028. The decoding layer L4 1028 can receive and process input from both the decoding layer L3 1026 and the encoding layer L3 1014 (e.g., using a skip connection between the encoding layer L3 1014 and the decoding layer L4 1028) to provide output to the decoding layer L5 1030. The decoding layer L5 1030 can receive and process input from both the decoding layer L4 1028 and the encoding layer L2 1012 (e.g., using a skip connection between the encoding layer L2 1012 and the decoding layer L5 1030) to provide output to the decoding layer L6 1032. The decoding layer L6 1032 can receive and process input from both the decoding layer L5 1030 and the encoding layer Ll 1010 (e.g., using a skip connection between the encoding layer Ll 1010 and the decoding layer L6 1032) to provide a prediction of the merchant image 250, which can then be output from the decoding layer L6 1032. The data provided by the skip connections between the encoding layers 1018, 1016, 1014, 1012, 1010 and the corresponding decoding layers 1024, 1026, 1028, 1030, 1032 can be used by each respective decoding layer to provide additional detail for generating a contribution to the prediction of the merchant image 250 by the decoding layer. In some examples, each of the decoding layers 1022, 1024, 1026, 1028, 1030, 1032 used to predict the merchant image 250 can include one or more convolutional layers, one or more activation layers, and possibly one or more input and / or output layers. In some examples, some or all of the layers 1010, 1012, 1014, 1016, 1018, 1020, 1022, 1024, 1026, 1028, 1030, 1032 can function as a convolutional encoder / decoder network.
[0099] In some implementations, the convolutional neural network 245 can be trained end-to-end with a loss on the high resolution relit image 260. For example, the loss can be computed as
[0100] Loss(UpsampledQuotient*HighResInputImage, HighResGroundTruthImage), where UnsampledQuotient is the un-sampled quotient image, HighResInputImage is the high resolution input image, and HighResGroundTruthImage is the high resolution ground truth image, to directly focus on the high resolution relit image 260, rather than trying to obtain an accurate quotient image 250. The convolutional neural network 245 can be trained using a combination of LI loss and adversarial loss. Generally, empirical data indicates that adversarial loss is an important factor for good generalization to in-the-wild images when the training data is limited (e.g., 15 subjects). However, LI loss achieves similar results with larger datasets (e.g., 70 subjects). Moreover, adversarial loss can become more difficult to train as the variation in viewpoint and subject clothing grows. In some implementations, adversarial loss can be selectively applied to facial portions of the image. Other loss metrics can likewise or alternatively be used. For example, during training of the convolutional neural network 245 to obtain the relit image 260, an L2 loss metric between the surface normal map prediction and the training image can be minimized.
[0101] As described herein, the convolutional neural network 245 can include perceptual loss processing. For example, the convolutional neural network 245 can use a generative adversarial network (GAN) loss function to determine whether a portion or all of an image is to be predicted to generate a surface normal map, and thus satisfy one or more perceptually-related conditions regarding the lighting of that portion of the image. In some examples, the output of the prediction can be fed back into the convolutional neural network 245 as input to perform iterative refinement. In some examples, the convolutional neural network 245 can use a deep supervision technique to provide constraints on intermediate layers. In some examples, the convolutional neural network 245 can have more, fewer, and / or different layers than those shown. Figure 10
[0102] Figure 11 is an example network illustrating a post-processing technique according to example embodiments. In one aspect, a multiplication / post-processing operation 1105 can be applied to the input image 210 and the quotient image 250 to output a relit image 260. As described herein, the quotient image 250 indicates a per-pixel multiplication factor that makes each pixel darker, brighter, or unchanged. The values of the quotient image 250 are unconstrained real numbers. However, during post-processing 1105, an increase in average foreground brightness or average foreground luminosity can be limited. Such an increase can be held fixed during training, and during post-processing 1105, a constraint can be added so that the relit image 260 is photo-realistic. Thus, in some embodiments, the quotient image 250 can be post-processed, resulting in a more realistic rendering of the relit image 260.
[0103] One or more post-processing techniques can be applied. For example, a highlight protection 1110 can be applied to enhance the input image 210. Bright portions of the input image 210 can be attenuated, providing a more realistic rendering of the input image 210 by restoring, for example, natural skin tones. Generally, the highlight protection 1110 provides a local adjustment of luminance. Further, for example, an exposure compensation 1115 can be applied. Assuming that the input image 210 was captured with a fixed exposure, the input image 210 can be normalized to a fixed foreground exposure value. Generally, input images can have a large range of exposure values. For example, exposure values can depend on light-colored clothing, dark-colored clothing, etc. Thus, underexposure and / or overexposure of the input image 210 can be compensated for, and the exposure can be compensated to further adjust the quotient image. Generally, the exposure compensation 1115 provides a global adjustment of luminance. As another example, a matting refinement 1120 can be applied to the input image 210 to smooth edges.
[0104] Figure 12 is an example interactive graphical user interface (UI) 1200 according to example embodiments. The UI 1200 can be displayed by a computing device (e.g., a display screen 1210 of a mobile phone). The user can be provided with the ability to adjust the light position 1220 by changing the input value via a user-friendly slider bar 1230. For example, the user can select an optional option of “portrait light” 1240. This allows the user to reposition the light position 1220 to relit the image. The user can also change the light intensity by using a user-friendly intensity scale 1250. As another example, an “auto” setting 1260 can automatically reposition the light position 1220 and the light intensity for optimal results. Adjusting the light position and the light intensity changes the synthetic lighting, and can enhance the input image based on one or more techniques described herein (e.g., with reference to the relit network 200).
[0105] Figure 13An example input-output image pair based on a surface orientation map is shown in accordance with example embodiments. For example, an input image 1310 can be processed to obtain a surface orientation map 1320. The surface orientation map 1320 can then be combined with light directions to output a relit image 1330.
[0106] Figure 14 An example relit pipeline is shown in accordance with example embodiments. An input image 1410 shows a portrait of a person. In some cases, the input image 1410 can be an image captured by a user with an image capture device. In other cases, the input image 1410 can be an image from a photo library of a user. The user can provide instructions to enhance the input image 1410. In response, the input image 1410 is processed for light estimation 1420. For example, the original environment lighting in the input image 1410 can be predicted. In some implementations, a convolutional neural network (e.g., convolutional neural network 225) can be used for light estimation 1420. A face detection operation can be performed on the input image 1420, and a face light estimation 1430 can be performed. In some instances, the network architecture described with reference to Figure 6 One or more aspects of the network architecture described can be used for face light estimation 1430.
[0107] Figure 15 An example input image with predicted face light estimation is shown in accordance with example embodiments. A face light estimation 1530 can be performed on each input image. Each input image and two sets of 3 spheres are shown in Figure 15
[0108] Referring again to Figure 14 Segmentation 1440 can be performed on input image 1410. For example, body part segmentation or alpha mask segmentation can be performed to detect foreground object 1450. Input image 1410, predicted face lighting 1430, and foreground object 1450 can be input to relighting network 1460. One or more aspects of relighting network 1460 can correspond to aspects of relighting network 200. For example, light directions for synthetic lighting can be automatically optimized. In some instances, aspects of this operation can be performed by a neural network (e.g., convolutional neural network 225). In some implementations, light directions can be input as user preferences (e.g., via user interface 1200). Further, for example, a surface orientation map (e.g., surface orientation map 220) can be predicted. In some instances, aspects of this operation can be performed by a neural network (e.g., convolutional neural network 215). Scalar products of the surface orientation map and light directions can be used to output a light visibility map (e.g., light visibility map 240). A quotient image 1470 can be output by relighting network 1460. In some instances, input image 1410 and quotient image 1470 can be multiplied to output a relit image 1490. In some instances, post-processing 1480 can be performed on input image 1410 based on quotient image 1470 to output relit image 1490.
[0109] Training a machine learning model for generating inferences / predictions
[0110] Figure 16 FIG. 1600 illustrates a training phase 1602 and an inference phase 1604 of a trained machine learning model 1632, in accordance with example embodiments. Some machine learning techniques involve training one or more machine learning algorithms on an input set of training data to identify patterns in the training data and provide output inferences and / or predictions about the (patterns in the) training data. The resulting trained machine learning algorithm can be referred to as a trained machine learning model. For example, Figure 16 Training phase 1602 is illustrated, in which one or more machine learning algorithms 1620 are trained on training data 1610 to become a trained machine learning model 1632. Then, during inference phase 1604, trained machine learning model 1632 can receive input data 1630 and one or more inference / prediction requests 1640 (possibly as part of input data 1630), and in response, provide one or more inferences and / or predictions 1650 as output.
[0111] As such, the trained machine learning model 1632 can include one or more models of one or more machine learning algorithms 1620. The machine learning algorithms 1620 can include, but are not limited to: artificial neural networks (e.g., convolutional neural networks, recurrent neural networks, Bayesian networks, hidden Markov models, Markov decision processes, logistic regression functions, support vector machines, suitable statistical machine learning algorithms, and / or heuristic machine learning systems described herein). The machine learning algorithms 1620 can be supervised or unsupervised, and can implement any suitable combination of online and offline learning.
[0112] In some examples, the machine learning algorithms 1620 and / or the trained machine learning model 1632 can be accelerated using a co-processor on a device such as a graphics processing unit (GPU), a tensor processing unit (TPU), a digital signal processor (DSP), and / or an application-specific integrated circuit (ASIC). The co-processor on such a device can be used to accelerate the machine learning algorithms 1620 and / or the trained machine learning model 1632. In some examples, the trained machine learning model 1632 can be trained, resident, and executed to provide inferences on a particular computing device, and / or can otherwise make inferences for a particular computing device.
[0113] During the training phase 1602, the machine learning algorithms 1620 can be trained by providing at least the training data 1610 as training input using unsupervised, supervised, semi-supervised, and / or reinforcement learning techniques. Unsupervised learning involves providing the machine learning algorithms 1620 with some (or all) of the training data 1610, and the machine learning algorithms 1620 determine one or more output inferences based on the provided some (or all) of the training data 1610. Supervised learning involves providing the machine learning algorithms 1620 with some of the training data 1610, where the machine learning algorithms 1620 determine one or more output inferences based on the provided some of the training data 1610, and the output inferences are accepted or corrected based on correct results associated with the training data 1610. In some examples, the supervised learning of the machine learning algorithms 1620 can be governed by a set of rules and / or a set of labels for the training input, and the set of rules and / or the set of labels can be used to correct inferences of the machine learning algorithms 1620.
[0114] Semi-supervised learning involves having correct results for only a portion of the training data 1610. During semi-supervised learning, supervised learning is used for the portion of the training data 1610 that has correct results, and unsupervised learning is used for the portion of the training data 1610 that does not have correct results. Reinforcement learning involves the machine learning algorithm 1620 receiving a reward signal with respect to a previous inference, where the reward signal can be a numerical value. During reinforcement learning, the machine learning algorithm 1620 can output an inference and receive a reward signal in response, where the machine learning algorithm 1620 is configured to attempt to maximize the numerical value of the reward signal. In some examples, reinforcement learning also utilizes a value function that provides a numerical value representing an expected sum of numerical values provided by the reward signal over time. In some examples, the machine learning algorithm 1620 and / or the trained machine learning model 1632 can be trained using other machine learning techniques, including but not limited to incremental learning and curriculum learning.
[0115] In some examples, the machine learning algorithm 1620 and / or the trained machine learning model 1632 can use transfer learning techniques. For example, transfer learning techniques can involve a trained machine learning model 1632 that is pre-trained on a set of data and is additionally trained using the training data 1610. More specifically, the machine learning algorithm 1620 can be pre-trained on data from one or more computing devices, and the resulting trained machine learning model is provided to a particular computing device, where the particular computing device is intended to execute the trained machine learning model during the inference phase 1604. The pre-trained machine learning model can then be additionally trained during the training phase 1602 using the training data 1610, which can be derived from the kernel and non-kernel data of the particular computing device. This further training of the machine learning algorithm 1620 and / or the pre-trained machine learning model using the training data 1610 of the data of the particular computing device can be performed using supervised or unsupervised learning. Once the machine learning algorithm 1620 and / or the pre-trained machine learning model has been trained on at least the training data 1610, the training phase 1602 can be completed. The trained result machine learning model can be used as at least one of the trained machine learning models 1632.
[0116] Specifically, once the training phase 1602 is completed, the trained machine learning models 1632 can be provided to the computing devices, if not already on the computing devices. The inference phase 1604 can begin after the trained machine learning models 1632 are provided to the particular computing devices.
[0117] During the inference phase 1604, the trained machine learning model 1632 can receive input data 1630 and generate and output one or more corresponding inferences and / or predictions 1650 about the input data 1630. As such, the input data 1630 can be used as input to the trained machine learning model 1632 for providing corresponding inferences and / or predictions 1650 to the kernel components and non-kernel components. For example, the trained machine learning model 1632 can generate inferences and / or predictions 1650 in response to one or more inference / prediction requests 1640. In some examples, the trained machine learning model 1632 can be executed by a portion of other software. For example, the trained machine learning model 1632 can be executed by an inference or prediction daemon to be readily available to provide inferences and / or predictions upon request. The input data 1630 can include data from the particular computing device that executes the trained machine learning model 1632 and / or input data from one or more computing devices other than the particular computing device.
[0118] The input data 1630 can include a set of images provided by one or more sources. The set of images can include images of an object (such as a human face, where the images of the human face are taken under different lighting conditions), images of multiple objects, images that reside on the particular computing device, and / or other images. Other types of input data are also possible.
[0119] The inferences and / or predictions 1650 can include output images, output lighting models, output surface orientation maps, output light estimates, output commercial images, numerical values, and / or other output data produced by the trained machine learning model 1632 operating on the input data 1630 (and the training data 1610). In some examples, the trained machine learning model 1632 can use the output inferences and / or predictions 1650 as input feedback 1660. The trained machine learning model 1632 can also rely on past inferences as input in order to generate new inferences.
[0120] The convolutional neural networks 215, 225, 245 can be examples of machine learning algorithms 1620. After training, a trained version of the convolutional neural networks 215, 225, 245 can be an example of a trained machine learning model 1632. In this approach, an example of an inference / prediction request 1640 can be a request to predict a surface orientation map from an input image of an object, and a corresponding example of an inference and / or prediction 1650 can be an output surface orientation map. As another example, an example of an inference / prediction request 1640 can be a request to predict an ambient lighting from an input image of an object, and a corresponding example of an inference and / or prediction 1650 can be an output set of 3-D vectors predicting the application of a particular lighting direction to the input image. Further, for example, an example of an inference / prediction request 1640 can be a request to determine a merchandise image of an input image of an object, and a corresponding example of an inference and / or prediction 1650 can be an output image predicting the merchandise image as a per-pixel multiplication factor to be applied to the input image.
[0121] In some examples, a single computing device (“CD_SOLO”) can include a trained version of the convolutional neural network 215 after training of the convolutional neural network 215. Then, the computing device CD_SOLO can receive a request to predict a surface orientation map from a corresponding input image, and use the trained version of the convolutional neural network 215 to generate an output image predicting the surface orientation map. In some examples, a single computing device CD_SOLO can include a trained version of the convolutional neural network 225 after training of the convolutional neural network 225. Then, the computing device CD_SOLO can receive a request to predict a light direction from a corresponding input image, and use the trained version of the convolutional neural network 225 to generate an output image predicting the light direction. In some examples, a single computing device CD_SOLO can include a trained version of the convolutional neural network 245 after training of the convolutional neural network 245. Then, the computing device CD_SOLO can receive a request to predict a merchandise image from a corresponding input image, and use the trained version of the convolutional neural network 215 to generate an output image predicting the merchandise image.
[0122] In some examples, two or more computing devices, such as a first client device ("CD_CLI") and a server device ("CD_SRV"), can be used to provide an output image; for example, the first computing device CD_CLI can generate and send a request to the second computing device CD_SRV to apply a particular synthetic illumination to a corresponding input image. The CD_SRV can then use a trained version of the convolutional neural network 215, 225, 245 (possibly after training the convolutional neural network 215, 225, 245) to generate an output image that predicts the application of the particular synthetic illumination to the input image, and in response to the request from the CD_CLI. The CD_CLI can then provide the requested output image (e.g., using a user interface and / or display, a printed copy, an electronic communication, etc.) when the response to the request is received.
[0123] Example data network
[0124] Figure 17 A distributed computing architecture 1700 is depicted in accordance with example embodiments. The distributed computing architecture 1700 includes server devices 1708, 1710 configured to communicate with programmable devices 1704a, 1704b, 1704c, 1704d, 1704e via a network 1706. The network 1706 can correspond to a local area network (LAN), a wide area network (WAN), a WLAN, a WW AN, an enterprise intranet, the public Internet, or any other type of network configured to provide a communication path between networked computing devices. The network 1706 can also correspond to a combination of one or more LANs, WANs, enterprise intranets, and / or public Internets.
[0125] Although Figure 17Only five programmable devices are shown, but a distributed application architecture can serve tens, hundreds, or thousands of programmable devices. Moreover, the programmable devices 1704a, 1704b, 1704c, 1704d, 1704e (or any additional programmable devices) can be any kind of computing device, such as a mobile computing device, a desktop computer, a wearable computing device, a head-mounted device (HMD), a network terminal, a mobile computing device, etc. In some examples, as shown by the programmable devices 1704a, 1704b, 1704c, 1704e, the programmable devices can connect directly to the network 1706. In other examples, as shown by the programmable device 1704d, the programmable device can connect indirectly to the network 1706 via an associated computing device, such as the programmable device 1704c. In this example, the programmable device 1704c can act as an associated computing device to pass electronic communications between the programmable device 1704d and the network 1706. In other examples, as shown by the programmable device 1704e, the computing device can be part of and / or inside a vehicle (e.g., an automobile, a truck, a bus, a boat or ship, an airplane, etc.). In Figure 17 In other examples not shown in FIG. 17, the programmable devices can connect directly and indirectly to the network 1706.
[0126] The server devices 1708, 1710 can be configured to perform one or more services requested by the programmable devices 1704a-1704e. For example, the server devices 1708 and / or 1710 can provide content to the programmable devices 1704a-1704e. The content can include, but is not limited to, web pages, hypertext, scripts, binary data such as compiled software, images, audio, and / or video. The content can include compressed and / or uncompressed content. The content can be encrypted and / or unencrypted. Other types of content are also possible.
[0127] As another example, the server devices 1708 and / or 1710 can provide the programmable devices 1704a-1704e with access to software for databases, searching, computing, graphics, audio, video, web / internet applications, and / or other functionality. Many other examples of server devices are also possible.
[0128] Computing device architecture
[0129] Figure 18 is a block diagram of an example computing device 1800 according to example embodiments. In particular, Figure 18 The computing device 1800 shown in FIG. 18 can be configured to perform at least one function of and / or at least one function related to a convolutional neural network, a predicted surface orientation map, a predicted raw illumination model, a predicted merchant image, a convolutional neural network 215, 225, 245, and / or the method 2000.
[0130] Computing device 1800 can include a user interface module 1801, a network communication module 1802, one or more processors 1803, a data storage 1804, one or more cameras 1818, one or more sensors 1820, and a power system 1822, all of which can be linked together via system bus, network, or other connection mechanism 1805.
[0131] User interface module 1801 can operate to send data to and / or receive data from external user input / output devices. For example, user interface module 1801 can be configured to send and / or receive data to and / or from user input devices such as touchscreens, computer mice, keyboards, keypads, touchpads, trackballs, joysticks, voice recognition modules, and / or other similar devices. User interface module 1801 can also be configured to provide output to user display devices such as one or more cathode ray tubes (CRTs), liquid crystal displays, light-emitting diodes (LEDs), displays using digital light processing (DLP) technology, printers, light bulbs, and / or other similar devices now known or later developed. User interface module 1801 can also be configured to generate audible output using devices such as speakers, speaker jacks, audio output ports, audio output devices, headphones, and / or other similar devices. User interface module 1801 can also be configured with one or more haptic devices that can generate haptic output such as vibrations and / or other output that can be detected through touch and / or physical contact with computing device 1800. In some examples, user interface module 1801 can be used to provide a graphical user interface (GUI) for utilizing computing device 1800, such as, for example Figure 15 the illustrated graphical user interface.
[0132] Network communication module 1802 can include one or more devices that provide one or more wireless interfaces 1807 and / or one or more wired interfaces 1808 that can be configured to communicate via a network. Wireless interfaces 1807 can include one or more wireless transmitters, receivers, and / or transceivers, such as, for example, Bluetooth TM (Bluetooth) transceivers, Wi-Fi TM WiMAX TM LTE TMThe transceiver and / or other types of wireless transceivers configurable to communicate via wireless networks. The wired interface 1808 can include one or more wired transmitters, receivers, and / or transceivers, such as an Ethernet transceiver, a Universal Serial Bus (USB) transceiver, or similar transceivers configurable to communicate via twisted pair wires, a coaxial cable, a fiber optic link, or similar physical connections to a wired network.
[0133] In some examples, the network communication module 1802 can be configured to provide reliable, secure, and / or authenticated communications. For each communication described herein, information for facilitating reliable communications (e.g., guaranteed message delivery) can be provided, possibly as part of a message header and / or trailer (e.g., packet / message sequencing information, encapsulation headers and / or trailers, size / time information, and transmission verification information such as a cyclic redundancy check (CRC) and / or parity value). Communications can be secured (e.g., encoded or encrypted) and / or decrypted / decoded using one or more cryptographic protocols and / or algorithms, such as but not limited to the Data Encryption Standard (DES), the Advanced Encryption Standard (AES), the Rivest-Shamir-Adelman (RSA) algorithm, the Diffie-Hellman algorithm, the Secure Sockets Layer (SSL) or Transport Layer Security (TLS) secure sockets protocol, and / or the Digital Signature Algorithm (DSA), among others. Other encryption protocols and / or algorithms can also be used, or in addition to those listed herein, to secure (and then decrypt / decode) communications.
[0134] The one or more processors 1803 can include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processors, tensor processing units (TPUs), graphics processing units (GPUs), application specific integrated circuits, etc.). The one or more processors 1803 can be configured to execute computer-readable instructions 1806 contained in the data storage 1804 and / or other instructions described herein.
[0135] The data storage 1804 can include one or more non-transitory computer- readable storage media (media) and / or storage architecture which can be read and / or accessed by at least one of the one or more processors 1803. The one or more computer-readable storage media can include volatile and / or non- volatile, removable and / or non-removable media implemented in a method or technology for storage and / or access of information, such as computer readable instructions, data structures, program modules, logic elements / circuits, or other data. The one or more computer- readable storage media can be implemented using a single physical device or article
[0136] The data storage 1804 can include computer readable instructions 1806 and possibly additional data. In some examples, the data storage 1804 can include storage required to perform at least a portion of the methods, scenarios, and techniques described herein and / or at least a portion of the functionality of the devices and networks described herein. In some examples, the data storage 1804 can include storage for a trained neural network model 1812 (e.g., a model of a trained convolutional neural network such as the convolutional neural networks 215, 225, 245). In particular in these examples, the computer readable instructions 1806 can include instructions that, when executed by the processor 1803, enable the computing device 1800 to provide some or all of the functionality of the trained neural network model 1812.
[0137] In some examples, the computing device 1800 can include one or more cameras 1818. The camera 1818 can include one or more image capture devices, such as still and / or video cameras, equipped to capture light and record the captured light in one or more images; that is, the camera 1818 can generate images of captured light. The one or more images can be one or more still images and / or one or more images utilized in a video image. The camera 1818 can capture light and / or electromagnetic radiation as visible light, infrared radiation, ultraviolet light, and / or one or more other frequencies of light emission.
[0138] In some examples, the computing device 1800 can include one or more sensors 1820. The sensors 1820 can be configured to measure conditions within the computing device 1800 and / or conditions in the environment of the computing device 1800 and provide data regarding those conditions. For example, the sensors 1820 can include one or more of: (i) sensors for obtaining data regarding the computing device 1800, such as, but not limited to, a thermometer for measuring a temperature of the computing device 1800, a battery sensor for measuring power of one or more batteries of the power system 1822, and / or other sensors that measure conditions of the computing device 1800; (ii) identification sensors that identify other objects and / or devices, such as, but not limited to, a radio frequency identification (RFID) reader, a proximity sensor, a one-dimensional barcode reader, a two-dimensional barcode (e.g., a quick response (QR) code) reader, and a laser tracker, where the identification sensors can be configured to read identifiers, such as RFID tags, barcodes, QR codes, and / or other devices and / or objects configured to read and provide at least identification information; (iii) sensors that measure a position and / or movement of the computing device 1800, such as, but not limited to, a tilt sensor, a gyroscope, an accelerometer, a Doppler sensor, a GPS device, a sonar sensor, a radar device, a laser displacement sensor, and a compass; (iv) environmental sensors for obtaining data indicative of an environment of the computing device 1800, such as, but not limited to, an infrared sensor, an optical sensor, a light sensor, a biological sensor, a capacitive sensor, a touch sensor, a temperature sensor, a wireless sensor, a radio sensor, a motion sensor, a microphone, a sound sensor, an ultrasonic sensor, and / or a smoke sensor; and / or (v) force sensors that measure one or more forces (e.g., inertial forces and / or gravitational forces) acting around the computing device 1800, such as, but not limited to, one or more sensors that measure one or more of a one-dimensional or multi-dimensional force, a torque, a ground force, a frictional force, and / or a zero moment point (ZMP) sensor that identifies a ZMP and / or a ZMP position. Many other examples of sensors 1820 are also possible.
[0139] The power system 1822 can include one or more batteries 1824 and / or one or more external power interfaces 1826 for providing power to the computing device 1800. Each of the one or more batteries 1824 can act as a source of stored power for the computing device 1800 when electrically coupled to the computing device 1800. The one or more batteries 1824 of the power system 1822 can be configured to be portable. Some or all of the one or more batteries 1824 can be easily removed from the computing device 1800. In other examples, some or all of the one or more batteries 1824 can be internal to the computing device 1800 and thus can not be easily removed from the computing device 1800. Some or all of the one or more batteries 1824 can be rechargeable. For example, a rechargeable battery can be charged via a wired connection between the battery and another power source, such as by one or more power sources that are external to the computing device 1800 and that are connected to the computing device 1800 via the one or more external power interfaces. In other examples, some or all of the one or more batteries 1824 can be non-rechargeable batteries.
[0140] The one or more external power interfaces 1826 of the power system 1822 can include one or more wired power interfaces, such as a USB cable and / or a power cord, that enable a wired power source to be connected to one or more power sources that are external to the computing device 1800. The one or more external power interfaces 1826 can include one or more wireless power interfaces, such as a Qi wireless charger, that enable a wireless power connection to one or more external power sources, for example, via the Qi wireless charger. Once a power connection to an external power source is established using the one or more external power interfaces 1826, the computing device 1800 can draw power from the external power source through the established power connection. In some examples, the power system 1822 can include related sensors, such as a battery sensor associated with the one or more batteries or other types of power sensors.
[0141] Cloud-based server
[0142] Figure 19 A network 1706 of computing clusters 1909a, 1909b, 1909c arranged as a cloud-based server system is depicted in accordance with example embodiments. The computing clusters 1909a, 1909b, 1909c can be cloud-based devices that store program logic and / or data of cloud-based applications and / or services, such as performing at least one function of and / or related to the convolutional neural networks 215, 225, 245 and / or the method 2400.
[0143] In some embodiments, the computing clusters 1909a, 1909b, 1909c can be a single computing device that resides at a single computing center. In other embodiments, the computing clusters 1909a, 1909b, 1909c can include multiple computing devices in a single computing center, or even multiple computing devices in multiple computing centers located in different geographic locations. For example, Figure 19 Each of the computing clusters 1909a, 1909b, and 1909c are depicted as residing at different physical locations.
[0144] In some embodiments, the data and services at the computing clusters 1909a, 1909b, 1909c can be encoded as computer-readable information stored in a non-transitory, tangible computer-readable medium (or computer-readable storage medium), and accessible to other computing devices. In some embodiments, the computing clusters 1909a, 1909b, 1909c can be stored on a single disk drive or other tangible storage medium, or can be implemented across multiple disk drives or other tangible storage media located in one or more different geographic locations.
[0145] Figure 19 A cloud-based server system according to example embodiments is depicted. In Figure 19 In some embodiments, the functionality of the convolutional neural networks 215, 225, 245 and / or the computing devices can be distributed among the computing clusters 1909a, 1909b, 1909c. The computing cluster 1909a can include one or more computing devices 1900a, a cluster storage array 1910a, and a cluster router 1911a connected by a local cluster network 1912a. Similarly, the computing cluster 1909b can include one or more computing devices 1900b, a cluster storage array 1910b, and a cluster router 1911b connected by a local cluster network 1912b. Likewise, the computing cluster 1909c can include one or more computing devices 1900c, a cluster storage array 1910c, and a cluster router 1911c connected by a local cluster network 1912c.
[0146] In some embodiments, each of the computing clusters 1909a, 1909b, and 1909c can have the same number of computing devices, the same number of cluster storage arrays, and the same number of cluster routers. However, in other embodiments, each of the computing clusters can have a different number of computing devices, a different number of cluster storage arrays, and a different number of cluster routers. The number of computing devices, cluster storage arrays, and cluster routers in each of the computing clusters can depend on the one or more computing tasks assigned to each of the computing clusters.
[0147] For example, in computing cluster 1909a, computing device 1900a can be configured to perform a convolutional neural network, a confidence learning, and / or various computing tasks of a computing device. In one embodiment, the convolutional neural network, the confidence learning, and / or various functions of the computing device can be distributed among one or more of computing devices 1900a, 1900b, 1900c. Computing devices 1900b and 1900c in respective computing clusters 1909b and 1909c can be configured similarly to computing device 1900a in computing cluster 1909a. On the other hand, in some embodiments, computing devices 1900a, 1900b, and 1900c can be configured to perform different functions.
[0148] In some embodiments, computing tasks and stored data associated with a convolutional neural network and / or a computing device can be distributed among computing devices 1900a, 1900b, and 1900c based at least in part on processing requirements of the convolutional neural network and / or the computing device, processing capabilities of computing devices 1900a, 1900b, 1900c, latency of network links among computing devices in each computing cluster and among computing clusters themselves, and / or other factors that can contribute to cost, speed, fault tolerance, resilience, efficiency, and / or other design goals of the overall system architecture.
[0149] Cluster storage arrays 1910a, 1910b, 1910c of computing clusters 1909a, 1909b, 1909c can be data storage arrays that include disk array controllers configured to manage read and write access to groups of hard disk drives. The disk array controllers (individually or in conjunction with their respective computing devices) can also be configured to manage backup or redundant copies of data stored in the cluster storage arrays to prevent disk drives of one or more computing clusters from being inaccessible to one or more computing devices and / or other cluster storage array failures and / or network failures.
[0150] Similar to the manner in which functions of a convolutional neural network and / or a computing device can be distributed across computing devices 1900a, 1900b, 1900c of computing clusters 1909a, 1909b, 1909c, various active portions and / or backup portions of these components can be distributed across cluster storage arrays 1910a, 1910b, 1910c. For example, some cluster storage arrays can be configured to store a portion of data of a convolutional neural network and / or a computing device, while other cluster storage arrays can store other portions of data of the convolutional neural network and / or the computing device. Further, for example, some cluster storage arrays can be configured to store data of a first convolutional neural network, while other cluster storage arrays can store data of a second and / or third convolutional neural network. Further, some cluster storage arrays can be configured to store backup versions of data stored in other cluster storage arrays.
[0151] Cluster routers 1911a, 1911b, 1911c in computing clusters 1909a, 1909b, 1909c can include networking equipment configured to provide internal and external communications for the computing clusters. For example, cluster router 1911a in computing cluster 1909a can include one or more internet exchange and routing devices configured to provide (i) local area network communications between computing devices 1900a and cluster storage array 1910a via local cluster network 1912a, and (ii) wide area network communications between computing cluster 1909a and computing clusters 1909b and 1909c via wide area network link 1913a to network 1706. Cluster routers 1911b and 1911c can include network equipment similar to cluster router 1911a, and cluster routers 1911b and 1911c can perform networking functions for computing clusters 1909b and 1909b similar to the networking functions performed by cluster router 1911a for computing cluster 1909a.
[0152] In some embodiments, the configuration of cluster routers 1911a, 1911b, 1911c can be based at least in part on data communication needs of the computing devices and cluster storage array, data communication capabilities of the network equipment in cluster routers 1911a, 1911b, 1911c, latency and throughput of local cluster networks 1912a, 1912b, 1912c, latency, throughput, and cost of wide area network links 1913a, 1913b, 1913c, and / or other factors that can contribute to cost, speed, fault tolerance, resilience, efficiency, and / or other design criteria that moderate the system architecture.
[0153] Example method of operation
[0154] Figure 20 is a flowchart of a method 2000 according to an example embodiment. Method 2000 can be performed by a computing device such as computing device 1700. Method 2000 can begin at block 2010, where the computing device can apply a geometric model to an input image to determine a surface orientation map indicative of a distribution of illumination on an object in the input image based on surface geometry of the object, such as discussed above at least in the context of Figures 2-5 .
[0155] At block 2020, the computing device can apply an ambient light estimation model to the input image to determine a direction of synthetic illumination to be applied to the input image to enhance at least a portion of the input image, such as discussed above at least in the context of Figure 2 , Figure 6 and Figure 7 .
[0156] At block 2030, the computing device can apply a light energy model based on the surface orientation map and the direction of the synthetic lighting to determine a quotient image indicating an amount of light energy to apply to each pixel of the input image, such as discussed above at least in the context of Figure 2 、 Figures 9-11 .
[0157] At block 2040, the computing device can enhance the portion of the input image based on the quotient image, such as discussed above at least in the context of Figures 2-15 .
[0158] Some embodiments involve training a neural network to perform one or more of: (1) applying a geometry model to predict a surface orientation map; (2) applying an ambient light estimation model to determine a direction of synthetic lighting; or (3) applying a light energy model to predict a quotient image. In some embodiments, the neural network is trained to apply a geometry model to predict a surface orientation map, and embodiments can include predicting a surface orientation map by using the trained neural network. In some embodiments, the neural network is trained to apply an ambient light estimation model to determine a direction of synthetic lighting, and embodiments can include predicting a direction of synthetic lighting by using the trained neural network. In some embodiments, the neural network is trained to apply a light energy model to predict a quotient image, and embodiments can include predicting a quotient by using the trained neural network.
[0159] In some embodiments, applying an ambient light estimation model includes detecting, by the computing device, a pose of an object in the input image. Determining a direction of synthetic lighting can be based on the pose.
[0160] Some embodiments include generating, by the computing device, a light visibility map based on the surface orientation map and the direction of the synthetic lighting. The quotient image can be based on the light visibility map.
[0161] In some embodiments, enhancing the portion of the input image includes determining, by the computing device, a request to enhance the portion of the input image. Such embodiments can further include sending the request to enhance the portion of the input image from the computing device to a second computing device, the second computing device including the trained neural network. Such embodiments can further include, after sending the request, the computing device receiving an output image from the second computing device, the output image applying the quotient image to enhance the portion of the input image.
[0162] In some embodiments, training the neural network includes using a plurality of images of the object. The plurality of images can illuminate the object with a plurality of lighting profiles.
[0163] Some embodiments include receiving, by the computing device, a user preference for a direction of synthetic lighting to apply to a particular input image.
[0164] In some embodiments, the object can have a property of diffusely reflecting light.
[0165] In some embodiments, the computing device can include a camera. Such embodiments include generating, using the camera, an input image of the object. Such embodiments can also include receiving, at the computing device, the generated input image from the camera.
[0166] Some embodiments include providing, using the computing device, the portion of the enhanced input image.
[0167] Some embodiments include predicting, using an ambient light estimation model, an illumination profile for the input image. Such embodiments include providing, using the computing device, the predicted illumination profile.
[0168] In some embodiments, predicting the illumination profile includes generating a high dynamic range (HDR) lighting environment based on low dynamic range (LDR) images of a set of reference objects, each reference object of the set of reference objects having a respective bidirectional reflectance distribution function (BRDF). In such embodiments, the set of reference objects can include one or more of a specular sphere, a matte silver sphere, or a gray diffuse sphere.
[0169] In some embodiments, predicting includes obtaining, at the computing device, the trained neural network. These embodiments include applying, at the computing device, the obtained neural network.
[0170] In some embodiments, training the neural network includes training, at the computing device, the neural network.
[0171] Some embodiments include adjusting the merchant image to apply one or more of the following to a particular input image: (i) compensation for exposure levels in the input image, (ii) compensation for brightness levels in the input image, or (iii) matting refinement.
[0172] Additional example embodiments
[0173] The following clauses are provided as further description of the disclosure.
[0174] Clause 1. A computer-implemented method comprising: training a neural network to apply an ambient light estimation model to an input image to optimize a direction of synthetic lighting to be applied to the input image to enhance at least a portion of the input image; detecting, by a computing device, a pose of an object in a particular input image; and predicting, by the computing device, based on the pose, a direction of synthetic lighting to be applied to the particular input image by using the trained neural network to apply the predicted ambient light estimation model to the particular input image.
[0175] Clause 2. The computer-implemented method of clause 1, further comprising applying the optimized direction of synthetic lighting to the particular input image to enhance the portion of the particular input image.
[0176] Clause 3. The computer-implemented method of clause 2, wherein applying the optimized direction of synthetic lighting to the particular input image comprises determining, by the computing device, a request to enhance the particular input image; sending the request to enhance the particular input image from the computing device to a second computing device, the second computing device comprising the trained neural network; and after sending the request, the computing device receiving, from the second computing device, an output image of applying the optimized direction of synthetic lighting to the particular input image.
[0177] Clause 4. The computer-implemented method of any of clauses 1-3, wherein the lighting profile of the input image is modeled using a light estimation model, and wherein predicting an optimal direction of added synthetic lighting further comprises estimating an original lighting profile of the input image.
[0178] Clause 5. The computer-implemented method of any of clauses 1-4, wherein the particular input image comprises a portrait, and the method further comprises detecting a face in the portrait; and determining a pose of the subject by determining a head pose of the face.
[0179] Clause 6. The computer-implemented method of any of clauses 1-5, wherein training the neural network to apply the ambient light estimation model comprises training the neural network using a plurality of images of the subject, wherein the plurality of images illuminate the subject with a plurality of lighting profiles.
[0180] Clause 7. The computer-implemented method of any of clauses 1-6, wherein the neural network is a convolutional neural network.
[0181] Clause 8. The computer-implemented method of any of clauses 1-7, wherein the subject comprises a light diffusing object.
[0182] Clause 9. The computer-implemented method of any of clauses 1-8, wherein the computing device comprises a camera, and the method further comprises generating the particular input image using the camera; and receiving the generated particular input image at the computing device from the camera.
[0183] Clause 10. The computer-implemented method of any of clauses 1-9, further comprising providing the optimized direction of synthetic lighting using the computing device.
[0184] Clause 11. The computer-implemented method of any of clauses 1-10, wherein the lighting profile of the input image is estimated using a light estimation model, and wherein the method further comprises providing, using the computing device, a prediction of the raw lighting profile.
[0185] Clause 12. The computer-implemented method of clause 11, wherein predicting the lighting profile comprises generating a high dynamic range (HDR) illumination environment based on a set of reference objects, each reference object of the set of reference objects having a corresponding bidirectional reflectance distribution function (BRDF).
[0186] Clause 13. The computer-implemented method of clause 12, wherein the set of reference objects comprises one or more of a specular sphere, a matte silver sphere, or a gray diffuse sphere.
[0187] Clause 14. The computer-implemented method of any of clauses 1-13, wherein optimizing the direction of the synthetic lighting by using the trained neural network comprises obtaining, at the computing device, the trained neural network; and determining, by the computing device, the direction of the synthetic lighting using the obtained neural network.
[0188] Clause 15. The computer-implemented method of clause 14, wherein training the neural network comprises training, at the computing device, the neural network.
[0189] Clause 16. The computer-implemented method of any of clauses 1-15, further comprising applying, by the computing device, a geometric model to the particular input image to determine a surface orientation map indicative of a lighting distribution on the object in the particular input image based on surface geometry of the object.
[0190] Clause 17. The computer-implemented method of any of clauses 1-16, further comprising applying, by the computing device, a light energy model based on the surface geometry of the object in the particular image and the optimized direction of the synthetic lighting to determine a merchandise image indicative of an amount of light energy to be applied to each pixel of the particular input image.
[0191] Clause 18. The computer-implemented method of clause 17, further comprising enhancing the portion of the particular input image based on the merchandise image.
[0192] Clause 19. A computing device comprising: one or more processors; and data storage, wherein the data storage has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to perform functions comprising the computer-implemented method of any of clauses 1-18.
[0193] Clause 20. An article of manufacture including one or more computer-readable media having stored thereon computer-readable instructions that, when executed by one or more processors of a computing device, cause the computing device to perform functions including the computer-implemented method of any of clauses 1-18.
[0194] Clause 21. The article of manufacture of clause 20, wherein the one or more computer- readable media comprise one or more non-transitory computer-readable media.
[0195] Clause 22. A computing device comprising means for performing the computer- implemented method of any of clauses 1-18.
[0196] Clause 23. A computer-implemented method comprising: training a neural network to apply a geometric model to an input image to predict, based on surface geometry of an object, a surface orientation map indicative of a distribution of illumination on the object in the input image; predicting, by a computing device, the surface orientation map based on particular surface geometry of a particular object by applying the geometric model to a particular object in a particular input image using the trained neural network; receiving, by the computing device, an indication of a direction of synthetic illumination to be applied to the particular input image to enhance at least a portion of the particular input image; and generating, by the computing device, a light visibility map for the particular input image based on the predicted surface orientation map and the indicated direction of synthetic illumination, wherein the light visibility map is indicative of the synthetic illumination to be applied to the particular input image based on the particular surface geometry of the particular object.
[0197] Clause 24. The computer-implemented method of clause 23, further comprising: applying the generated light visibility map to the particular input image to enhance at least a portion of the particular input image.
[0198] Clause 25. The computer-implemented method of clause 24, wherein applying the generated light visibility map to the particular input image comprises: determining, by the computing device, a request to enhance the particular input image; sending the request to enhance the particular input image from the computing device to a second computing device, the second computing device including the trained neural network; and after sending the request, receiving, by the computing device from the second computing device, an output image of the particular input image with the generated light visibility map applied to the particular input image.
[0199] Clause 26. The computer-implemented method of any of clauses 23-25, wherein training the neural network to apply the geometric model comprises: training the neural network using a plurality of images of the object, wherein the plurality of images illuminate the object using a plurality of lighting profiles.
[0200] Clause 27. The computer-implemented method of any of clauses 23-26, wherein the neural network is a convolutional neural network.
[0201] Clause 28. The computer-implemented method of any of clauses 23-27, wherein the object comprises an object that diffusely reflects light.
[0202] Clause 29. The computer-implemented method of any of clauses 23-28, wherein the computing device comprises a camera, and the method further comprises: generating, using the camera, a particular input image of the object; and receiving, at the computing device from the camera, the generated particular input image.
[0203] Clause 30. The computer-implemented method of any of clauses 23-29, wherein receiving the indication of the direction of the synthetic lighting comprises: receiving, by the computing device, a user preference of the direction of the synthetic lighting.
[0204] Clause 31. The computer-implemented method of any of clauses 23-29, wherein receiving the indication of the direction of the synthetic lighting comprises: training a second neural network to apply an ambient light estimation model to the input image to optimize the direction of the synthetic lighting; detecting, by the computing device, a pose of the particular object in the particular input image; and optimizing, by the computing device, the direction of the synthetic lighting based on the pose by applying, using the trained second neural network, the ambient light estimation model to the particular input image.
[0205] Clause 32. The computer-implemented method of any of clauses 23-29, further comprising: providing, using the computing device, the predicted surface orientation map.
[0206] Clause 33. The computer-implemented method of any of clauses 23-32, wherein the lighting profile of the input image is modeled using a raw light model, and wherein the method further comprises: providing, using the computing device, a prediction of the raw light model.
[0207] Clause 34. The computer-implemented method of clause 33, wherein predicting the lighting profile comprises: generating a high dynamic range (HDR) lighting environment based on a set of reference objects, each reference object in the set of reference objects having a corresponding bidirectional reflectance distribution function (BRDF), based on a low dynamic range (LDR) image of the set of reference objects.
[0208] Clause 35. The computer-implemented method of clause 34, wherein the set of reference objects comprises one or more of a specular sphere, a matte silver sphere, or a gray diffuse sphere.
[0209] Clause 36. The computer-implemented method of any of clauses 23-35, wherein predicting the surface orientation map by using the trained neural network comprises: obtaining, at the computing device, the trained neural network; and determining, by the computing device, the surface orientation map using the obtained neural network.
[0210] Clause 37. The computer-implemented method of clause 36, wherein training the neural network comprises: training, at the computing device, the neural network.
[0211] Clause 38. The computer-implemented method of any of clauses 23-37, further comprising: applying, by the computing device, a light energy model based on the surface orientation map and a direction of the synthetic lighting to determine a merchandise image indicating an amount of light energy to be applied to each pixel of the particular input image.
[0212] Clause 39. The computer-implemented method of clause 38, further comprising: enhancing, based on the merchandise image, the portion of the input image.
[0213] Clause 40. A computing device comprising: one or more processors; and data storage, wherein the data storage has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to perform functions comprising the computer-implemented method of any of clauses 23-39.
[0214] Clause 41. An article of manufacture including one or more computer-readable media having stored thereon computer-readable instructions that, when executed by one or more processors of a computing device, cause the computing device to perform functions comprising the computer-implemented method of any of clauses 23-39.
[0215] Clause 42. The article of manufacture of clause 41, wherein the one or more computer- readable media comprise one or more non-transitory computer-readable media.
[0216] Clause 43. A computing device comprising means for performing the computer- implemented method of any of clauses 23-39.
[0217] Clause 44. A computer-implemented method comprising: training a neural network to apply a light energy model to an input image to predict a merchandise image indicating an amount of light energy to be applied to each pixel of the input image; receiving, by a computing device, a light visibility map for a particular input image, wherein the light visibility map indicates synthetic lighting to be applied to the particular input image based on surface geometry of objects in the particular input image; and predicting, by the computing device, the merchandise image by using the trained neural network to apply the light energy model to the particular input image.
[0218] Clause 45. The computer-implemented method of clause 44, further comprising applying the predicted merchant image to the particular input image to enhance at least a portion of the particular input image.
[0219] Clause 46. The computer-implemented method of any of clauses 44 or 45, further comprising adjusting the merchant image to compensate for an exposure level in the particular input image.
[0220] Clause 47. The computer-implemented method of any of clauses 44-46, further comprising adjusting the merchant image to compensate for a brightness level in the particular input image.
[0221] Clause 48. The computer-implemented method of any of clauses 44-47, further comprising adjusting the merchant image to apply a matting refinement to the particular input image.
[0222] Clause 49. The computer-implemented method of any of clauses 45-49, wherein predicting the merchant image comprises: determining, by the computing device, a request to enhance the particular input image; sending the request to enhance the particular input image from the computing device to a second computing device, the second computing device comprising the trained neural network; and after sending the request, the computing device receiving, from the second computing device, an output image that applies the predicted merchant image to the particular input image.
[0223] Clause 50. The computer-implemented method of any of clauses 44-49, wherein training the neural network to apply the light energy model comprises training the neural network using a plurality of images of the object, wherein the plurality of images illuminate the object using a plurality of lighting profiles.
[0224] Clause 51. The computer-implemented method of any of clauses 44-50, wherein the neural network is a convolutional neural network.
[0225] Clause 52. The computer-implemented method of any of clauses 44-51, wherein the object comprises an object that diffusely reflects light.
[0226] Clause 53. The computer-implemented method of any of clauses 44-52, wherein the computing device comprises a camera, and the method further comprises: generating, using the camera, the particular input image of the object; and receiving, at the computing device, the generated particular input image from the camera.
[0227] Clause 54. The computer-implemented method of any of clauses 44-53, further comprising receiving, by the computing device, a user preference for a direction of synthetic lighting to apply to the particular input image.
[0228] Clause 55. The computer-implemented method of any of clauses 44-54, further comprising training a second neural network to apply the ambient light estimation model to the input image to optimize a direction of the synthetic lighting; detecting, by the computing device, a pose of the object in the particular input image; and optimizing, by the computing device, the direction of the synthetic lighting by applying the ambient light estimation model to the particular input image using the trained second neural network based on the pose.
[0229] Clause 56. The computer-implemented method of any of clauses 44-55, further comprising applying, by the computing device, a geometry model to the particular input image to determine a surface orientation map indicative of a lighting distribution on the object in the particular input image based on surface geometry of the object, and wherein the light visibility map is based on the surface orientation map.
[0230] Clause 57. The computer-implemented method of any of clauses 44-56, wherein the lighting profile of the input image is modeled using the ambient light estimation model, and wherein the method further comprises providing, using the computing device, a prediction of the lighting profile.
[0231] Clause 58. The computer-implemented method of any of clauses 44-57, wherein predicting the shading image by using the trained neural network comprises obtaining, at the computing device, the trained neural network; and determining, by the computing device, the shading image using the obtained neural network.
[0232] Clause 59. The computer-implemented method of clauses 44-58, wherein training the neural network comprises training, at the computing device, the neural network.
[0233] Clause 60. A computing device comprising: one or more processors; and data storage, wherein the data storage has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to perform functions comprising the computer-implemented method of any of clauses 44-59.
[0234] Clause 61. An article of manufacture comprising one or more computer-readable media having computer-readable instructions stored thereon that, when executed by one or more processors of a computing device, cause the computing device to perform functions comprising the computer-implemented method of any of clauses 44-59.
[0235] Clause 62. The article of manufacture of clause 61, wherein the one or more computer- readable media comprise one or more non-transitory computer-readable media.
[0236] Clause 63. A computing device comprising means for performing the computer- implemented method of any of clauses 44-59.
[0237] Clause 64. A computer-implemented method comprising: applying, by a computing device, a geometry model to an input image to determine, based on surface geometry of an object, a surface orientation map indicative of a distribution of illumination on the object in the input image; applying, by the computing device, an ambient light estimation model to the input image to determine an optimal direction of synthetic illumination to be applied to the input image to enhance at least a portion of the input image; applying, by the computing device, a light energy model based on the surface orientation map and the optimal direction of synthetic illumination to determine a quotient image indicative of an amount of light energy to be applied to each pixel of the input image; and enhancing the portion of the input image based on the quotient image.
[0238] Clause 65. The computer-implemented method of clause 64, wherein applying the geometry model comprises: training a neural network to apply the geometry model to predict the surface orientation map; and predicting the surface orientation map by using the trained neural network.
[0239] Clause 66. The computer-implemented method of any of clauses 64 or 65, wherein applying the ambient light estimation model comprises: training a neural network to apply the ambient light estimation model to optimize the direction of synthetic illumination; and optimizing, by the computing device, the direction of synthetic illumination using the trained neural network.
[0240] Clause 67. The computer-implemented method of any of clauses 64-66, wherein applying the light energy model comprises: training a neural network to apply the light energy model to predict the quotient image; and predicting the quotient using the trained neural network.
[0241] Clause 68. The computer-implemented method of any of clauses 64-67, further comprising: generating, by the computing device, a light visibility map based on the surface orientation map and the direction of synthetic illumination, and wherein the quotient image is based on the light visibility map.
[0242] Clause 69. The computer-implemented method of any of clauses 64-68, wherein enhancing the portion of the input image comprises: determining, by the computing device, a request to enhance the input image; sending the request to enhance the input image from the computing device to a second computing device, the second computing device comprising the trained neural network; and after sending the request, the computing device receiving, from the second computing device, an output image that applies the quotient image to enhance the portion of the input image.
[0243] Clause 70. The computer-implemented method of any of clauses 65-68, further comprising: training the neural network using a plurality of images of the object, wherein the plurality of images illuminate the object using a plurality of lighting profiles.
[0244] Clause 71. The computer-implemented method of any of clauses 65-68 or 70, wherein the neural network is a convolutional neural network.
[0245] Clause 72. The computer-implemented method of any of clauses 64-71, wherein the object comprises an object that diffusely reflects light.
[0246] Clause 73. The computer-implemented method of any of clauses 64-72, wherein the computing device comprises a camera, and the method further comprises: generating, using the camera, an input image of the object; and receiving, at the computing device, the generated input image from the camera.
[0247] Clause 74. The computer-implemented method of any of clauses 64-73, further comprising: providing, using the computing device, an augmented input image.
[0248] Clause 75. The computer-implemented method of any of clauses 64-74, wherein the lighting profile of the input image is modeled using an ambient light estimation model, and wherein the method further comprises: providing, using the computing device, a prediction of the lighting profile.
[0249] Clause 76. The computer-implemented method of clause 75, wherein predicting the lighting profile comprises generating a high dynamic range (HDR) lighting environment based on a set of reference objects, each reference object in the set of reference objects having a corresponding bidirectional reflectance distribution function (BRDF).
[0250] Clause 77. The computer-implemented method of clause 76, wherein the set of reference objects comprises one or more of a specular sphere, a matte silver sphere, or a gray diffuse sphere.
[0251] Clause 78. The computer-implemented method of any of clauses 65-67 or 70-71, wherein predicting using the trained neural network comprises: obtaining, at the computing device, the trained neural network; and applying, at the computing device, the obtained neural network.
[0252] Clause 79. The computer-implemented method of any of clauses 65-66 or 70-71, wherein training the neural network comprises: training, at the computing device, the neural network.
[0253] Clause 80. A computing device comprising: one or more processors; and data storage, wherein the data storage has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to perform functions comprising the computer-implemented method of any of clauses 64-79.
[0254] Clause 81. An article of manufacture including one or more computer-readable media having stored thereon computer-readable instructions that, when executed by one or more processors of a computing device, cause the computing device to perform functions including the computer-implemented method of any of clauses 64-79.
[0255] Clause 82. The article of manufacture of clause 81, wherein the one or more computer- readable media comprise one or more non-transitory computer-readable media.
[0256] Clause 83. A computing device comprising means for performing the computer- implemented method of any of clauses 64-79.
[0257] The present disclosure is not limited to the particular embodiments described in this application, which are intended as illustrations of various aspects. Numerous modifications and variations are apparent to those skilled in the art in light of this description, which is to be understood as both an example and an embodiment. In addition to the examples listed here, other equivalents of the methods and apparatus within the scope of the disclosure will be apparent to those skilled in the art in view of the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims.
[0258] The above detailed description describes various features and functions of the disclosed systems, devices, and methods with reference to the accompanying drawings. In the drawings, like symbols typically identify corresponding components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments can be used, and other changes can be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and with reference to the accompanying drawings, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.
[0259] With respect to any or all of the diagrams, scenes flowcharts, and flowcharts discussed in the present document, according to an example embodiment, each block and / or communication can represent a processing, and / or a transfer of information. Alternative embodiments include those in which additional blocks and / or functions are added, some are removed, and / or some are combined. In some embodiments, the blocks and / or functions are performed in a different order than shown in the diagrams, scenes flowcharts, and flowcharts. In some embodiments, the blocks and / or functions are performed in parallel, rather than sequentially. In some embodiments, the blocks and / or functions are performed by different entities than shown in the diagrams, scenes flowcharts, and flowcharts. In some embodiments, the diagrams, scenes flowcharts, and flowcharts are combined with other diagrams, scenes flowcharts, and flowcharts. Figure One In some embodiments, the diagrams, scenes flowcharts, and flowcharts are combined with other diagrams, scenes flowcharts, and flowcharts.
[0260] The blocks representing information processing can correspond to electrical circuitry that can be configured to perform the specific logical functions described herein. Alternatively or additionally, the blocks representing information processing can correspond to modules, segments, or portions of program code (including associated data) that can be executed by a processor to implement the specific logical functions or acts in the methods or techniques. The program code and / or associated data can be stored on any type of computer readable medium, including storage devices such as magnetic or optical disks or other storage devices.
[0261] Computer readable media also can include non-transitory computer readable media, such as non-transitory computer readable media that store data for short periods of time, like register memory, processor cache and Random Access Memory (RAM). Computer readable media also can include non-transitory computer readable media that store program code and / or data for longer periods of time, like secondary or persistent long term memory, like read only memory (ROM), Optical disc or hard disc storage. The computer readable media can also be considered a computer readable storage medium, or a tangible storage device.
[0262] Further, the blocks that represent one or more information transmissions can correspond to information transmissions between software and / or hardware modules within the same physical device. However, other information transmissions can be between software modules and / or hardware modules in different physical devices.
[0263] While several aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of explanation and are not intended to be limiting, with the true scope being indicated by the appended claims.
Claims
1. A computer-implemented method comprising: applying, by a computing device, a geometry model to an input image to determine, based on surface geometry of an object, a surface orientation map indicative of a distribution of illumination on the object in the input image; applying, by the computing device, an ambient light estimation model to the input image to determine a direction of synthetic illumination to be applied to the input image to enhance at least a portion of the input image; applying, by the computing device, based on the surface orientation map and the direction of synthetic illumination, a light energy model to determine a merchandise image indicative of an amount of light energy to be applied to each pixel of the input image; and enhancing, based on the merchandise image, the portion of the input image.
2. The computer-implemented method of claim 1, further comprising: training a neural network to perform one or more of: (1) applying the geometry model to predict the surface orientation map, (2) applying the ambient light estimation model to determine the direction of synthetic illumination, or (3) applying the light energy model to predict the merchandise image. the neural network is trained to apply the geometry model to predict the surface orientation map, and the method further comprises:
3. The computer-implemented method of claim 2, wherein, predicting, using the trained neural network, the surface orientation map. the neural network is trained to apply the ambient light estimation model to determine the direction of synthetic illumination, and the method further comprises:
4. The computer-implemented method of claim 2, wherein, predicting, using the trained neural network, the direction of synthetic illumination. applying the ambient light estimation model comprises:
5. The computer-implemented method of any of claims 1-4, wherein, detecting, by the computing device, a pose of the object in the input image, and wherein the determination of the direction of synthetic illumination is based on the pose. the neural network is trained to apply the light energy model to predict the merchandise image, and the method further comprises:
6. The computer-implemented method of claim 2, wherein, predicting, using the trained neural network, the merchandise image.
7. The computer-implemented method of any one of claims 1-6, further comprising: generating, by the computing device, a light visibility map based on the surface orientation map and the direction of synthetic illumination, and wherein the merchandise image is based on the light visibility map. enhancing the portion of the input image comprises:
8. The computer-implemented method of any of claims 1-7, wherein, determining, by the computing device, a request to enhance the portion of the input image; sending, from the computing device to a second computing device, a request to enhance the portion of the input image, the second computing device comprising the trained neural network; and after sending the request, receiving, by the computing device from the second computing device, an output image that applies the merchandise image to enhance the portion of the input image. training the neural network comprises utilizing a plurality of images of the object, wherein the plurality of images illuminate the object with a plurality of lighting profiles.
9. The computer-implemented method of any of claims 2-8, wherein, applying the ambient light estimation model comprises:
10. The computer-implemented method of any of claims 1-9, wherein, receiving, by the computing device, a user preference for the direction of synthetic illumination to be applied to a particular input image. the object has a diffuse reflectance characteristic for light.
11. The computer-implemented method of any of claims 1-10, wherein, the computing device comprises a camera, and the method further comprises:
12. The computer-implemented method of any of claims 1-11, wherein, generating, using the camera, the input image of the object; and receiving, at the computing device, the generated input image from the camera.
13. The computer-implemented method of any one of claims 1-12, further comprising: providing, using the computing device, the enhanced portion of the input image.
14. The computer-implemented method of any one of claims 1-13, further comprising: predicting, using the ambient light estimation model, a lighting profile for the input image; and providing the predicted lighting distribution using a computing device.
15. The computer-implemented method of claim 14, wherein, predicting the lighting profile includes: generating a high dynamic range (HDR) lighting environment based on low dynamic range (LDR) images of a set of reference objects, each reference object of the set of reference objects having a corresponding bidirectional reflectance distribution function (BRDF).
16. The computer-implemented method of claim 15, wherein, the set of reference objects includes one or more of a specular sphere, a matte silver sphere, or a gray diffuse sphere.
17. The computer-implemented method of any one of claims 3, 4, 6, or 14, wherein, predicting includes: obtaining the trained neural network at a computing device; and applying the obtained neural network at the computing device.
18. The computer-implemented method of claim 2, wherein, training the neural network includes training the neural network at a computing device.
19. The computer-implemented method of any of claims 1-18, further comprising: adjusting the merchant image to apply one or more of the following to the particular input image: (i) compensation for exposure levels in the input image, (ii) compensation for brightness levels in the input image, or (iii) matting refinement.
20. A computing device comprising: one or more processors; and a data storage device, wherein the data storage device has stored thereon computer- executable instructions that, when executed by the one or more processors, cause the computing device to perform functions comprising the computer-implemented method of any of claims 1-19.
21. An article of manufacture comprising one or more computer-readable media having stored thereon computer-readable instructions that, when executed by one or more processors of a computing device, cause the computing device to perform functions comprising the computer- implemented method of any of claims 1-19.
Citation Information
Patent Citations
Implementation of an advanced image formation process as a network layer and its applications
CN108537864A
Computer system and method for improved gloss representation in digital images
CN109564701A