Photograph re-illumination using deep neural network and confidence learning

By using deep neural networks and confidence learning techniques, the convolutional neural network is trained to adjust image lighting, which solves the problem of poor image lighting adjustment in the prior art and improves image quality.

CN120070296APending Publication Date: 2025-05-30GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510038721.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-10-22
Filing Date
2019-04-01
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively adjust the lighting of captured images, especially in images of human faces and similar objects, resulting in poor image quality.

Method used

The deep neural network is used to train the convolutional neural network through confidence learning, receive input images and specific lighting model data, and generate adjusted output images to solve the problem of poor lighting.

Benefits of technology

Flexible adjustments to image lighting are achieved, the actual and perceived quality of the image is improved, and the portrait of the character looks more natural and clear.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070296A_ABST
    Figure CN120070296A_ABST
Patent Text Reader

Abstract

Apparatus and methods related to applying an illumination model to an image of an object are provided. A neural network may be trained to apply the illumination model to the input image. Training of the neural network may utilize confidence learning based on light predictions and predictive confidence values associated with illumination of the input image. A computing device may receive an input image of an object and data regarding a particular lighting model to be applied to the input image. The computing device may determine an output image of the object by applying the particular lighting model to an input image of the object using the trained neural network.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application for invention titled "Photo Relighting using Deep Neural Networks and Confidence Learning" with the international filing date of April 1, 2019, Chinese application number 201980060345.7.

[0002] Cross - reference to related applications

[0003] This application claims the priority of U.S. Patent Application No. 62 / 735,506, titled "Photo Relighting using Deep Neural Networks and Confidence Learning", filed on September 24, 2018, and U.S. Patent Application No. 62 / 749,081, titled "Photo Relighting using Deep Neural Networks and Confidence Learning", filed on October 22, 2018. For all purposes, the entire contents thereof are incorporated herein by reference in their entirety. Background art

[0004] Many modern computing devices (including mobile phones, personal computers, and tablet computers) include image - capturing devices, such as still and / or video cameras. The image - capturing devices can capture images, such as images including people, animals, landscapes, and / or objects.

[0005] Some image - capturing devices and / or computing devices can correct or modify the captured images. For example, some image - capturing devices can provide "red - eye" correction, which removes artifacts such as the red appearance of the eyes of people and animals that may occur in images captured using strong lights such as flash lighting. After the captured image is corrected, the corrected image can be saved, displayed, transmitted, printed on paper, and / or otherwise utilized.

[0006] In one aspect, a computer - implemented method is provided. A neural network is trained to apply a lighting model to an input image. The training of the neural network utilizes confidence learning based on light prediction and prediction confidence values associated with the lighting of the input image. A computing device receives an input image of an object and data regarding a specific lighting model to be applied to the input image. The computing device determines an output image of the object by applying the specific lighting model to the input image of the object using the trained neural network.

[0007] In another aspect, a computing device is provided. The computing device includes one or more processors and a data storage device. Computer-executable instructions are stored on the data storage device, which, when executed by the one or more processors, cause the computing device to perform functions. The functions include: training a neural network to apply an illumination model to an input image by using confidence learning based on light prediction and prediction confidence values associated with the illumination of the input image; receiving an input image of an object and data regarding a specific illumination model to be applied to the input image; and determining an output image of the object by applying the specific illumination model to the input image of the object by using the trained neural network.

[0008] In another aspect, an article of manufacture is provided. The article of manufacture includes one or more computer-readable media having computer-readable instructions stored thereon, which, when executed by one or more processors of a computing device, cause the computing device to perform functions. The functions include: training a neural network to apply an illumination model to an input image by using confidence learning based on light prediction and prediction confidence values associated with the illumination of the input image; receiving an input image of an object and data regarding a specific illumination model to be applied to the input image; and determining an output image of the object by applying the specific illumination model to the input image of the object by using the trained neural network.

[0009] In another aspect, a computing device is provided. The computing device includes: means for training a neural network to apply an illumination model to an input image by using confidence learning based on light prediction and prediction confidence values associated with the illumination of the input image; means for receiving an input image of an object and data regarding a specific illumination model to be applied to the input image; and means for determining an output image of the object by applying the specific illumination model to the input image of the object by using the trained neural network.

[0010] The foregoing summary is illustrative only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the figures and the following detailed description and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 An image with imperfect illumination in accordance with an example embodiment is shown.

[0012] Figure 2 and Figure 3 An example image for training a convolutional neural network to change image illumination in accordance with an example embodiment is shown.

[0013] Figure 4is a diagram depicting the training of a convolutional neural network for changing the illumination of an image according to an example embodiment.

[0014] Figure 5 is according to an example embodiment Figure 4 block diagram of the convolutional neural network.

[0015] Figure 6 is a diagram showing according to an example embodiment Figure 4 confidence learning used during the training of the convolutional neural network.

[0016] Figures 7 - 18 shows an example image of a face generated by Figure 4 the convolutional neural network according to an example embodiment.

[0017] Figure 19 is a diagram showing the training and inference phases of a machine learning model according to an example embodiment.

[0018] Figure 20 depicts a distributed computing architecture according to an example embodiment.

[0019] Figure 21 is a block diagram of a computing device according to an example embodiment.

[0020] Figure 22 depicts a network of a computing cluster arranged as a cloud-based server system according to an example embodiment.

[0021] Figure 23 is a flowchart of a method according to an example embodiment. DETAILED DESCRIPTION

[0022] This application relates to using machine learning techniques (such as, but not limited to, neural network techniques) to change the illumination of an image of an object (such as an object depicting a face). When a user of a mobile computing device takes an image of an object (such as a person), the resulting image may not always have ideal illumination. For example, the image may be too bright or too dark, the light may come from an unwanted direction, or the illumination may include different colors, which gives the image an unwanted tint. In addition, even if the image does have the desired illumination for a moment, the user may later want to change the illumination. Thus, there arises a technical problem related to image processing involving adjusting the illumination of an already acquired image.

[0023] To allow a user to control the illumination of an image, particularly an image of a face and similar objects, the techniques described herein apply a model based on a convolutional neural network to correct the illumination of the image. The techniques described herein include: receiving an input image and data regarding a specific illumination model to be applied to the input image; using a convolutional neural network to predict an output image applying the data regarding the specific illumination model to be applied to the input image; and generating an output based on the output image. The input and output images can be high-resolution images, such as images of several million pixels in size captured by a camera of a mobile computing device. The convolutional neural network can handle well the input images captured under various natural and artificial lighting conditions. In some examples, the model of the trained convolutional neural network can work on various computing devices, including but not limited to, mobile computing devices (e.g., smart phones, tablets, cellular phones, laptop computers), fixed computing devices (e.g., desktop computers), and server computing devices. Thus, the convolutional neural network can apply a specific illumination model to the input image, thereby adjusting the illumination of the input image and solving the technical problem of adjusting the illumination of an already acquired image.

[0024] The input image can be illuminated, and the illumination of the input image can be represented by an original illumination model. An illumination model, such as the above specific illumination model or the original illumination model, can be represented by a grid of lighting cells, where each lighting cell contains data related to the illumination of at least a portion of the corresponding image. The data related to the illumination of at least a portion of the corresponding image can represent one or more colors, intensities, albedos, directions, and / or surface normals of the illumination of that portion of the corresponding image. For example, an input image of 512×512 pixels can have a corresponding original illumination model represented by a 16×32 grid of lighting cells. Other examples of input images and grids of lighting cells are possible, e.g., images of different sizes, larger illumination models, etc. In addition, more, fewer, and / or different data can be stored in the lighting cells of the lighting cell grid.

[0025] A training dataset of images can be used to train a neural network, such as a convolutional neural network, to apply an illumination model to an image of an object, such as a face. In some examples, the neural network can be arranged as an encoder / decoder neural network.

[0026] Although the examples described herein relate to determining an illumination model and applying it to an image of an object with a face, a neural network can be trained to determine an illumination model and apply it to images of other objects, such as objects that reflect light similar to a face. Due to direct reflection of light, a face typically diffusely reflects light and has some specular highlights. For example, specular highlights may be caused by direct light reflection from the surface of the eyes, glasses, jewelry, etc. In many face images, the area ratio of such specular highlights is relatively small compared to the area of the facial surface that diffusely reflects light. Thus, a neural network can be trained to apply an illumination model to images of other objects that diffusely reflect light, and these diffusely reflecting objects can have some relatively small specular highlights (e.g., a tomato or a wall painted with a matte finish). Images in a training dataset can show one or more specific objects using illumination provided under multiple different conditions, such as illumination provided from different directions, illumination provided at different intensities (e.g., brighter and dimmer illumination), illumination provided by light sources of different colors, illumination provided by different numbers of light sources, etc.

[0027] Confidence learning can be used to train a neural network, where predictions can be weighted by one (or more) confidence values. For example, a neural network can generate light predictions for each of multiple "patches" or portions of an image. Then, as part of confidence learning, the light predictions for the patches can be mathematically combined (e.g., multiplied) with the confidence values of the predictions, which are also generated by the neural network. In some examples, the confidence values can be output values explicitly predicted and / or otherwise provided by the neural network. In some examples, the confidence values can be implicit values based on one or more weights of the neural network used to provide the predictions (e.g., a particular weight value, a maximum weight value, an average weight value, a minimum weight value, or some other weight value determined by combining some or all of the weights in the neural network). Confidence learning allows the neural network to weight illumination predictions based on the confidence values, resulting in better light predictions by weighting more confident predictions compared to less confident predictions, thereby improving the quality of the output illumination prediction and thus the quality of the predicted output image.

[0028] Once trained, the neural network can receive an input image and information about a desired illumination model. The trained neural network can process the input image to determine a prediction of the original illumination model of the illumination provided for capturing the input image. The trained neural network can also process the image to apply a specific illumination model to the original image and predict an output image where the specific illumination model has been applied to the input image. Then, the trained neural network can provide an output that includes the predicted output image and / or the predicted original illumination model. In some examples, the neural network can be trained to predict only the original illumination model or only the output image.

[0029] In one example, a trained neural network (a copy thereof) can reside on a mobile computing device. The mobile computing device can include a camera that can capture an input image of an object, such as a portrait of a person's face. A user of the mobile computing device can view the input image and determine that the input image should be relit. The user can then provide the input image and information about how the input image should be relit to the trained neural network residing on the mobile computing device. In response, the trained neural network can generate a predicted output image that shows the input image relit as indicated by the user, and then output the output image (e.g., provide the output image for display by the mobile computing device). In other examples, the trained neural network does not reside on the mobile computing device; rather, the mobile computing device provides the input image and information about how the input image should be relit to a trained neural network located remotely (e.g., via the Internet or another data network). The remotely located convolutional neural network can process the input image and information about how the input image should be relit as described above, and provide an output image to the mobile computing device that shows the input image relit as indicated by the user. In other examples, non-mobile computing devices can also use a trained neural network to relight an image, including an image that is not captured by the camera of the computing device.

[0030] In some examples, the trained neural network can work in conjunction with other neural networks (or other software) and / or be trained to identify whether an input image of an object is poorly lit. Then, when it is determined that the input image is poorly lit, the trained neural network described herein can apply a corrective lighting model to the poorly lit input image, thereby correcting the poor lighting of the input image. The corrective lighting model can be selected based on user input and / or can be predetermined. For example, a lighting model from user input or a predetermined lighting model can provide "flat light" or a lighting model that provides minimal contrast in the image to illuminate the object (e.g., a model that minimizes the difference between the "highlights" of bright lighting and the "shadows" of dim lighting in the image).

[0031] In some examples, a trained neural network can take an input image and one or more lighting models as input and provide one or more resulting output images. The trained neural network can then determine the one or more resulting output images by applying each of the plurality of lighting models to the input image. For example, the one or more lighting models can include a plurality of lighting models that represent one (or more) light sources that change position, lighting color, and / or other characteristics in each of the plurality of lighting models. More specifically, the plurality of lighting models can represent one or more light sources, where at least one light source changes position (e.g., by a predetermined amount) between the provided models. In such a scenario, the resulting output images represent the input image as shown when the changing light source appears to rotate or otherwise move around one (or more) objects depicted in the input image. Similarly, the changing light source can change color (e.g., a predetermined distance in color space) between the provided models such that the resulting output images represent the input image illuminated with light of various colors. The plurality of output images can be provided as still images and / or video images. By having the trained neural network apply a plurality of lighting models to an image (or, relatedly, having the trained neural network apply one lighting model to a plurality of input images), other effects can be produced.

[0032] In this way, the techniques described herein can improve images by applying more desirable and / or selectable lighting models to the images, thereby improving their actual and / or perceived quality. Improving the actual and / or perceived quality of images, including portraits of people, can provide an emotional benefit to those who believe their photos look better. These techniques are flexible, so a wide variety of lighting models can be applied to images of faces and other objects, particularly other objects with similar lighting characteristics. Additionally, by varying the lighting model, different aspects of the image can be highlighted, which can lead to a better understanding of the objects depicted in the image.

[0033] Techniques for image relighting using a neural network

[0034] Figure 1 Image 100 with imperfect lighting is shown in accordance with an example embodiment. Image 100 includes Image 110, Image 120, and Image 130. Image 110 is an image that includes a face with inconsistent lighting, where the left side of the face is illuminated more brightly than the right side of the face. Image 120 is an image that includes a face with a relatively large number of shadows (including several shadows that obscure the face). Image 130 is an image that includes a face photographed with a dark lighting that has a slightly green tint that can be considered "moody". Other examples of images with imperfect lighting and other types of imperfect lighting are also possible.

[0035] Figure 2 and Figure 3 shows an example image for training a convolutional neural network to change image illumination according to an example embodiment. Figure 2 The upper part of shows image 210, which shows a scene with several light sources. Figure 2 The lower part of shows images 220, 230, 240, 250, which are face images captured under different illumination conditions. Each of images 220, 230, 240, 250 is an image of the same person and is captured from the same perspective and distance, e.g., an image captured by a camera at the same distance and orientation as the person.

[0036] The image of the person can be captured when illuminated by each individual light source depicted in image 210. For example, image 220 can be captured when illuminated by the first light source depicted in image 210, image 230 can be captured when illuminated by the second light source depicted in image 210, image 240 can be captured when illuminated by the third light source depicted in image 210, and so on. Each of images 220, 230, and 240 can be multiplied or otherwise combined with the color of each corresponding light source. Figure 2 shows that the pixels of image 220 can be multiplied by data representing the color of the first light source, the pixels of image 230 can be multiplied by data representing the color of the second light source, and the pixels of image 240 can be multiplied by data representing the color of the third light source. Then, after each image is combined with the color of the corresponding light source, the resulting color images can be added or otherwise combined to obtain image 250, which represents an image of the person taken using all several light sources of the scene shown in image 210, as Figure 2 shown. Thus, image 250 appears to be taken using the illumination provided in the scene shown in image 210. In some examples, some or all of images 220, 230, 240, 250 can be provided as part of a one - light - at - a - time (OLAT) dataset.

[0037] Figure 3Indoor images 310, 320, 330 and outdoor images 330, 350, 360 representing indoor lighting conditions are shown, which can be used to train the convolutional neural network described herein. Images including images 310, 320, 330, 340, 350, 360 can be part of a training data set of images, where the training data set can include these images representing indoor and lighting conditions, as well as one or more OLAT data sets, other object images such as faces, and possibly other images. Then, the resulting training data set of images can be used to train the convolutional neural network described herein.

[0038] Figure 4 is a diagram depicting the training of a convolutional neural network 430 for changing image lighting according to an example embodiment. The convolutional neural network 430 can be a fully convolutional neural network as described herein. During training, the convolutional neural network 430 can receive as input one or more input training images and an input target lighting model. For example, Figure 4 shows a convolutional neural network 430 trained on an input original image 410 and an input target lighting model (TLM) 420. During training, the convolutional neural network 430 is guided to generate predictions of a target image 440 and a prediction of an original lighting model 450. The target image 440 can be a predicted (or generated) image produced by applying the input target lighting model 420 to the input original image 410. Thus, the target image 440 can be a prediction of how the original image 410 would look illuminated by the input target lighting model 420, rather than a prediction of the actual lighting condition that was used to illuminate the original image 410 when the original image 410 was initially captured. In this way, the target image 440 predicts how the original image 410 would be re-illuminated by the target lighting model 420.

[0039] The original lighting model 450 is a lighting model that predicts the actual lighting condition used to illuminate the original image 410. The lighting model can include a grid or other arrangement of lighting model data related to the lighting of part or all of one or more images. The lighting model data can include, but is not limited to, data representing one or more colors, intensities, albedos, directions, and / or surface normals of the lighting of part or all of one or more images. Thus, each of the lighting models 420, 450 can include lighting model data related to the lighting of at least a portion of the corresponding image. The target lighting model 420 can relate to the lighting of the target image 440, and the original lighting model 450 can relate to the lighting of the original image 410.

[0040] Figure 4Also shown are example dimensions of the image and illumination models - each of the original image and the target image is an image sized 512×512 pixels, and each of the target illumination model 420 and the original illumination model 450 is a 16×32 cell grid of illumination model data. Other dimensions for the original image, original illumination model, target image, and target illumination model are possible.

[0041] Figure 5 is a block diagram of a convolutional neural network 430 according to an example embodiment. The convolutional neural network 430 can receive the original image 410 and the target illumination model 420 as inputs, as shown at the upper left and lower right respectively. The convolutional neural network 430 can process the original image 410 and the target illumination model 420. The original image 410 can be processed to determine a prediction of the original illumination model (OLM) 450 and provide an input to determine a prediction of the target image 440. The target illumination model 420 can be used to determine a prediction of the target image 440.

[0042] The convolutional neural network 430 can include node layers for processing the original image 410 and the target illumination model 420. Example layers can include, but are not limited to, an input layer, a convolutional layer, an activation layer, a pooling layer, a fully connected layer, and an output layer. The input layer can store input data, such as pixel data of the original image 410 and inputs from other layers of the convolutional neural network 430. The convolutional layer can compute the output of neurons connected to local regions in the input. In some cases, the convolutional layer can act as a transposed convolutional or deconvolution layer to apply a filter to a relatively small input, thereby expanding / upsampling the relatively small input into a larger output. The activation layer can determine whether the output of the previous layer is "activated" or actually provided (e.g., provided to a subsequent layer). The pooling layer can downsample the input. For example, the convolutional neural network 430 can include one or more pooling layers to downsample the input by a predetermined factor (e.g., factor 2) in the horizontal and / or vertical dimensions. The fully connected layer can determine scores related to the prediction. Such scores can include, but are not limited to, scores related to the predicted original illumination model and / or scores related to the predicted target image. The output layer can provide the output of the convolutional neural network 430 to software and / or hardware interfacing with the convolutional neural network 430; for example, hardware and / or software for displaying, printing, communicating, and / or otherwise providing the target image 440. Other layers, such as batch normalization layers, can also be in the convolutional neural network 430. Layers 510, 512, 514, 516, 520, 522, 524, 530, 532, 534, 540, 542, 544, 546 can include one or more input layers, output layers, convolutional layers, activation layers, pooling layers, fully connected layers, and / or other layers described herein.

[0043] In Figure 5 and Figure 6 , gray blocks are used to show the layers of the convolutional neural network 430 involved in processing the original image 410 and the target illumination model 420 to determine the prediction of the target image 440. In addition, yellow blocks are used to show the information of the convolutional neural network 430 used to determine the prediction of the original illumination model and in processing the target illumination model 420. More specifically, the convolutional neural network 430 may include original layers 510, 512, 514, 516, which are arranged in the order of layers L1, L2, L3, L4 respectively. Each layer continuously convolves its input and provides its output to the consecutive layer until reaching the original layer L4 516.

[0044] The output layer L4 516 may be associated with the first original illumination model (OLM) information layer 520, and the first original illumination model information layer 520 may provide an output to the second original illumination model information layer 522, and the second original illumination model information layer 522 may in turn provide an output to the third original illumination model information layer 524. The third original illumination model information layer 524 may include an output layer to provide the predicted original illumination model 450. For example, the original illumination model information layers 520, 522, 524 may include one or more fully connected layers for predicting the original illumination model 450. In some examples, the original illumination model information layer 520 may receive illumination-related features of the original image 410 determined by the original layers 510, 512, 514, 516. For example, the original layer 516 may output or otherwise provide illumination-related features to the original illumination model information layer 520.

[0045] The first target illumination model (TLM) information layer 530 may act as an input layer to receive the target illumination model 420 as an input. The target illumination model information layer 530 may provide an output to the second target illumination model information layer 532, and the second target illumination model information layer 532 may in turn provide an output to the third target illumination model information layer 534. The target illumination model information layer 534 may include an output layer to provide illumination features related to the target illumination model 420. For example, the target illumination model information layers 530, 532, 534 may include fully connected layers for predicting the original illumination model 450. In some examples, the original illumination model information layer 520 may receive illumination-related features of the original image 410 determined by the original layers 510, 512, 514, 516. For example, the original layer 516 may output or otherwise provide illumination-related features to the original illumination model information layer 520.

[0046] In Figure 5 and Figure 6In [the figure], the depicted layer(s) may include one or more actual layers. For example, the original layer L1 510 may have one or more input layers, one or more activation layers, and / or one or more additional layers. As another example, the original layer L2 512, the original layer L3 514, and / or the original layer L4 516 may include one or more convolutional layers, one or more activation layers (e.g., having a one-to-one relationship with one or more convolutional layers), one or more pooling layers, and / or one or more additional layers.

[0047] In some examples, some or all of the pooling layers in the convolutional neural network 430 may downsample the input by a common factor in both the horizontal and vertical dimensions without downsampling the depth dimension associated with the input. The depth dimension may store data of pixel colors (red, green, blue) and / or data representing scores. For example, assume the size of the original image 410 is 512×512 and the depth is D, each of the original layers 510, 512, 514, 516 includes a pooling layer, and each pooling layer in the original layers 510, 512, 514, 516 downsamples the original image 410 by a factor of 2 in both the horizontal and vertical dimensions. In this case, the output size of the original layer 510 will be 256×256×D, the output size of the original layer 512 will be 128×128×D, the output size of the original layer 514 will be 64×64×D, and the output size of the original layer 516 will be 32×32×D. One or more (pooling) layers of the convolutional neural network 430 may also use other common factors for downsampling other than two (two).

[0048] The original layer L1 510 may receive and process the original image 410 and provide an output to the original layer L2 512. The original layer L2 512 may process the output of the original layer L1 and provide an output to the original layer L3 514. The original layer L3 514 may process the output of the original layer L2 and provide an output to the original layer L4 516. The original layer L4 516 may process the output of the original layer L3. At least a part of the output of the original layer L4 516 may be provided as an input to the original illumination model information layer 520.

[0049] The convolutional neural network 430 may use the original illumination model information layers 520, 522, and 524 to predict the original illumination model 450. The original illumination model 450 may be output by the illumination model information layer 524 of the convolutional neural network 430. The convolutional neural network 430 may use confidence learning to train the original illumination model information layers 520, 522, and 524 to predict the original illumination model 450. Confidence learning is discussed in more detail at least in Figure 6 the context of [here].

[0050] To predict the target image 440, the convolutional neural network 430 can use the target illumination model information layers 530, 532, and 534 to process the target illumination model 420. The output of the target illumination model information layer 534 can be provided as an input to the target layer L1 540 together with the data provided by copying (e.g., using a skip connection between the original layer L4 516 and the target layer L1 540) the data from the original layer L4 516 to start predicting the target image 440. The target layer L2 542 can receive and process inputs from the target layer L1 540 and the original layer L3 514 (e.g., using a skip connection between the original layer L3 514 and the target layer L2 542) to provide an output to the target layer L3 544. The target layer L3 544 can receive and process inputs from the target layer L2 542 and the original layer L2 512 (e.g., using a skip connection between the original layer L2 512 and the target layer L3 544) to provide an output to the target layer L4 546. The target layer L4 546 can receive and process inputs from the target layer L3 544 and the original layer L1 510 to provide a prediction of the target image 440, which can then be output from the target layer L4 546. The data provided by the skip connections between the original layers 516, 514, 512, 510 and the corresponding target layers 540, 542, 544, 546 can be used by each corresponding target layer to provide additional details for generating the contribution of the target layer to the prediction of the target image 440. In some examples, each of the target layers 540, 542, 544, 546 used to predict the target image 440 can include one or more convolutional layers (possibly performing transposed convolution / deconvolution), one or more activation layers, and possibly one or more input and / or output layers. In some examples, some or all of the layers 510, 512, 514, 516, 520, 522, 524, 530, 532, 534, 540, 542, 544, 546 can act as a convolutional encoder / decoder network.

[0051] A loss measure can be used during the training of the convolutional neural network 430. For example, during the training of the convolutional neural network 430 for predicting the target image 440, the L2 loss measure between the target image prediction and the training image can be minimized. As another example, during the training of the convolutional neural network 430 for predicting the original illumination model 450, the log L1 loss measure between the original illumination model prediction and the training illumination model data can be minimized. Other loss measures can also be used, or alternatively other loss measures can be used.

[0052] In some examples, the convolutional neural network 430 may include perceptual loss processing. For example, the convolutional neural network 430 may use a generative adversarial net (GAN) loss function to determine whether a part or all of an image will be predicted to be actually illuminated by a specific illumination model, so as to meet one or more perception-related conditions of the illumination of this part of the image. In some examples, cycle loss may be used to feed the predicted target image and / or the original illumination model back into the convolutional neural network 430 to generate and / or further refine the predicted target image and / or the original illumination model. In some examples, the convolutional neural network 430 may utilize deep supervision techniques to provide constraints on intermediate layers. In some examples, the convolutional neural network 430 may have more, fewer, and / or different layers than Figure 5 shown.

[0053] Figure 6 FIG. is a block diagram showing confidence learning 630 used by the convolutional neural network 430 during training according to an example embodiment. As described above, the original illumination model information layers 520, 522, and 524 may be used to predict the original illumination model 450. Confidence learning may be used during the training of one or more of the original illumination model information layers 520, 522, and 524.

[0054] To determine the original illumination model 450, the original illumination model information layers 520, 522, and 524 may determine predictions regarding illumination features, such as predictions regarding light direction. Figure 6 The upper part of FIG. shows the position "patch 1" of the original image 410, and the possible predicted normal direction 620 of the ray of light falling on the face at patch 1 is shown using a red arrow. The possible predicted normal direction 620 faces outward from the depicted face at patch 1 because the light falling on patch 1 may be emitted from outside the depicted face, and thus has a normal incidence from the face outward on the depicted face. In contrast, Figure 6 the non-predicted normal direction 622 depicted using a blue arrow in FIG. would be associated with light from inside the face depicted at patch 1, which is less likely. After the training of the convolutional neural network 430, light from the non-predicted normal direction 622 should not be predicted. However, during the training of the original illumination model information layers 520, 522, and / or 524, an illumination model indicating light from the non-predicted normal direction 622 may be predicted, where the confidence that the light may come from the non-predicted normal direction 622 decreases as the training progresses.

[0055] Confidence learning 630 can be used to apply such confidence information during the training of convolutional neural network 430. For example, during the training of the original illumination model information layers 520, 522, and / or 524, the convolutional neural network 430 (e.g., the original layer L4 516 and / or the original illumination model information layers 520, 522, 524) can determine the original illumination information 610. The original illumination information 610 can include illumination model predictions at a patch or portion of the image; for example, the light prediction 640 regarding the normal direction and / or other attributes of the light falling on patch 1 of the face depicted in the original image 410. Additionally, the convolutional neural network 430 (e.g., the original layer L4 516 and / or the original illumination model information layers 520, 522, 524) can determine a confidence value 650 associated with the prediction 640 of the illumination model at patch 1.

[0056] The convolutional neural network 430 can be used to explicitly predict the confidence value 650 or implicitly provide the confidence value 650 based on some or all of the weights of the convolutional neural network 430. In some examples, one or more of the original layer L4 516 and / or the original illumination model information layers 520, 522, 524 can be used to explicitly predict the confidence value 650 or implicitly provide the confidence value 650 based on some or all of the weights of one or more of the original layer L4 516 and / or the original illumination model information layers 520, 522, 524. More specifically, one or more weights of one or more of the original layer L4 516 and / or the original illumination model information layers 520, 522, 524 can be used as the confidence value 650. Then, confidence learning 630 can involve multiplying the light prediction 640 by the predicted confidence value 650 and / or otherwise mathematically combining them to determine an updated light prediction 660. Compared to using the light prediction 640 during training, using the updated light prediction 660 generated by confidence learning 630 can result in emphasizing relatively-confident predictions over relatively-non-confident predictions during the training of the convolutional neural network 430, thus providing additional use and feedback on the confidence of the illumination prediction. Other examples and / or uses of confidence learning are also possible.

[0057] Figures 7 - 18 An example image of a human face generated by the convolutional neural network 430 according to an example embodiment is shown. In particular, Figures 7 - 9 An example image of a human face related to the illumination model prediction made by the convolutional neural network 430 is shown, Figures 10 - 18 An example image of a human face related to the target model prediction made by the convolutional neural network 430 is shown. Generally, Figures 7 - 18It is shown that the convolutional neural network 430 can generate accurate predictions of the illumination model and relit face images in a wide range of lighting environments.

[0058] Figure 7 Image 700 is shown, which is captured when illuminated by light modeled by the "Groundtruth original light" indicated at the lower left corner of Image 700. The convolutional neural network 430 predicts the illumination model shown by the "Predicted original light" at the lower right corner of Image 700. Figure 7 Both the groundtruth original light and the predicted original light depicted in Figure 7 are based on the environment map shown on the left. The environment map indicates that the upper parts of both the groundtruth original light and the predicted original light are related to the light from the rear of the environment depicted in Image 700, and the lower parts of both the groundtruth original light and the predicted original light are related to the light from the front of the environment depicted in Image 700. The environment map also indicates that the left parts of both the groundtruth original light and the predicted original light are related to the light on the left side of the face depicted in Image 700, which corresponds to the right side of Image 700, and also indicates that the right parts of both the groundtruth original light and the predicted original light are related to the light on the right side of the face depicted in Image 700, which corresponds to the left side of Image 700. In Figure 7 depicted and used for Figure 7 the same environment map is also depicted in the illumination model shown in Figures 8 - 18 and used for the illumination model shown in Figures 8 - 18 shown in

[0059] Figure 7 It is shown that both the groundtruth original light and the predicted original light of Image 700 have bright parts in the upper left, which indicates that most of the light in the environment of Image 700 is predicted to be and actually comes from behind the face depicted in Image 700 and falls on the left side of the face, which is shown on the right side of Image 700. Image 700 confirms that the prediction made by the predicted original light is that Image 700 is illuminated more brightly on the right side than on the left side.

[0060] Figure 8 The lower left part of Image 800 shows the groundtruth original light of the image, while the lower right part of Image 800 shows the predicted original light of the image, where the predicted original light of Image 800 is generated by the convolutional neural network 430. Both the groundtruth original light and the predicted original light of Image 800 have bright parts in the upper right, which indicates that most of the light in the environment of Image 800 is predicted to be and actually comes from behind the face depicted in Image 800 and falls on the right side of the face, which is shown on the left side of Image 800. Image 800 confirms that the prediction made by the predicted original light is that Image 800 is illuminated more brightly on the left side than on the right side.

[0061] Figure 9 The lower left portion of Image 900 shows the true original light of the image, while the lower right portion of Image 900 shows the predicted original light of the image, where the predicted original light of Image 900 is generated by the convolutional neural network 430. Both the true original light and the predicted original light of Image 900 have a relatively large bright portion in the upper right and a relatively small bright portion in the upper left, indicating that most of the light in the environment of Image 900 is predicted to be and actually comes from two light sources: a larger light source behind the face depicted in Image 900, and the light from this source falls on the right side of the face shown on the left side of Image 900; and a smaller light source also behind the face depicted in Image 900, and the light from this source falls on the left side of the face shown on the right side of Image 900. Image 900 confirms that the prediction made by the predicted original light is that Image 900 is illuminated more brightly on the left than on the right, like Image 800, but is illuminated more evenly across the entire face than Image 800.

[0062] Figures 10 - 18 Each of which shows a set of three images: the original image; the true target image with the corresponding target illumination model; and the predicted target image generated by applying the target illumination model to the original image by the training version of the convolutional neural network 430. For example, Figure 10 shows the original image 1010, the true target image 1020, and the predicted target image 1030, where the original image 1010 is shown with an environmental map of the original illumination model, and the environmental map indicates that the image 1010 is backlit with relatively uniform light. Both the true target image 1020 and the predicted target image 1030 are shown with an environmental map of the target illumination model showing three light sources, where one light source is more forward in the environment than the light used for the original image 1010. Both the true target image 1020 and the predicted target image 1030 show a similar illumination reflected by the target illumination model shown in each corresponding image.

[0063] Figure 11 shows the original image 1110, the true target image 1120, and the predicted target image 1130, where the original image 1110 is shown with an environmental map of the original illumination model, and the environmental map indicates that the image 1110 is backlit with a relatively small light source near the face depicted in the image 1110 with relatively dim light. Both the true target image 1120 and the predicted target image 1130 are shown with an environmental map of the target illumination model showing relatively bright backlighting compared to the original image 1110. Both the true target image 1120 and the predicted target image 1130 show a similar illumination reflecting the target illumination model shown in each corresponding image.

[0064] Figure 12Shows the original image 1210, the ground truth target image 1220, and the predicted target image 1230. Among them, the original image 1210 is shown with an environmental map of the original illumination model, and the environmental map indicates that the image 1210 is backlit with stronger light on the left side (shown on the right side of the image 1210) of the depicted face than on the right side (shown on the left side of the image 1210) of the depicted face. Both the ground truth target image 1220 and the predicted target image 1230 are shown with an environmental map of the target illumination model, and the environmental map shows a relatively large light source that dominates the illumination environment. Both the ground truth target image 1220 and the predicted target image 1230 show similar illumination that reflects the target illumination model shown in each corresponding image. However, the predicted target image 1230 is slightly darker than the target image 1220, which may reflect the relatively dark illumination of the input original image 1210.

[0065] Figure 13 Shows the original image 1310, the ground truth target image 1320, and the predicted target image 1330. Among them, the original image 1310 is shown with an environmental map of the original illumination model, and the environmental map indicates that the image 1310 is backlit with relatively uniform light. Both the ground truth target image 1320 and the predicted target image 1330 are shown with an environmental map of the target illumination model that shows two light sources, where the larger light source is on the right side of the face depicted in the image 1320. Both the ground truth target image 1320 and the predicted target image 1330 show similar illumination that reflects the target illumination model shown in each corresponding image.

[0066] Figure 14 Shows the original image 1410, the ground truth target image 1420, and the predicted target image 1430. Among them, the original image 1410 is shown with an environmental map of the original illumination model, and the environmental map indicates that the image 1410 is backlit with a relatively large white light source. Both the ground truth target image 1420 and the predicted target image 1430 are shown with an environmental map of the target illumination model that shows three light sources: a relatively large white light source on the left side of the face depicted in the image 1420, a relatively small white light source on the right side of the face depicted in the image 1420, and a relatively large yellow light source in the center of the illumination environment of the image 1420. Both the ground truth target image 1420 and the predicted target image 1430 show similar illumination that reflects the target illumination model shown in each corresponding image.

[0067] Figure 15Shows the original image 1510, the ground truth target image 1520, and the predicted target image 1530. Among them, the original image 1510 is shown with an environmental map of the original illumination model, and the environmental map indicates that the image 1510 is backlit with relatively uniform light approaching the left and right sides of the face depicted in the image 1510. Both the ground truth target image 1520 and the predicted target image 1530 are shown with an environmental map of the target illumination model, and the environmental map shows a light source on the left side of the face depicted in the image 1520. Both the ground truth target image 1520 and the predicted target image 1530 show a similar illumination reflecting the target illumination model shown in each corresponding image. However, the illumination on the face depicted in the image 1530 is darker than the illumination on the face depicted in the image 1520.

[0068] Figure 16 Shows the original image 1610, the ground truth target image 1620, and the predicted target image 1630. Among them, the original image 1610 is shown with an environmental map of the original illumination model, and the environmental map indicates that the image 1610 is mainly illuminated by a light source on the left side of the face depicted in the image 1610. Both the ground truth target image 1620 and the predicted target image 1630 are shown with an environmental map of the target illumination model dominated by a relatively bright light source on the right side of the face depicted in the image 1620. Both the ground truth target image 1620 and the predicted target image 1630 show a similar illumination reflecting the target illumination model shown in each corresponding image.

[0069] Figure 17 Shows the original image 1710, the ground truth target image 1720, and the predicted target image 1730. Among them, the original image 1710 is shown with an environmental map of the original illumination model, and the environmental map indicates that the image 1710 is mainly illuminated by a light source on the right side of the face depicted in the image 1710. Both the ground truth target image 1720 and the predicted target image 1730 are shown with an environmental map of the target illumination model showing three light sources, where two light sources are relatively large and used to backlight the face depicted in the image 1720 from the left and right sides, and the third relatively small light source is on the left side of the face depicted in the image 1720. Both the ground truth target image 1720 and the predicted target image 1730 show a similar illumination reflecting the target illumination model shown in each corresponding image.

[0070] Figure 18Shows the original image 1810, the ground truth target image 1820, and the predicted target image 1830, where the original image 1810 is shown with an environment map of the original lighting model, and the environment map indicates that the image 1810 is mainly illuminated by a single white light source relatively close to the face depicted in the image 1810. Both the ground truth target image 1820 and the predicted target image 1830 are shown with an environment map of the target lighting model showing three yellow light sources. Two of the light sources are relatively large and mainly backlight the face depicted in the image 1820 from the left side. The other light source is relatively small and is located on the right side of the face depicted in the image 1820 and close to the face. Both the ground truth target image 1820 and the predicted target image 1830 show a similar yellowish tint of illumination, which is reflected in the target lighting model shown in each respective image.

[0071] Training a machine learning model for generating inferences / predictions

[0072] Figure 19 FIG. 1900 shows a diagram illustrating a training phase 1902 and an inference phase 1904 of a trained machine learning model 1932 according to an example embodiment. Some machine learning techniques involve training one or more machine learning algorithms on an input set of training data to identify patterns in the training data and provide output inferences and / or predictions regarding the patterns in the training data. The resulting trained machine learning algorithms may be referred to as trained machine learning models. For example, Figure 19 FIG. 1902 shows the training phase, in which one or more machine learning algorithms 1920 are being trained on training data 1910 to become a trained machine learning model 1932. Then, during the inference phase 1904, the trained machine learning model 1932 can receive input data 1930 and one or more inference / prediction requests 1940 (possibly as part of the input data 1930) and provide, as a response, one or more inferences and / or predictions 1950 as output.

[0073] Thus, the trained machine learning model 1932 can include one or more models of one or more machine learning algorithms 1920. The machine learning algorithms 1920 can include, but are not limited to: artificial neural networks (e.g., convolutional neural networks, recurrent neural networks using the confidence learning techniques described herein), Bayesian networks, hidden Markov models, Markov decision processes, logistic regression functions, support vector machines, suitable statistical machine learning algorithms, and / or heuristic machine learning systems. The machine learning algorithms 1920 can be supervised or unsupervised and can implement any suitable combination of online and offline learning.

[0074] In some examples, the machine learning algorithm 1920 and / or the trained machine learning model 1932 can be accelerated using an on-device co-processor (e.g., a graphics processing unit (GPU), a tensor processing unit (TPU), a digital signal processor (DSP), and / or an application specific integrated circuit (ASIC)). Such an on-device co-processor can be used to accelerate the machine learning algorithm 1920 and / or the trained machine learning model 1932. In some examples, the trained machine learning model 1932 can be trained, reside, and execute to provide inference on a particular computing device, and / or can make inferences for a particular computing device.

[0075] During the training phase 1902, the machine learning algorithm 1920 can be trained by providing at least training data 1910 as training input using unsupervised, supervised, semi-supervised, and / or reinforcement learning techniques. Unsupervised learning involves providing a portion (or all) of the training data 1910 to the machine learning algorithm 1920, and the machine learning algorithm 1920 determining one or more output inferences based on the portion (or all) of the provided training data 1910. Supervised learning involves providing a portion of the training data 1910 to the machine learning algorithm 1920, the machine learning algorithm 1920 determining one or more output inferences based on the portion of the provided training data 1910, and accepting or correcting the output inferences based on the correct results associated with the training data 1910. In some examples, the supervised learning of the machine learning algorithm 1920 can be governed by a set of rules and / or a set of labels for the training input, and the set of rules and / or the set of labels can be used to correct the inferences of the machine learning algorithm 1920.

[0076] Semi-supervised learning involves having correct results for a portion rather than all of the training data 1910. During semi-supervised learning, supervised learning is used for the portion of the training data 1910 that has correct results, and unsupervised learning is used for the portion of the training data 1910 that does not have correct results. Reinforcement learning involves a machine learning algorithm 1920 receiving a reward signal regarding a previous inference, where the reward signal can be a numerical value. During reinforcement learning, the machine learning algorithm 1920 can output an inference and receive a reward signal as a response, where the machine learning algorithm 1920 is configured to attempt to maximize the numerical value of the reward signal. In some examples, reinforcement learning also utilizes a value function that provides a numerical value representing the expected sum of the numerical values provided by the reward signal over time. In some examples, other machine learning techniques can be used to train the machine learning algorithm 1920 and / or the trained machine learning model 1932, including but not limited to, incremental learning and curriculum learning.

[0077] In some examples, the machine learning algorithm 1920 and / or the trained machine learning model 1932 can use transfer learning techniques. For example, transfer learning techniques can involve a trained machine learning model 1932 being pre-trained on a set of data and being additionally trained using the training data 1910. More specifically, the machine learning algorithm 1920 can be pre-trained on data from one or more computing devices, and the resulting trained machine learning model is provided to a computing device CD1, where CD1 is intended to execute the trained machine learning model during the inference phase 1904. Then, during the training phase 1902, the trained machine learning model can be additionally trained using the training data 1910, where the training data 1910 can be derived from the kernel and non-kernel data of the computing device CD1. This further training of the machine learning algorithm 1920 and / or the pre-trained trained machine learning model using the training data 1910 from CD1's data can be performed using supervised or unsupervised learning. Once the machine learning algorithm 1920 and / or the pre-trained machine learning model has been trained at least on the training data 1910, the training phase 1902 can be completed. The resulting trained machine learning model can be used as at least one of the trained machine learning models 1932.

[0078] In particular, once the training phase 1902 has been completed, the trained machine learning model 1932 can be provided to the computing device, if not already on the computing device. The inference phase 1904 can begin after the trained machine learning model 1932 has been provided to the computing device CD1.

[0079] During the inference phase 1904, the trained machine learning model 1932 can receive input data 1930 and generate and output one or more corresponding inferences and / or predictions 1950 regarding the input data 1930. In this way, the input data 1930 can be used as input to the trained machine learning model 1932 for providing corresponding inferences and / or predictions 1950 to kernel components and non-kernel components. For example, the trained machine learning model 1932 can generate inferences and / or predictions 1950 in response to one or more inference / prediction requests 1940. In some examples, the trained machine learning model 1932 can be executed as part of other software. For example, the trained machine learning model 1932 can be executed by an inference or prediction daemon to be readily available to provide inferences and / or predictions upon request. The input data 1930 can include data from the computing device CD1 on which the trained machine learning model 1932 is executed and / or input data from one or more computing devices other than CD1.

[0080] The input data 1930 can include a collection of images provided by one or more sources. The collection of images can include images of objects (e.g., faces, where the images of the faces are taken under different lighting conditions), images of multiple objects, images resident on the computing device CD1, and / or other images. Other types of input data are also possible.

[0081] The inferences and / or predictions 1950 can include output images, output lighting models, numerical values, and / or other output data generated by the trained machine learning model 1932 operating on the input data 1930 (and the training data 1910). In some examples, the trained machine learning model 1932 can use the output inferences and / or predictions 1950 as input feedback 1960. The trained machine learning model 1932 can also rely on past inferences as input for generating new inferences.

[0082] The convolutional neural network 430 can be an example of the machine learning algorithm 1920. After training, the trained version of the convolutional neural network 430 can be an example of the trained machine learning model 1932. In this method, an example of the inference / prediction request 1940 can be a request to apply a specific lighting model to an input image of an object, and a corresponding example of the inferences and / or predictions 1950 can be a prediction of the output image when the specific lighting model is applied to the input image.

[0083] In some examples, a computing device CD_SOLO can include a trained version of a convolutional neural network 430, perhaps after training the convolutional neural network 430. Then, the computing device CD_SOLO can receive a request to apply a particular lighting model to a corresponding input image and use the trained version of the convolutional neural network 430 to generate an output image that predicts the application of the particular lighting model to the input image. In some of these examples, the request received by the CD_SOLO to predict the output image of the application of the particular lighting model to the input image can include a request for the original lighting model or be replaced by a request for the original lighting model, each of which can model the lighting that illuminates the corresponding input image. Then, the CD_SOLO can use the trained version of the convolutional neural network 430 to generate the output image and / or the original lighting model according to the request.

[0084] In some examples, two or more computing devices CD_CLI and CD_SRV can be used to provide the output image; for example, a first computing device CD_CLI can generate and send a request to apply a particular lighting model to a corresponding input image to a second computing device CD_SRV. Then, the CD_SRV can use the trained version of the convolutional neural network 430, perhaps after training the convolutional neural network 430, to generate an output image that predicts the application of the particular lighting model to the input image and respond to the request for the output image from the CD_CLI. Then, upon receiving the response to the request, the CD_CLI can provide the requested output image (e.g., using a user interface and / or display, printed copy, electronic communication, etc.). In some examples, the request for the output image that predicts the application of the particular lighting model to the input image can include a request for the original lighting model or be replaced by a request for the original lighting model, each of which can model the lighting that illuminates the corresponding input image. Then, the CD_SRV can use the trained version of the convolutional neural network 430 to generate the output image and / or the original lighting model according to the request. Other examples for generating the output image that predicts the application of the particular lighting model to the input image and / or for generating the original lighting model using the trained version of the convolutional neural network 430 are also possible.

[0085] Example data network

[0086] Figure 20Depicts a distributed computing architecture 2000 according to an example embodiment. The distributed computing architecture 2000 includes server devices 2008, 2010 that are configured to communicate with programmable devices 2004a, 2004b, 2004c, 2004d, 2004e via a network 2006. The network 2006 can correspond to a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a wireless wide area network (WWAN), a corporate intranet, the public Internet, or any other type of network configured to provide a communication path between networked computing devices. The network 2006 can also correspond to a combination of one or more LANs, WANs, corporate intranets, and / or the public Internet.

[0087] Although Figure 20 only five programmable devices are shown, the distributed application architecture can serve dozens, hundreds, or thousands of programmable devices. Additionally, the programmable devices 2004a, 2004b, 2004c, 2004d, 2004e (or any additional programmable devices) can be any kind of computing device, such as a mobile computing device, a desktop computer, a wearable computing device, a head-mounted device (HMD), a network terminal, a mobile computing device, etc. In some examples, as shown by programmable devices 2004a, 2004b, 2004c, 2004e, the programmable devices can be directly connected to the network 2006. In other examples, as shown by programmable device 2004d, the programmable device can be indirectly connected to the network 2006 via an associated computing device (such as programmable device 2004c). In this example, programmable device 2004c can act as the associated computing device to relay electronic communications between programmable device 2004d and the network 2006. In other examples, as shown by programmable device 2004e, the computing device can be part of and / or within a vehicle, such as a car, a truck, a bus, a ship or vessel, an airplane, etc. In Figure 20 other examples not shown, the programmable devices can be directly and indirectly connected to the network 2006.

[0088] As requested by the programmable devices 2004a - 2004e, the server devices 2008, 2010 can be configured to perform one or more services. For example, the server devices 2008 and / or 2010 can provide content to the programmable devices 2004a - 2004e. The content can include, but is not limited to, web pages, hypertext, scripts, binary data such as compiled software, images, audio, and / or video. The content can include compressed and / or uncompressed content. The content can be encrypted and / or unencrypted. Other types of content are also possible.

[0089] As another example, server devices 2008 and / or 2010 can provide access to software for databases, search, computing, graphics, audio, video, World Wide Web / Internet utilization, and / or other functions to programmable devices 2004a - 2004e. Many other examples of server devices are also possible.

[0090] Computing Device Architecture

[0091] Figure 21 is a block diagram of an example computing device 2100 according to an example embodiment. Specifically, Figure 21 the computing device 2100 shown in can be configured to perform at least one function of and / or functions related to a convolutional neural network, confidence learning, predicting a target image, predicting an original illumination model, convolutional neural network 430, confidence learning 630, and / or method 2300.

[0092] The computing device 2100 can include a user interface module 2101, a network communication module 2102, one or more processors 2103, a data store 2104, one or more cameras 2118, one or more sensors 2120, and a power system 2122, all of which can be linked together via a system bus, network, or other connection mechanism 2105.

[0093] The user interface module 2101 is operable to send data to and / or receive data from external user input / output devices. For example, the user interface module 2101 can be configured to send data to and / or receive data from user input devices such as a touch screen, computer mouse, keyboard, keypad, touchpad, trackball, joystick, voice recognition module, and / or other similar devices. The user interface module 2101 can also be configured to provide output to a user display device such as one or more cathode ray tubes (CRTs), liquid crystal displays, light emitting diodes (LEDs), displays using digital light processing (DLP) technology, printers, light bulbs, and / or other similar devices now known or later developed. The user interface module 2101 can also be configured to generate an audible output using devices such as speakers, speaker jacks, audio output ports, audio output devices, headphones, and / or other similar devices. The user interface module 2101 can also be configured with one or more haptic devices that can generate a haptic output such as vibration and / or other outputs detectable by touch and / or physical contact with the computing device 2100. In some examples, the user interface module 2101 can be used to provide a graphical user interface (GUI) for utilizing the computing device 2100.

[0094] The network communication module 2102 may include one or more devices that provide one or more wireless interfaces 2107 and / or one or more wireline interfaces 2108, which may be configured to communicate via a network. The wireless interface 2107 may include one or more wireless transmitters, receivers, and / or transceivers, such as Bluetooth TM transceivers, transceivers, Wi-Fi TM transceivers, WiMAX TM transceivers, and / or other similar types of wireless transceivers that may be configured to communicate via a wireless network. The wireline interface 2108 may include one or more wireline transmitters, receivers, and / or transceivers, such as Ethernet transceivers, universal serial bus (USB) transceivers, or similar transceivers, which may be configured to communicate with a wireline network via a twisted pair, coaxial cable, fiber optic link, or similar physical connection.

[0095] In some examples, the network communication module 2102 may be configured to provide reliable, secure, and / or authenticated communication. For each communication described herein, information may be provided to facilitate reliable communication (e.g., guaranteed message delivery), possibly as part of a message header and / or footer (e.g., packet / message sequencing information, encapsulation headers and / or footers, size / time information, and transmission verification information such as cyclic redundancy check (CRC) and / or parity values). One or more encryption protocols and / or algorithms may be used to secure (e.g., encode or encrypt) and / or decrypt / decode the communication, such as, but not limited to, Data Encryption Standard (DES), Advanced Encryption Standard (AES), Rivest-Shamir-Adelman (RSA) algorithm, Diffie-Hellman algorithm, secure socket protocols such as Secure Sockets Layer (SSL) or Transport Layer Security (TLS), and / or Digital Signature Algorithm (DSA). Other encryption protocols and / or algorithms may also be used, or additional to those listed herein, to protect (and then decrypt / decode) the communication.

[0096] The one or more processors 2103 may include one or more general-purpose processors, and / or one or more dedicated processors (e.g., digital signal processors, tensor processing units (TPU), graphics processing units (GPU), application specific integrated circuits, etc.). The one or more processors 2103 may be configured to execute computer-readable instructions 2106 contained in the data storage 2104 and / or other instructions described herein.

[0097] Data storage 2104 may include one or more non-transitory computer-readable storage media that may be read and / or accessed by at least one of one or more processors 2103. The one or more computer-readable storage media may include volatile and / or non-volatile storage components, such as optical, magnetic, organic, or other memory or disk storage, which may be integrated, in whole or in part, with at least one of one or more processors 2103. In some examples, data storage 2104 may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage unit), while in other examples, data storage 2104 may be implemented using two or more physical devices.

[0098] Data storage 2104 may include computer-readable instructions 2106 and possibly additional data. In some examples, data storage 2104 may include storage required to execute at least a portion of the methods, scenarios, and techniques described herein and / or at least a portion of the functionality of the devices and networks described herein. In some examples, data storage 2104 may include storage of a trained neural network model 2112 (e.g., a model of a trained convolutional neural network such as convolutional neural network 430). In a particular example of these examples, computer-readable instructions 2106 may include instructions that, when executed by processor 2103, enable computing device 2100 to provide some or all of the functionality of trained neural network model 2112.

[0099] In some examples, computing device 2100 may include one or more cameras 2118. Cameras 2118 may include one or more image capture devices, such as still and / or video cameras, that are configured to capture light and record the captured light in one or more images; that is, cameras 2118 may generate images of the captured light. The one or more images may be one or more still images and / or one or more images used in video images. Cameras 2118 may capture light and / or electromagnetic radiation emitted as visible light, infrared radiation, ultraviolet light, and / or one or more other frequencies of light.

[0100] In some examples, computing device 2100 may include one or more sensors 2120. The sensors 2120 may be configured to measure conditions within the computing device 2100 and / or conditions in the environment of the computing device 2100 and provide data regarding such conditions. For example, the sensors 2120 may include one or more of the following: (i) sensors for obtaining data regarding the computing device 2100, such as, but not limited to, a thermometer for measuring the temperature of the computing device 2100, a battery sensor for measuring the power of one or more batteries of the power system 2122, and / or other sensors for measuring the conditions of the computing device 2100; (ii) identification sensors for identifying other objects and / or devices, such as, but not limited to, a radio frequency identification (RFID) reader, a proximity sensor, a one-dimensional barcode reader, a two-dimensional barcode (e.g., quick response (QR) code) reader, and a laser tracker, where the identification sensors may be configured to read identifiers such as RFID tags, barcodes, QR codes, and / or other devices and / or objects configured to be read and provide at least identification information; (iii) sensors for measuring the position and / or movement of the computing device 2100, such as, but not limited to, an inclinometer, a gyroscope, an accelerometer, a Doppler sensor, a GPS device, a sonar sensor, a radar device, a laser displacement sensor, and a compass; (iv) environmental sensors for obtaining data indicative of the environment of the computing device 2100, such as, but not limited to, an infrared sensor, an optical sensor, a light sensor, a biosensor, a capacitance sensor, a touch sensor, a temperature sensor, a wireless sensor, a radio sensor, a motion sensor, a microphone, a sound sensor, an ultrasonic sensor, and / or a smoke sensor; and / or (v) force sensors for measuring one or more forces acting on the computing device 2100 (e.g., inertial forces and / or gravity), such as, but not limited to, one or more of the following sensors that measure: force in one or more dimensions, torque, ground force, friction, and / or a zero moment point (ZMP) sensor for identifying and / or the position of the ZMP. Many other examples of the sensors 2120 are possible.

[0101] The power system 2122 may include one or more batteries 2124 and / or one or more external power interfaces 2126 for providing power to the computing device 2100. When electrically coupled to the computing device 2100, each of the one or more batteries 2124 can act as a source of stored power for the computing device 2100. One or more batteries 2124 of the power system 2122 may be configured to be portable. Some or all of the one or more batteries 2124 may be easily removable from the computing device 2100. In other examples, some or all of the one or more batteries 2124 may be internal to the computing device 2100 and thus may not be easily removable from the computing device 2100. Some or all of the one or more batteries 2124 may be rechargeable. For example, a rechargeable battery may be recharged via a wired connection between the battery and another power source, such as one or more power sources external to the computing device 2100 and connected to the computing device 2100 via one or more external power interfaces. In other examples, some or all of the one or more batteries 2124 may be non-rechargeable batteries.

[0102] One or more external power interfaces 2126 of the power system 2122 may include one or more wired power interfaces, such as a USB cable and / or a power cord, which enable a wired power connection to one or more power sources external to the computing device 2100. One or more external power interfaces 2126 may include one or more wireless power interfaces, such as a Qi wireless charger, which enables a wireless power connection to one or more external power sources, such as via the Qi wireless charger. Once a power connection to an external power source is established using one or more external power interfaces 2126, the computing device 2100 can draw power from the external power source to the established power connection. In some examples, the power system 2122 may include associated sensors, such as battery sensors associated with one or more batteries or other types of power sensors.

[0103] Cloud-based server

[0104] Figure 22 A network 2006 depicting computing clusters 2209a, 2209b, 2209c arranged as a cloud-based server system according to an example embodiment is shown. The computing clusters 2209a, 2209b, 2209c may be cloud-based devices that store program logic and / or data for cloud-based applications and / or services; for example, performing at least one function of and / or related to a convolutional neural network, confidence learning, predicting target images, predicting an original illumination model, convolutional neural network 430, confidence learning 630, and / or method 2300.

[0105] In some embodiments, the computing clusters 2209a, 2209b, 2209c can be a single computing device residing in a single computing center. In other embodiments, the computing clusters 2209a, 2209b, 2209c can include multiple computing devices in a single computing center, or even multiple computing devices in multiple computing centers located in diverse geographical locations. For example, Figure 22 depicts each of the computing clusters 2209a, 2209b, and 2209c residing in different physical locations.

[0106] In some embodiments, the data and services of the computing clusters 2209a, 2209b, 2209c can be encoded as computer-readable information stored in a non-transitory tangible computer-readable medium (or computer-readable storage medium) and accessible by other computing devices. In some embodiments, the computing clusters 2209a, 2209b, 2209c can be stored on a single disk drive or other tangible storage medium, or can be implemented on multiple disk drives or other tangible storage mediums located in one or more different geographical locations.

[0107] Figure 22 depicts a cloud-based server system according to an example embodiment. In Figure 22 which, the functions of the convolutional neural network, confidence learning, and / or the computing devices can be distributed among the computing clusters 2209a, 2209b, 2209c. The computing cluster 2209a can include one or more computing devices 2200a, a cluster storage array 2210a, and a cluster router 2211a connected by a local cluster network 2212a. Similarly, the computing cluster 2209b can include one or more computing devices 2200b, a cluster storage array 2210b, and a cluster router 2211b connected by a local cluster network 2212b. Likewise, the computing cluster 2209c can include one or more computing devices 2200c, a cluster storage array 2210c, and a cluster router 2211c connected by a local cluster network 2212c.

[0108] In some embodiments, each of the computing clusters 2209a, 2209b, and 2209c can have an equal number of computing devices, an equal number of cluster storage arrays, and an equal number of cluster routers. However, in other embodiments, each computing cluster can have a different number of computing devices, a different number of cluster storage arrays, and a different number of cluster routers. The number of computing devices, cluster storage arrays, and cluster routers in each computing cluster can depend on one or more computing tasks assigned to each computing cluster.

[0109] For example, in computing cluster 2209a, computing device 2200a may be configured to perform convolutional neural networks, confidence learning, and / or various computing tasks of the computing device. In one embodiment, the convolutional neural networks, confidence learning, and / or various functions of the computing device may be distributed among one or more of computing devices 2200a, 2200b, 2200c. Computing devices 2200b and 2200c in corresponding computing clusters 2209b and 2209c may be configured similarly to computing device 2200a in computing cluster 2209a. On the other hand, in some embodiments, computing devices 2200a, 2200b, and 2200c may be configured to perform different functions.

[0110] In some embodiments, based at least in part on the processing requirements of convolutional neural networks, confidence learning, and / or the computing device, the processing capabilities of computing devices 2200a, 2200b, 2200c, the latency of the network links between the computing devices within each computing cluster and between the computing clusters themselves, and / or other factors contributing to cost, speed, fault tolerance, resilience, efficiency, and / or other design goals of the overall system architecture, the computing tasks and stored data associated with convolutional neural networks, confidence learning, and / or the computing device may be distributed across computing devices 2200a, 2200b, and 2200c.

[0111] The cluster storage arrays 2210a, 2210b, 2210c of computing clusters 2209a, 2209b, 2209c may be data storage arrays that include a disk array controller configured to manage read and write access to a group of hard disk drives. The disk array controller, either alone or in combination with their respective computing devices, may also be configured to manage backup or redundant copies of the data stored in the cluster storage array to protect against disk drive or other cluster storage array failures and / or network failures that prevent one or more computing devices from accessing one or more cluster storage arrays.

[0112] Similar to the way in which functions such as convolutional neural networks, confidence learning, and / or computing devices can be distributed across the computing devices 2200a, 2200b, 2200c of the computing clusters 2209a, 2209b, 2209c, the various active and / or backup portions of these components can be distributed across the cluster storage arrays 2210a, 2210b, 2210c. For example, some cluster storage arrays can be configured to store a portion of the data of convolutional neural networks, confidence learning, and / or computing devices, while other cluster storage arrays can store other portions of the data of convolutional neural networks, confidence learning, and / or computing devices. Additionally, some cluster storage arrays can be configured to store backup versions of the data stored in other cluster storage arrays.

[0113] The cluster routers 2211a, 2211b, 2211c in the computing clusters 2209a, 2209b, 2209c can include networking devices configured to provide internal and external communication to the computing clusters. For example, the cluster router 2211a in the computing cluster 2209a can include one or more Internet switching and routing devices configured to provide (i) local area network communication between the computing device 2200a and the cluster storage array 2210a via the local cluster network 2212a, and (ii) wide area network communication between the computing cluster 2209a and the computing clusters 2209b and 2209c via the wide area network link 2213a to the network 2006. The cluster routers 2211b and 2211c can include networking devices similar to the cluster router 2211a, and the cluster routers 2211b and 2211c can perform networking functions similar to those performed by the cluster router 2211a for the computing cluster 2209a for the computing clusters 2209b and 2209b.

[0114] In some embodiments, the configuration of the cluster routers 2211a, 2211b, 2211c can be at least partially based on the data communication requirements of the computing devices and cluster storage arrays, the data communication capabilities of the networking devices in the cluster routers 2211a, 2211b, 2211c, the latency and throughput of the local cluster networks 2212a, 2212b, 2212c, the latency, throughput, and cost of the wide area network links 2213a, 2213b, 2213c, and / or other factors that can contribute to regulating the cost, speed, fault tolerance, resilience, efficiency, and / or other design criteria of the system architecture.

[0115] Example methods of operation

[0116] Figure 23is a flowchart of a method 2300 according to an example embodiment. The method 2300 may be performed by a computing device such as the computing device 2100. The method 2300 may begin at block 2310, where the computing device may train a neural network to apply an illumination model to an input image using confidence learning based on a light prediction and a prediction confidence value associated with the illumination of the input image, as discussed above at least in Figures 2 - 6 the context of

[0117] In some examples, training a neural network to apply an illumination model to an input image using confidence learning may include training a convolutional neural network to apply an illumination model to an input image using confidence learning, as discussed at least in Figure 5 and Figure 6 the context of. In some of these examples, the illumination of the input image may be modeled using an original light model. In such examples, training a convolutional neural network to apply an illumination model to an input image using confidence learning may include training a convolutional neural network using confidence learning based on a light prediction and a prediction confidence value associated with the original light model, as discussed above at least in Figure 5 and Figure 6 the context of. In some of these examples, training a convolutional neural network using confidence learning based on a light prediction and a prediction confidence value associated with the original light model includes training a convolutional neural network using confidence learning based on a light prediction of the original light model for a portion of the input image and a prediction confidence of the illumination prediction for that portion of the input image, as discussed above at least in Figure 6 the context of.

[0118] In some examples, training a neural network to apply an illumination model to an input image may include training the neural network using multiple images of an object, where the multiple images illuminate the object using multiple illumination models, as discussed above at least in Figures 2 - 5 the context of. In some examples, training the neural network may include training the neural network at the computing device, as discussed above at least in Figure 19 the context of.

[0119] At block 2320, the computing device may receive an input image of an object and data regarding a particular illumination model to be applied to the input image, as discussed above at least in Figures 4 - 18 the context of.

[0120] In some examples, the illumination of the input image may be modeled using an original light model. In such examples, determining the output image may further include using the trained neural network to determine the output image and a prediction of the original light model, as discussed above at least in Figures 4 - 18 the context of.

[0121] In some examples, the object can include an object that diffuses light, such as discussed at least in Figures 4 - 18 the context of. In some examples, the object can include a human face, such as discussed at least in Figures 4 - 18 the context of.

[0122] In some examples, the computing device can include a camera. In such examples, receiving an input image of the object can include generating an input image of the object using the camera and receiving the generated input image from the camera at the computing device, such as discussed at least in Figure 2 and Figure 3 the context of.

[0123] In some examples, the input image of the object can be a single image of the object, such as discussed at least in Figures 4 - 18 the context of.

[0124] At block 2330, the computing device can determine an output image of the object by applying a specific lighting model to the input image of the object using a trained neural network, such as discussed at least in Figures 4 - 18 the context of.

[0125] In some examples, receiving an input image of the object and data about a specific lighting model to be applied to the input image can include receiving an input image of the object and data about a plurality of specific lighting models to be applied to the input image, and determining the output image can include determining a plurality of output images by applying each of the plurality of specific lighting models to the input image, such as discussed at least in Figure 2 and Figure 3 the context of.

[0126] In some examples, determining an output image of the object by using a trained neural network can include obtaining a trained neural network at the computing device; and using the obtained neural network by the computing device to determine an output image of the object, such as discussed at least in Figure 19 the context of.

[0127] In some examples, method 2300 can further include using the computing device to provide the output image, such as discussed at least in Figures 4 - 18 the context of.

[0128] In some examples, the illumination of the input image can be modeled using a raw light model. In such examples, method 2300 can further include using the computing device to provide a prediction of the raw light model, such as discussed at least in Figures 4 - 18 the context of.

[0129] In some examples, determining an output image of an object by using a trained neural network can include: a computing device determining a request to apply a specific lighting model to an input image; sending, from the computing device to a second computing device, the request to apply the specific lighting model to the input image, the second computing device including the trained neural network; and after sending the request, the computing device receiving, from the second computing device, an output image of applying the specific lighting model to the input image of the object, as discussed at least in Figure 19 the context of.

[0130] Additional example embodiments

[0131] The following clauses are provided as further description of the present disclosure.

[0132] Clause 1 - A computer-implemented method, comprising: training a neural network to apply a lighting model to an input image by using confidence learning based on light prediction and prediction confidence values associated with the lighting of the input image; receiving, at a computing device, an input image of an object and data regarding a specific lighting model to be applied to the input image; and determining, by the computing device, an output image of the object by applying the specific lighting model to the input image of the object by using the trained neural network.

[0133] Clause 2 - The computer-implemented method according to Clause 1, wherein the lighting of the input image is modeled using an original light model, and wherein determining the output image further includes determining a prediction of the output image and the original light model by using the trained neural network.

[0134] Clause 3 - The computer-implemented method according to Clause 1 or Clause 2, wherein training the neural network to apply the lighting model to the input image by using confidence learning includes training a convolutional neural network to apply the lighting model to the input image by using confidence learning.

[0135] Clause 4 - The computer-implemented method according to Clause 3, wherein the lighting of the input image is modeled using an original light model, and wherein training the convolutional neural network to apply the lighting model to the input image by using confidence learning includes training the convolutional neural network by using confidence learning based on light prediction and prediction confidence values associated with the original light model.

[0136] Clause 5 - The computer-implemented method according to Clause 4, wherein training the convolutional neural network by using confidence learning based on light prediction and prediction confidence values associated with the original light model includes training the convolutional neural network by using confidence learning based on light prediction of the original light model of a portion of the input image and prediction confidence of the lighting of the portion of the input image.

[0137] Clause 6 - The computer-implemented method according to any one of Clauses 1 - 5, wherein training a neural network to apply an illumination model to an input image includes training the neural network using a plurality of images of an object, wherein the plurality of images illuminate the object using a plurality of illumination models.

[0138] Clause 7 - The computer-implemented method according to any one of Clauses 1 - 6, wherein the object includes an object that diffuses light.

[0139] Clause 8 - The computer-implemented method according to any one of Clauses 1 - 7, wherein the object includes a human face.

[0140] Clause 9 - The computer-implemented method according to any one of Clauses 1 - 8, wherein the computing device includes a camera, and wherein receiving an input image of an object includes: generating an input image of the object using the camera; and receiving, at the computing device, the generated input image from the camera.

[0141] Clause 10 - The computer-implemented method according to any one of Clauses 1 - 9, further comprising: providing, using the computing device, an output image.

[0142] Clause 11 - The computer-implemented method according to any one of Clauses 1 - 10, wherein the illumination of the input image is modeled using an original light model, and wherein the method further comprises: providing, using the computing device, a prediction of the original light model.

[0143] Clause 12 - The computer-implemented method according to any one of Clauses 1 - 11, wherein receiving an input image of an object and data regarding a specific illumination model to be applied to the input image includes receiving an input image of the object and data regarding a plurality of specific illumination models to be applied to the input image, and wherein determining the output image includes determining a plurality of output images by applying each of the plurality of specific illumination models to the input image.

[0144] Clause 13 - The computer-implemented method according to any one of Clauses 1 - 12, wherein the input image of the object is a single image of the object.

[0145] Clause 14 - The computer-implemented method according to any one of Clauses 1 - 13, wherein determining an output image of an object by using a trained neural network includes: obtaining, at the computing device, the trained neural network; and determining, by the computing device, the output image of the object using the obtained neural network.

[0146] Clause 15 - The computer-implemented method according to Clause 14, wherein training the neural network includes training the neural network at the computing device.

[0147] Article 16 - The computer-implemented method according to any one of Articles 1 - 15, wherein determining an output image of an object by using a trained neural network includes: determining, by a computing device, a request to apply a specific lighting model to an input image; sending, from the computing device, the request to apply the specific lighting model to the input image to a second computing device, the second computing device including the trained neural network; and after sending the request, receiving, by the computing device, from the second computing device an output image of applying the specific lighting model to the input image of the object.

[0148] Article 17 - A computing device, comprising: one or more processors; and a data storage device, wherein the data storage device stores computer-executable instructions thereon, which when executed by the one or more processors, cause the computing device to perform functions including the computer-implemented method according to any one of Articles 1 - 16.

[0149] Article 18 - An article of manufacture, comprising one or more computer-readable media having computer-readable instructions stored thereon, which when executed by one or more processors of a computing device, cause the computing device to perform functions including the computer-implemented method according to any one of Articles 1 - 16.

[0150] Article 19 - The article of manufacture according to Article 18, wherein the one or more computer-readable media include one or more non-transitory computer-readable media.

[0151] Article 20 - A computing device, comprising: means for performing the computer-implemented method according to any one of Articles 1 - 16.

[0152] The present disclosure is not limited to the specific embodiments described in this application, which are intended to be illustrative of various aspects. It will be apparent to those skilled in the art that many modifications and variations can be made without departing from its spirit and scope. In addition to those enumerated herein, functionally equivalent methods and apparatuses within the scope of the present disclosure will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims.

[0153] The foregoing detailed description has described various features and functions of the disclosed systems, devices, and methods with reference to the accompanying drawings. In the drawings, like symbols generally identify like components unless the context dictates otherwise. The illustrative embodiments described in the detailed description, the drawings, and the claims are not meant to be limiting. Other embodiments may be utilized and other changes may be made without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein and illustrated in the drawings, can be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are expressly contemplated herein.

[0154] Regarding any and all ladder diagrams, scenarios, and flowcharts in the drawings and as discussed herein, according to example embodiments, each block and / or communication may represent the processing of information and / or the transmission of information. Alternative embodiments are also included within the scope of these example embodiments. In these alternative embodiments, for example, depending on the functions involved, functions described as blocks, transmissions, communications, requests, responses, and / or messages may not be performed in the order shown or discussed, including substantially concurrently or in the reverse order. Additionally, more or fewer blocks and / or functions may be used with any of the ladder diagrams, scenarios, and flowcharts discussed herein, and these ladder diagrams, scenarios, and flowcharts may be combined with each other, in part or in whole. Figure One and these ladder diagrams, scenarios, and flowcharts may be used starting from, and these ladder diagrams, scenarios, and flowcharts may be combined with each other, in part or in whole.

[0155] Blocks representing information processing may correspond to circuits that may be configured to perform specific logical functions described herein or techniques. Alternatively or additionally, blocks representing information processing may correspond to a module, segment, or portion of program code (including associated data). The program code may include one or more instructions executable by a processor to implement a specific logical function or action in a method or technique. The program code and / or associated data may be stored on any type of computer-readable medium, such as a storage device including a magnetic disk or hard drive or other storage medium.

[0156] Computer-readable media may also include non-transitory computer-readable media, such as non-transitory computer-readable media for storing data for a short time, such as register memory, processor cache, and random access memory (RAM). Computer-readable media may also include non-transitory computer-readable media for storing program code and / or data for a longer time, such as auxiliary or persistent long-term storage, such as read-only memory (ROM), optical disk or magnetic disk, compact disc read-only memory (CD-ROM). Computer-readable media may also be any other volatile or non-volatile storage system. Computer-readable media may be considered, for example, a computer-readable storage medium or a tangible storage device.

[0157] In addition, blocks representing one or more information transmissions may correspond to information transmissions between software and / or hardware modules in the same physical device. However, other information transmissions may occur between software modules and / or hardware modules in different physical devices.

[0158] Although various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those of ordinary skill in the art. The various aspects and embodiments disclosed herein are provided for illustrative purposes and are not intended to be limiting, the true scope being indicated by the appended claims.

Claims

1. A computer-implemented method, comprising: receiving a training data set comprising a plurality of images, wherein each image of the plurality of images is associated with a corresponding illumination model, and wherein a given illumination model corresponding to a given image indicates the position of one or more environmental light sources relative to an object in the given image; training a neural network based on the training data set by: receiving, by a computing device, an input image and data regarding a target illumination model, predicting an initial illumination model associated with the input image, and predicting a relighting of the input image by replacing the initial illumination model with the target illumination model; and providing the trained neural network.

2. The computer-implemented method according to claim 1, wherein, the training is based on a cycle loss.

3. The computer-implemented method according to claim 1, wherein, the training is based on an L2 loss metric.

4. The computer-implemented method according to claim 1, wherein, the training is based on a log L1 loss metric.

5. The computer-implemented method according to claim 1, wherein, the training utilizes a deep supervision technique to constrain one or more intermediate layers of the neural network.

6. The computer-implemented method according to claim 1, wherein, the training is based on a generative adversarial network loss function.

7. The computer-implemented method according to claim 1, wherein, the training utilizes confidence learning, which is based on light predictions and prediction confidence values associated with the illumination of the input image.

8. The computer-implemented method according to claim 1, wherein, the object includes a reflection property of diffuse light.

9. The computer-implemented method according to claim 1, wherein, the object includes a human face.

10. The computer-implemented method according to claim 1, wherein, training the neural network includes training the neural network at a computing device.

11. The computer-implemented method according to claim 1, wherein, the initial illumination model, the target illumination model, and the given illumination model include data representing one or more of the following: (i) color, (ii) intensity, (iii) albedo, (iv) light direction, (v) surface normal, or (vi) one or more light sources, wherein at least one light source has a different position between the initial illumination model and the target illumination model.

12. The computer-implemented method according to claim 1, wherein, the plurality of images includes one or more objects under a plurality of different illumination conditions, the plurality of different illumination conditions including one or more of the following: (i) a first illumination provided from different directions, (ii) a second illumination provided at different intensities, (iii) a third illumination provided by light sources of different colors, or (iv) a fourth illumination provided by different numbers of light sources.

13. A computer-implemented method, comprising: receiving, by a computing device, an input image of an object and data regarding a target illumination model to be applied to the object; Prediction by a trained neural network of (i) an initial illumination model indicating the position of one or more environmental light sources relative to the object in the input image, and (ii) the relighting of the object by applying a target illumination model, the neural network having been trained by: Receiving a given input image and data regarding a given target illumination model, Predicting a given initial illumination model associated with the given input image, and Predicting a given relighting of the given input image by replacing the given initial illumination model with the given target illumination model; And Providing, by the computing device, an output image including the relighting of the object.

14. The computer-implemented method according to claim 13, Wherein, The object includes the reflection property of diffuse light.

15. The computer-implemented method according to claim 13, Wherein, The object includes a human face.

16. The computer-implemented method according to claim 13, Wherein, The relighting of the object is modeled using the initial illumination model predicted by the trained neural network, and wherein the method further includes: Providing the initial illumination model predicted by the trained neural network.

17. The computer-implemented method according to claim 13, Wherein, Providing the output image includes: Determining, by the computing device, a request to apply the target illumination model to the input image; Sending the request to apply the target illumination model to the input image from the computing device to a second computing device, the second computing device including a trained neural network; and After sending the request, receiving, by the computing device, from the second computing device the output image of applying the target illumination model to the input image.

18. The computer-implemented method according to claim 13, Wherein, Providing the output image includes: Obtaining, at the computing device, a trained neural network; and Determining the output image of the object by using the obtained neural network.

19. The computer-implemented method according to claim 13, Wherein, The computing device includes an image capture device, and wherein receiving the input image includes: Capturing the input image using the image capture device.

20. A computing device, Comprising: One or more processors; And A data storage device, wherein computer-executable instructions are stored on the data storage device, and when executed by the one or more processors, the instructions cause the computing device to perform functions including: Receiving, by the computing device, an input image of an object and data regarding a target illumination model to be applied to the object; Predicting, by a trained neural network, (i) an initial illumination model indicating the position of one or more environmental light sources relative to the object in the input image, and (ii) the relighting of the object by applying the target illumination model, the neural network having been trained by: Receiving a given input image and input data regarding a given target illumination model, Predicting a given initial illumination model associated with the given input image, and Predicting a given relighting of the given input image by replacing the given initial illumination model with the given target illumination model; and The computing device provides an output image including relighting of an object.

Citation Information

Patent Citations

  • Simulation lighting device and simulation lighting method

    CN103678761A

  • Editing digital images utilizing a neural network with an in-network rendering layer

    US20180253869A1