Dynamically Estimating Lighting Parameters of Locations in Augmented Reality Scenes Using Neural Networks

By using local lighting estimation neural network in augmented reality system and dynamically adjusting the lighting parameters of virtual objects, the problem of unnatural lighting in the prior art is solved, and a more realistic and flexible lighting effect is achieved.

CN111723902BActive Publication Date: 2025-05-13ADOBE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201911159261.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-21
Filing Date
2019-11-22
Publication Date
2025-05-13
Estimated Expiration
2039-11-22

AI Technical Summary

Technical Problem

Existing augmented reality systems have difficulty adjusting lighting conditions in real time to reflect changes in the real world, especially in three-dimensional scenarios, resulting in unnatural lighting of virtual objects.

Method used

The local illumination estimation neural network is used to extract the global feature map of the digital scene and generate local position indicators, and the lighting parameters are dynamically adjusted to match the specific position and viewing angle of the virtual object.

Benefits of technology

It realizes real-time and natural adjustment of the lighting conditions of virtual objects in an augmented reality system, improves the authenticity and flexibility of lighting, and can quickly respond to scene changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111723902B_ABST
    Figure CN111723902B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to dynamically estimating lighting parameters of locations in an augmented reality scene using a neural network. The present disclosure relates to methods, non-transitory computer-readable media, and systems for estimating lighting parameters of specific locations within a digital scene for augmented reality using a local lighting estimation neural network. For example, based on a request to draw a virtual object in a digital scene, the system generates location-specific lighting parameters for a specified location within the digital scene using a local lighting estimation neural network. The system also draws a modified digital scene including a virtual object at the specified location based on the parameters. The system generates such location-specific lighting parameters to spatially change and adapt lighting conditions for different locations within the digital scene. Since the request to draw a virtual object is real-time (or near real-time), the system can quickly generate different location-specific lighting parameters in response to the drawing request, and these parameters accurately reflect the lighting conditions of different locations within the digital scene.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Augmented reality systems typically depict digitally enhanced images or other scenes with computer-simulated objects. To depict such scenes, augmented reality systems sometimes render real objects and computer-simulated objects with shadows and other lighting conditions. Many augmented reality systems attempt to seamlessly render virtual objects that are composited with objects in the real world. To achieve convincing synthesis, augmented reality systems must illuminate virtual objects with consistent lighting that matches the physical scene. Because the real world is constantly changing (e.g., objects move, lighting changes), augmented reality systems that capture lighting conditions in advance are typically unable to adjust lighting conditions to reflect changes in the real world.

[0002] Despite advances in estimating lighting conditions for digitally augmented scenes, several technical limitations still prevent conventional augmented reality systems from realistically depicting lighting conditions on a computing device. These limitations include changing lighting conditions as the digitally augmented scene changes, rapidly rendering or adjusting lighting conditions in real-time (or near real-time), and faithfully capturing lighting variations across the scene. These limitations are exacerbated in three-dimensional scenes, where every location at a given moment can receive a different amount of light from the full 360-degree directional range. Both directional dependence and lighting variations across the scene play a critical role when trying to faithfully and convincingly render synthetic objects into a scene.

[0003] For example, some conventional augmented reality systems are unable to depict the lighting conditions of computer-simulated objects in real time (or near real time). In some cases, conventional augmented reality systems use an ambient light model (i.e., only a single constant term without directional information) to estimate the light received by an object from its environment. For example, conventional augmented reality systems often use simple heuristics to create lighting conditions, such as by relying on the average brightness value of pixels of an object (or surrounding an object) to create lighting conditions in an ambient light model. Such an approximation does not capture directional variations in lighting, and may not produce a reasonable approximation of ambient lighting under many conditions - resulting in unrealistic and unnatural lighting. Such lighting makes computer-simulated objects appear unrealistic or out of place in a digitally enhanced scene. For example, in some cases, when the light used for a computer-simulated object comes from outside the perspective (or viewpoint) shown in the digitally enhanced image, conventional systems cannot accurately depict lighting on the object.

[0004] In addition to the challenges of depicting realistic lighting, in some cases, conventional augmented reality systems are unable to flexibly adjust or change lighting conditions for specific computer-simulated objects in a scene. For example, some augmented reality systems determine the lighting conditions of a digitally enhanced image as a group of objects or the entire image, rather than the lighting conditions of a specific object or location within the digitally enhanced image. Because such lighting conditions are generally applicable to a group of objects or the entire image, conventional systems cannot adjust the lighting conditions for specific objects, or can only do so by re-determining the lighting conditions for the entire digitally enhanced image, which is an inefficient use of computing resources.

[0005] Independent of technical limitations that affect the realism or flexibility of lighting in augmented reality, conventional augmented reality systems are sometimes unable to quickly estimate lighting conditions for objects in a digitally augmented scene. For example, some conventional augmented reality systems receive user input that defines baseline parameters such as image geometry or material properties, and estimate parametric lighting for the digitally augmented scene based on the baseline parameters. Although some conventional systems can apply such user-defined parameters to accurately estimate lighting conditions, such systems can neither quickly estimate parametric lighting nor apply image geometry-specific lighting models to other scenes with different light sources and geometries. Summary of the invention

[0006] The present disclosure describes embodiments of methods, non-transitory computer-readable media, and systems that, in addition to providing other benefits, solve the aforementioned problems. For example, based on a request to draw a virtual object in a digital scene, the disclosed system generates location-specific lighting parameters for a specified location within the digital scene using a local illumination estimation neural network. In certain implementations, the disclosed system draws a modified digital scene that includes a virtual object at a specified location illuminated according to the location-specific lighting parameters. As described below, the disclosed system can generate such location-specific lighting parameters to spatially change the lighting for different locations within the digital scene. Therefore, because the request to draw the virtual object is real-time (or near real-time), the disclosed system can quickly generate different location-specific lighting parameters based on such drawing requests that accurately reflect the lighting conditions at different locations of the digital scene.

[0007] For example, in some embodiments, the disclosed system identifies a request to draw a virtual object at a specified location within a digital scene. The disclosed system extracts a global feature map from the digital scene using a first set of network layers of a local illumination estimation neural network. The system also generates a local position indicator for the specified location and modifies the global feature map of the digital scene based on the local position indicator. Based on the modified global feature map, the system generates location-specific lighting parameters for the specified location using a second set of layers of the local illumination estimation neural network. In response to the drawing request, the system draws the modified digital scene, which includes the virtual object at the specified location illuminated according to the location-specific lighting parameters.

[0008] Additional features and advantages of the disclosed methods, non-transitory computer-readable media, and systems are set forth in the following description and may be apparent from, or may be disclosed from, practice of the exemplary embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The detailed description refers to the accompanying drawings which are briefly described below.

[0010] Figure 1 An augmented reality system and a lighting estimation system are illustrated that use a local lighting estimation neural network to generate location-specific lighting parameters for a specified location within a digital scene, and draw a modified digital scene including a virtual object at the specified location according to the parameters of one or more embodiments.

[0011] Figure 2 Illustrated are digital training scenes and corresponding cubemaps in accordance with one or more embodiments.

[0012] Figure 3 Graphs illustrating degrees of position-specific spherical harmonic training coefficients in accordance with one or more embodiments.

[0013] Figure 4A The diagram illustrates an illumination estimation system that trains a local illumination estimation neural network to generate location-specific spherical harmonic coefficients for each specified location within a digital scene, in accordance with one or more embodiments.

[0014] Figure 4B The diagram illustrates an illumination estimation system that uses a trained local illumination estimation neural network to generate location-specific spherical harmonic coefficients for a specified location within a digital scene in accordance with one or more embodiments.

[0015] Figure 5A The diagram illustrates an illumination estimation system that trains a local illumination estimation neural network to generate location-specific spherical harmonic coefficients for each specified location within a digital scene using coordinates of feature maps from layers of the neural network, in accordance with one or more embodiments.

[0016] Figure 5B The diagram illustrates an illumination estimation system that uses a trained local illumination estimation neural network to generate location-specific spherical harmonic coefficients for a specified location in a digital scene using coordinates of feature maps from layers of the neural network, in accordance with one or more embodiments.

[0017] Figure 6A-6C The diagram illustrates a computing device that renders a digital scene including virtual objects at a designated location and a new designated location according to location-specific lighting parameters and new location-specific lighting parameters, respectively, in response to different rendering requests, according to one or more embodiments.

[0018] Fig. 7A The diagram illustrates a digital scene and corresponding red, green, and blue representations and light intensity representations rendered according to base real lighting parameters and location-specific lighting parameters generated by a lighting estimation system in accordance with one or more embodiments.

[0019] Figure 7B The diagram illustrates a digital scene including a virtual object at a specified location illuminated according to base real lighting parameters and a digital scene including a virtual object at a specified location illuminated according to location specific lighting parameters generated by a lighting estimation system, according to one or more embodiments.

[0020] Figure 8 FIG. 1 is a block diagram illustrating an environment in which a lighting estimation system may operate in accordance with one or more embodiments.

[0021] Fig. 9 The diagram illustrates an example of a method according to one or more embodiments. Figure 8 Schematic diagram of the lighting estimation system.

[0022] Fig.10 A flow diagram illustrating a series of actions for training a local lighting estimation neural network to generate location-specific lighting parameters in accordance with one or more embodiments.

[0023] Fig.11 A flow diagram illustrating a series of actions for applying a trained local illumination estimation neural network to generate location-specific illumination parameters in accordance with one or more embodiments.

[0024] Fig.12 The figure illustrates a block diagram of an exemplary computing device for implementing one or more embodiments of the present disclosure. DETAILED DESCRIPTION

[0025] The present disclosure describes one or more embodiments of an illumination estimation system that uses a local illumination estimation neural network to estimate illumination parameters for a specific location within a digital scene for augmented reality. For example, based on a request to draw a virtual object in a digital scene, the illumination estimation system generates location-specific illumination parameters for a specified location within the digital scene using a local illumination estimation neural network. In some implementations, the illumination estimation system also draws a modified digital scene including a virtual object at the specified location according to the parameters. In some embodiments, the illumination estimation system generates such location-specific illumination parameters to spatially change and adapt illumination conditions for different locations within the digital scene. Because the request to draw the virtual object is in real-time (or near real-time), the illumination estimation system can quickly generate different location-specific illumination parameters that accurately reflect illumination conditions at different locations within the digital scene or reflect illumination conditions from different perspectives of the digital scene in response to the drawing request. The illumination estimation system can also quickly generate different location-specific illumination parameters that reflect changes in illumination or other conditions.

[0026] For example, in some embodiments, the lighting estimation system identifies a request to draw a virtual object at a specified location within a digital scene. To draw such a scene, the lighting estimation system extracts a global feature map from the digital scene using a first set of network layers of a local lighting estimation neural network. The lighting estimation system also generates a local position indicator for the specified location and modifies the global feature map of the digital scene based on the local position indicator. Based on the modified global feature map, the lighting estimation system generates location-specific lighting parameters for the specified location using a second set of layers of the local lighting estimation neural network. In response to the drawing request, the lighting estimation system draws the modified digital scene, which includes the virtual object illuminated at the specified location according to the location-specific lighting parameters.

[0027] By using location-specific lighting parameters, in some embodiments, the lighting estimation system can both illuminate virtual objects from different perspectives of the scene and quickly update lighting conditions for different locations, different perspectives, changes in lighting to the scene or virtual objects in the scene, or other environmental changes. For example, in some cases, the lighting estimation system generates location-specific lighting parameters that capture the lighting conditions of the location of the virtual object from various perspectives within the digital scene. After identifying a position adjustment request to move the virtual object to a new specified location, the lighting estimation system can also generate a new local position indicator for the new specified location and modify the global feature map for the digital scene using a neural network to output new lighting parameters for the new specified location. When identifying or otherwise responding to a change in the lighting conditions of the digital scene, the lighting estimation system can also update the global feature map of the digital scene to output new lighting parameters for the new lighting conditions. For example, when the viewer's viewpoint changes (e.g., the camera moves in the scene), when the lighting in the scene changes (e.g., lights are added, dimmed, blocked, exposed), when objects in the scene or the scene itself changes, the lighting estimation system can dynamically determine or update the lighting parameters.

[0028] To generate location-specific lighting parameters, the lighting estimation system can use different types of local position indicators. For example, in some embodiments, the lighting estimation system identifies (as a local position indicator) local position coordinates representing a specified location within the digital scene. In contrast, in some implementations, the lighting estimation system identifies one or more local position indicators from features extracted by different layers of the local lighting estimation neural network, such as pixels corresponding to the specified location from different feature maps extracted by the neural network layers.

[0029] When generating location-specific lighting parameters, the lighting estimation system can generate spherical harmonic coefficients that indicate lighting conditions for a specified location of a virtual object in a digital scene. When the digital scene is represented with low dynamic range ("LDR") lighting, such location-specific spherical harmonic coefficients can capture high dynamic range ("HDR") lighting at a location within the digital scene. When the virtual object changes position in the digital scene, the lighting estimation system can generate new location-specific spherical harmonic coefficients using a local lighting estimation neural network by requesting that the change in lighting be realistically depicted at the changed location of the virtual object.

[0030] As described above, in some embodiments, the lighting estimation system not only applies a local lighting estimation neural network, but can also optionally train such a network to generate location-specific lighting parameters. When training the neural network, in some implementations, the lighting estimation system extracts a global feature training map from a digital training scene using a first set of layers of the local lighting estimation neural network. The lighting estimation system also generates a local location training indicator for a specified location within the digital training scene, and modifies the global feature training map based on the local location training indicator for the specified location.

[0031] Based on the modified global feature training graph, the illumination estimation system generates location-specific illumination training parameters for the specified location using a second set of network layers of the local illumination estimation neural network. The illumination estimation system then modifies the network parameters of the local illumination estimation neural network based on a comparison of the location-specific illumination training parameters for the specified location within the digital training scene with the underlying real-world illumination parameters. By iteratively generating such location-specific illumination training parameters and adjusting the network parameters of the neural network, the illumination estimation system can train the local illumination estimation neural network to a convergence point.

[0032] As previously described, the lighting estimation system can specify locations using ground truth lighting parameters to facilitate training. To create such ground truth lighting parameters, in some embodiments, the lighting estimation system generates cube maps for various locations within the digital training scene. The lighting estimation system then projects the cube map of the digital training scene to ground truth spherical harmonic coefficients. Such ground truth spherical harmonic coefficients can be used for comparison when iteratively training the local lighting estimation neural network.

[0033] As described above, the disclosed lighting estimation system overcomes several technical deficiencies that have hampered conventional augmented reality systems. For example, the lighting estimation system improves the accuracy and realism with which existing augmented reality systems generate lighting conditions for specific locations within a digital scene. As described above, the lighting estimation system can create such realistic lighting in part by using a local lighting estimation neural network that is trained to generate location-specific spherical lighting parameters based on local location indicators for specified locations within a digital scene.

[0034] Unlike some conventional systems that use average brightness values ​​that result in unnatural brightness, the disclosed lighting estimation system can create lighting parameters with coordinate-level accuracy corresponding to local position indicators. In addition, unlike some conventional systems that cannot depict lighting from outside the perspective of the digital scene, the disclosed lighting estimation system can create lighting parameters that capture lighting conditions emitted from light sources outside the perspective of the digital scene. To achieve this accuracy, in some embodiments, the lighting estimation system generates location-specific spherical harmonic coefficients that effectively capture location-specific realistic and natural-looking lighting conditions from multiple viewpoints within the digital scene.

[0035] In addition to more realistically depicting lighting, in some embodiments, the lighting estimation system exhibits greater flexibility in drawing different lighting conditions for different locations relative to existing augmented reality systems. Unlike some conventional augmented reality systems, which are limited to redetermining lighting for a group of objects or an entire image, the lighting estimation system can flexibly adapt lighting conditions for different locations to which a virtual object is moved. For example, after identifying a position adjustment request for a moving virtual object, the disclosed lighting estimation system can modify an existing global feature map of a digital scene using a new local position indicator. By modifying the global feature map to reflect the new specified location, the lighting estimation system can generate new location-specific lighting parameters for the new specified location without having to redetermine lighting conditions for other objects or the entire image. This flexibility enables users to manipulate objects in augmented reality applications on mobile devices or other computing devices.

[0036] Being reality independent and flexible, the disclosed lighting estimation system can also increase the speed at which an augmented reality system can draw a digital scene with location-specific lighting for a virtual object. Unlike lighting models that rely on examining the geometry of an image or similar baseline parameters, the disclosed lighting estimation system estimates lighting using a neural network that requires relatively few inputs, namely indicators of the digital scene and the location of the virtual object. By training a local lighting estimation neural network to analyze such inputs, the lighting estimation system reduces the computational resources required to quickly generate lighting parameters for a specific location within a digital scene.

[0037] Now go to Figure 1 , which illustrates an augmented reality system 108 and a lighting estimation system 110 that use a neural network to estimate location-specific lighting parameters. In general, Figure 1 As shown, the lighting estimation system 110 identifies a request to draw a virtual object 106 at a specified location in the digital scene 102, and generates location-specific lighting parameters 114 for the specified location using the local lighting estimation neural network 112. Based on the request, the augmented reality system 108, together with the lighting estimation system 110, draws a modified digital scene 116 that includes the virtual object 106 at the specified location that is illuminated according to the location-specific lighting parameters 114. Figure 1 Augmented reality system 108 is depicted including lighting estimation system 110 and rendering modified digital scene 116 , but lighting estimation system 110 may alternatively render modified digital scene 116 alone.

[0038] As just noted, the lighting estimation system 110 identifies a request to draw a virtual object 106 at a specified location within the digital scene 102. For example, the lighting estimation system 110 may identify a digital request from a mobile device to draw a virtual pillow (or other virtual item) at a specific location on a piece of furniture (or another real item) depicted in a digital image. Regardless of the type of object or scene from the request, in some embodiments, the request to draw a digital scene includes an indication of a specified location to draw a virtual object.

[0039] As used in this disclosure, the term "digital scene" refers to a digital image, model, or depiction of an object. For example, in some embodiments, a digital scene includes a digital image of a real scene from a particular viewpoint or from multiple viewpoints. As another example, a digital scene may include a three-dimensional digital model of a scene. Regardless of the format, a digital scene may include a depiction of light from a light source. As just one example, a digital scene may include a digital image of a real room containing real walls, carpets, furniture, and people, with light emitted from a lamp or window. As discussed further below, a digital scene may be modified to include virtual objects in an adjusted or modified digital scene that depicts augmented reality.

[0040] Relatedly, the term "virtual object" refers to a computer-generated graphical object that does not exist in the physical world. For example, a virtual object may include an object created by a computer for use in an augmented reality application. Such a virtual object may be, but is not limited to, a virtual accessory, an animal, clothing, cosmetics, footwear, fixtures, furniture, furnishings, hair, a person, a human feature, a vehicle, or any other graphical object created by a computer. This disclosure generally uses the word "virtual" to designate a particular virtual object (e.g., a "virtual pillow" or "virtual shoes"), but generally refers to a real object without the word "real" (e.g., a "bed," a "sofa").

[0041] like Figure 1 As further shown, the lighting estimation system 110 identifies or generates a local position indicator 104 for a specified position in a drawing request. As used herein, the term "local position indicator" refers to a digital identifier for a position within a digital scene. For example, in some implementations, the local position indicator includes digital coordinates, pixels, or other indicia that indicate a specified position within a digital scene from a request to draw a virtual object. For illustration, the local position indicator can be a coordinate representing the specified position or a pixel (or coordinates of a pixel) corresponding to the specified position. In other embodiments, the lighting estimation system 110 can generate (and input) the local position indicator into the local lighting estimation neural network 112, or use the local lighting estimation neural network 112 to identify one or more local position indicators from a feature map.

[0042] In addition to generating the local position indicators, illumination estimation system 110 analyzes one or both of digital scene 102 and local position indicators 104 using local illumination estimation neural network 112. For example, in some cases, illumination estimation system 110 extracts a global feature map from digital scene 102 using a first set of layers of local illumination estimation neural network 112. Illumination estimation system 110 may further modify the global feature map of digital scene 102 based on local position indicators 104 for the specified location.

[0043] As used herein, the term "global feature map" refers to a multidimensional array or multidimensional vector representing features of a digital scene (e.g., a digital image or a three-dimensional digital model). For example, a global feature map for a digital scene can represent different visual or latent features of the entire digital scene, such as illumination or geometric features visible or embedded in the digital image or three-dimensional digital model. As described below, one or more layers of a local illumination estimation neural network output a global feature map of the digital scene.

[0044] The term "local illumination estimation neural network" refers to an artificial neural network that generates lighting parameters indicative of lighting conditions at a location within a digital scene. In particular, in certain implementations, the local illumination estimation neural network refers to an artificial neural network that generates a location-specific lighting parameter image indicative of lighting conditions for a specified location corresponding to a virtual object within the digital scene. In some embodiments, the local illumination estimation neural network includes some or all of the following network layers: one or more layers in a densely connected convolutional network ("DenseNet"), a convolutional layer, and a fully connected layer.

[0045] After modifying the global feature map, the lighting estimation system 110 uses a second set of network layers of the local lighting estimation neural network 112 to generate location-specific lighting parameters 114 based on the modified global feature map. As used in the present disclosure, the term "location-specific lighting parameters" refers to parameters that indicate lighting or illuminating a portion of a digital scene or lighting or illuminating a location in a digital scene. For example, in some embodiments, the location-specific lighting parameters define, specify, or otherwise indicate lighting or shading of pixels corresponding to a specified location of the digital scene. Such location-specific lighting parameters may define the shading or hue of a pixel of a virtual object at the specified location. In some embodiments, the location-specific lighting parameters include spherical harmonic coefficients that indicate lighting conditions at a specified location within the digital scene for the virtual object. Thus, the location-specific lighting parameters may be a function corresponding to the surface of a sphere.

[0046] like Figure 1As further shown, the augmented reality system 108, in addition to generating such lighting parameters, renders the modified digital scene 116, which includes the virtual object 106 illuminated at the specified location according to the location-specific lighting parameters 114. For example, in some embodiments, the augmented reality system 108 overlays or otherwise integrates a computer-generated image of the virtual object 106 within the digital scene 102. As part of the rendering, the augmented reality system 108 selects and renders pixels of the virtual object 106 that reflect the lighting, shading, or appropriate hue indicated by the location-specific lighting parameters 114.

[0047] As suggested above, in some embodiments, the lighting estimation system 110 uses a cubemap for the digital scene to project ground truth lighting parameters for specified locations of the digital scene. Figure 2 An example of a cubemap corresponding to a digital scene is shown in FIG. Figure 2 As shown, the digital training scene 202 includes viewpoints of objects illuminated by light sources. To generate ground truth lighting parameters for training, in some cases, the lighting estimation system 110 selects and identifies locations in the digital training scene 202. The lighting estimation system 110 also generates cubemaps 204a-204d corresponding to the identified locations, wherein each cubemap represents the identified location within the digital training scene 202. The lighting estimation system 110 then projects the cubemaps 204a-204d to ground truth spherical harmonic coefficients for training the local lighting estimation neural network.

[0048] The lighting estimation system 110 optionally generates or prepares a digital training scene, such as the digital training scene 202, by modifying an image of a realistic scene or a computer-generated scene. For example, in some cases, the lighting estimation system 110 modifies a three-dimensional scene from the Princeton University SUNCG dataset, as described by Shuran Song et al., “Semantic Scene Completion from a Single Depth Image,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017), the entire contents of which are incorporated herein by reference. The scenes in the SUNCG dataset typically include realistic room and furniture layouts. Based on the SUNCG dataset, the lighting estimation system 110 computes a physically based rendering of the scene image. In some such cases, the lighting estimation system 110 uses the Mitsuba framework to compute physically based rendering, as described by Yinda Zhang et al., “Physically-Based Rendering for Indoor Scene Understanding Using Convolutional Neural Networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017) (hereinafter referred to as “Zhang”), the entire contents of which are incorporated herein by reference.

[0049] To eliminate some of the inaccuracies and biases in this rendering, in some embodiments, the lighting estimation system 110 changes the calculation method of Zhang's algorithm in some aspects. First, the lighting estimation system 110 digitally removes lights that appear inconsistent with indoor scenes, such as area lights for floors, walls, and ceilings. Second, instead of using a single panorama for outdoor lighting as in Zhang, the researchers randomly selected a panorama from a dataset of 200 HDR outdoor panoramas and applied a random rotation around the Y axis of the panorama. Third, instead of assigning the same intensity to each indoor area light, the lighting estimation system 110 randomly selects light intensities between one hundred and five hundred candelas with a uniform distribution. However, in some implementations of generating digital training scenes, the lighting estimation system 110 uses the same rendering method and spatial resolution as described by Zhang.

[0050] Because indoor scenes may include arbitrary distributions of light sources and light intensities, the lighting estimation system 110 normalizes the rendering of each physically based digital scene. When normalizing the rendering, the lighting estimation system 110 uses the following equation:

[0051]

[0052] In equation (1), I represents the original image rendering with HDR, I′ represents the re-exposed image rendering, m is set to a value of 0.8, and P 90 represents 90% of the original image drawing I. By re-exposing the original image drawing I, the re-exposed image drawing I′ still contains HDR values. The lighting estimation system 110 further applies a gamma tone mapping operation with random values ​​between 1.8 and 2.2 to the re-exposed image drawing I′, and clips all values ​​greater than 1. By applying the gamma tone mapping operation, the lighting estimation system 110 can produce an image with a saturated bright window and improved scene contrast.

[0053] As described above, lighting estimation system 110 identifies sample locations within a digital training scene, such as digital training scene 202. Figure 2 As shown, digital training scene 202 includes four spheres that identify four sample locations identified by lighting estimation system 110. For illustrative purposes, digital training scene 202 includes the spheres. Therefore, the spheres represent visualizations of the sample locations identified by lighting estimation system 110, rather than objects in digital training scene 202.

[0054] To identify such sample locations, in some embodiments, lighting estimation system 110 identifies four different quadrants of digital training scene 202 with a margin of 20% of the image resolution from the image border. In each quadrant, lighting estimation system 110 identifies a sample location.

[0055] As further noted above, the lighting estimation system 110 generates cube maps 204a-204d based on sample positions from the digital training scene 202. Figure 2 As shown, each cube map 204a-204d includes six visual parts in HDR, which represent six different perspectives of the digital training scene 202. Cube map 204a, for example, includes visual parts 206a-206e. In some cases, each cube map 204a-204d includes a 64×64×6 resolution. Such a resolution has little effect on spherical harmonic functions of the fifth order or less and facilitates fast drawing of cube maps.

[0056] To draw the cubemaps 204a-204d, the lighting estimation system 110 may apply a two-stage principal sample space Metropolis light transport ("PSSMLT") with 512 direct samples. When generating the visual portion of the cubemap, such as the visual portion 206c, the lighting estimation system 110 translates the surface position along the surface normal by 10 centimeters to minimize the risk of placing a portion of the cubemap within the surface of another object. In some implementations, the lighting estimation system 110 uses the same method to identify sample locations in a digital training scene and generate corresponding cubemaps as a prelude to determining the underlying true position-specific spherical harmonic coefficients.

[0057] For example, after generating the cubemaps 204a-204d, the lighting estimation system 110 projects the cubemaps 204a-204d to the ground truth position-specific spherical harmonic coefficients for each identified position in the digital training scene 202. In some cases, the ground truth position-specific spherical harmonic coefficients include coefficients of the fifth order. To calculate such spherical harmonics, in some embodiments, the lighting estimation system 110 applies a least squares method to project the cubemaps.

[0058] For example, the lighting estimation system 110 may project the cubemap using the following equation:

[0059]

[0060] In equation (2), f represents the light intensity in each direction shown by the visible part of the cube map, where the light intensity is weighted corresponding to the solid angle of the pixel position. Denotes a spherical harmonic function of degree l and order m. In some cases, for each cubemap, the illumination estimation system 110 computes spherical harmonic coefficients of degree 5 (or some other order) for each color channel (eg, third order), resulting in 36×3 spherical harmonic coefficients.

[0061] In some embodiments, in addition to generating the ground truth position-specific spherical harmonic coefficients, the lighting estimation system 110 also enhances the digital training scene in specific ways. First, the lighting estimation system 110 randomly scales the exposure to a uniform distribution between 0.2 and 4. Second, the lighting estimation system 110 randomly sets the gamma value used for the tone mapping operator between 1.8 and 2.2. Third, the lighting estimation system 110 inverts the viewpoint of the digital training scene on the X-axis. Similarly, the lighting estimation system 110 enhances the digital training scene by inverting the negative order harmonics (such as the sign ) to transform the underlying real spherical harmonic coefficients to match the inverted viewpoint.

[0062] As further suggested above, in some implementations, the illumination estimation system 110 may use spherical harmonic coefficients of different orders. For example, the illumination estimation system 110 may generate five-order fundamental true spherical harmonic coefficients for each color channel, or five-order position-specific spherical harmonic coefficients for each color channel. Figure 3 Illustrated are lighting conditions according to different orders of spherical harmonics, visual representations of various orders of spherical harmonics, and a complete environment map for a location within a digital scene.

[0063] like Figure 3 As shown, visual representations 302a, 302b, 302c, 302d, and 302e correspond to first, second, third, fourth, and fifth order spherical harmonics, respectively. In particular, each of visual representations 302a-302e includes one or more spheres in a row that visually represent different orders. Figure 3 As shown, with each increase in order, the spherical harmonic coefficients indicate more detailed lighting conditions.

[0064] To illustrate, Figure 3 Lighting depictions 306a-306e of a complete environment map 304 for a location within a digital scene are included. Lighting depictions 306a, 306b, 306c, 306d, and 306e correspond to first, second, third, fourth, and fifth order spherical harmonic coefficients, respectively. As the order of spherical harmonics increases, lighting depictions 306a-306e better capture the lighting shown in the complete environment map 304. As shown in lighting depictions 306a-306e, spherical harmonic coefficients can implicitly capture occlusion and geometry of the digital scene.

[0065] As suggested above, the lighting estimation system 110 may use a variety of architectures and inputs for the local lighting estimation neural network. Figure 4A and Figure 4B An example of an illumination estimation system 110 is depicted that trains and applies, respectively, a local illumination estimation neural network to generate location-specific illumination parameters. Figure 5A and 5B Another embodiment of the lighting estimation system 110 is depicted, which trains and applies a local lighting estimation neural network to generate location-specific lighting parameters. Both embodiments can use digital training scenes and corresponding spherical harmonic coefficients from a cube map, such as Figure 2 and Figure 3 as described.

[0066] For example, Figure 4AAs shown, the illumination estimation system 110 iteratively trains the local illumination estimation neural network 406. As an overview of the training iterations, the illumination estimation system 110 extracts a global feature training map from the digital training scene using a first set of network layers 408 of the local illumination estimation neural network 406. The illumination estimation system 110 also generates a local position training indicator for a specified position within the digital training scene, and modifies the global feature training map based on the local position training indicator.

[0067] Based on the modification of the global feature training map reflected in the combined feature training map, the illumination estimation system 110 generates position-specific spherical harmonic coefficients for the specified position using the second set of network layers 416 of the local illumination estimation neural network 406. The illumination estimation system 110 then modifies the network parameters of the local illumination estimation neural network 406 based on a comparison of the position-specific spherical harmonic coefficients for the specified position in the digital training scene with the underlying true spherical harmonic coefficients.

[0068] For example, Figure 4A As shown, the lighting estimation system 110 feeds the digital training scene 402 to the local lighting estimation neural network 406. After receiving the digital training scene 402 as training input, the first set of network layers 408 extracts a global feature training map 410 from the digital training scene 402. In some such embodiments, the global feature training map 410 represents visual features of the digital training scene 402, such as the location of light sources and the global geometry of the digital training scene 402.

[0069] As described above, in some implementations, the first set of network layers 408 includes layers of DenseNet, such as the various lower layers of DenseNet. For example, the first set of network layers 408 may include a convolutional layer, followed by a dense block, and (in some cases) one or more groups of convolutional layers, pooling layers, and dense blocks. In some cases, the first set of network layers 408 includes layers from DenseNet 120, as described in "Densely Connected Convolutional Layers" (hereinafter referred to as "Huang") by G. Huang et al. in the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017), the entire contents of which are incorporated herein by reference. The lighting estimation system 110 optionally initializes network parameters for each layer of the DenseNet using weights trained on ImageNet, as described in "ImageNet Large Scale Visual Recognition Challenge" by Olga Russakovsky et al., International Journal of Computer Vision, Vol. 30, No. 3, 211-252 (2015) (hereinafter referred to as "Russakovsky"), the entire contents of which are incorporated herein by reference. Regardless of the architecture for initializing network parameters for the first set of network layers 408, the first set of network layers 408 optionally outputs a global feature training map 410 in the form of a dense feature map corresponding to the digital training scene 402.

[0070] In an alternative to the DenseNet layers, the first set of network layers 408 includes an encoder from a convolutional neural network (“CNN”) that includes several convolutional layers followed by four residual layers. In some such embodiments, the first set of network layers 408 includes an encoder described by Marc-André Gardner et al., “Learning to Predict Indoor Illumination from a Single Image,” ACM Transactions on Graphics (2017), Vol. 36, No. 6 (hereinafter referred to as “Gardner”), the entire contents of which are incorporated herein by reference. Thus, as an encoder, the first set of network layers 408 optionally outputs a global feature training map 410 in the form of an encoded feature map of the digital training scene 402.

[0071] like Figure 4AAs further shown, the illumination estimation system 110 identifies or generates a local position training indicator 404 for a specified position within the digital training scene 404. For example, the illumination estimation system 110 may identify local position training coordinates representing the specified position within the digital training scene 402. Such coordinates may represent two-dimensional coordinates of a sample position within the digital training scene 402 that correspond to a cube map (and underlying real spherical harmonic coefficients) generated by the illumination estimation system 110.

[0072] After having identified the local position training indicators 404, the illumination estimation system 110 uses the local position training indicators 404 to modify the global feature training map 410. In some cases, for example, the illumination estimation system 110 uses the local position training indicators 404 to mask the global feature training map 410. For example, the illumination estimation system 110 optionally generates a masked feature training map based on the local position training indicators 404, such as by applying a vector encoder to the local position training indicators 404 (e.g., by one-hot encoding). In some implementations, the masked feature training map includes an array of values ​​that indicates the local position training indicators 404 for specified locations within the digital training scene 402, such as one or more values ​​of a number indicating coordinates of the specified locations within the digital training scene 402, and other values ​​(e.g., the number zero) indicating coordinates of other locations within the digital training scene 402.

[0073] like Figure 4A As further shown, the lighting estimation system 110 multiplies the global feature training map 410 and the masked feature training map of the local position training indicator 404 to generate a masked dense feature training map 412. Because the masked feature training map and the global feature training map 410 may include the same spatial resolution, in some embodiments, the masked feature training map effectively masks the global feature training map 410 to create the masked dense feature training map 412. The masked dense feature training map 412 accordingly represents the local feature map for the specified location within the digital training scene 402.

[0074] In some embodiments, when generating the masked dense feature training map 412, the illumination estimation system 110 concatenates the global feature training map 410 and the masked dense feature training map 412 to form a combined feature training map 414. For example, the illumination estimation system 110 couples the global feature training map 410 and the masked dense feature training map 412 together to form a dual feature map or a stacked feature map as the combined feature training map 414. Alternatively, in some implementations, the illumination estimation system 110 combines rows of values ​​from the global feature training map 410 with rows of values ​​from the masked dense feature training map 412 to form the combined feature training map 414. However, any suitable concatenation method may be used.

[0075] like Figure 4A As further shown, the lighting estimation system 110 feeds the combined feature training map 414 to a second set of network layers 416 of the local lighting estimation neural network 406. Figure 4A As shown, the second set of network layers 416 includes one or both of network layers 418 and 420, and may include a regressor. In some implementations, for example, the second set of network layers 416 includes a convolutional layer as network layer 418 and a fully connected layer as network layer 420. The second set of network layers 416 optionally includes a maximum pooling layer between the convolutional layers constituting network layer 418 and a dropout layer before the fully connected layer constituting network layer 420. As part of a training iteration with such layers, the lighting estimation system 110 optionally passes the combined feature training map 414 through several convolutional layers and several fully connected layers with batch normalization and exponential learning unit (“ELU”) learning functions.

[0076] After passing the combined feature training map 414 through the second set of network layers 416, the local illumination estimation neural network 406 outputs position-specific spherical harmonic training coefficients 422. Consistent with the above disclosure, the position-specific spherical harmonic training coefficients 422 indicate the lighting conditions of a specified position within the digital training scene 402. For example, the position-specific spherical harmonic training coefficients 422 indicate the lighting conditions of a specified position within the digital training scene 402 identified by the local position training indicator 404.

[0077] After generating the position-specific spherical harmonic training coefficients 422, the lighting estimation system 110 compares the position-specific spherical harmonic training coefficients 422 with the base real spherical harmonic coefficients 426. As used in the present disclosure, the term "base real spherical harmonic coefficients" refers to spherical harmonic coefficients that are empirically determined based on one or more cubemaps. For example, the base real spherical harmonic coefficients 426 represent spherical harmonic coefficients projected from a cubemap corresponding to a position within the digital training scene 402 identified by the local position training indicator 404.

[0078] If Figure 4A As further indicated, the illumination estimation system 110 uses a loss function 424 to compare the position-specific spherical harmonic training coefficients 422 with the base true spherical harmonic coefficients 426. In some embodiments, the illumination estimation system 110 uses a mean square error (“MSE”) function as the loss function 424. Alternatively, in some implementations, the illumination estimation system 110 uses an L2 loss function, a mean absolute error function, a mean absolute percentage error function, a drawing loss function, a root mean square error function, or other suitable loss function as the loss function 424.

[0079] After determining the loss from the loss function 424, the lighting estimation system 110 modifies the network parameters (e.g., weights or values) of the local lighting estimation neural network 406 to reduce the loss of the loss function 424 in subsequent training iterations using back propagation indicated by the arrow from the loss function 434 to the local lighting estimation neural network 406. For example, the lighting estimation system 110 may increase or decrease the weights or values ​​from some (or all) of the first set of network layers 408 or the second set of network layers 416 within the local lighting estimation neural network 406 to reduce or minimize the loss in subsequent training iterations.

[0080] After modifying the network parameters of the local illumination estimation neural network 406 for the initial training iteration, the illumination estimation system 110 can perform additional training iterations. In subsequent training iterations, for example, the illumination estimation system 110 extracts additional global feature training maps for additional digital training scenes, generates additional local position training indicators for specified locations within the additional digital training scenes, and modifies the additional global feature training maps based on the additional local position training indicators. Based on the additional combined feature training maps, the illumination estimation system 110 generates additional position-specific spherical harmonic training coefficients for the specified locations.

[0081] The illumination estimation system 110 then modifies the network parameters of the local illumination estimation neural network 406 based on the loss from the loss function 424 to compare the additional location-specific spherical harmonic training coefficients with the additional base true spherical harmonic coefficients for the specified location in the additional digital training scene. In some cases, the illumination estimation system 110 performs training iterations until the values ​​or weights of the local illumination estimation neural network 406 do not change significantly between training iterations or a convergence criterion is satisfied.

[0082] To reach a convergence point, the lighting estimation system 110 is optionally trained using mini-batches of 20 digital training scenes and the Adam optimizer Figure 4A or Figure 5A The local illumination estimation neural network 406 is shown (the latter is described below), with the optimizer minimizing the loss during training iterations with a learning rate of 0.0001 and a weight decay of 0.0002. During some such training experiments, the illumination estimation system 110 reaches a convergence point more quickly when position-specific spherical harmonic training coefficients are generated with a degree greater than zero (e.g., five).

[0083] The lighting estimation system 110 also uses the trained local lighting estimation neural network to generate location specific lighting parameters. Figure 4B An example of such an application is depicted in FIG. Figure 4BAs shown, the lighting estimation system 110 identifies a request to draw a virtual object 432 at a specified location in the digital scene 428. The lighting estimation system 110 uses the local lighting estimation neural network 406 to analyze both the digital scene 428 and the local position indicator 430 to generate a position-specific spherical harmonic coefficient 440 for the specified location. Based on the request, the lighting estimation system 110 draws a modified digital scene 442 that includes the virtual object 432 illuminated at the specified location according to the position-specific spherical harmonic coefficient 440.

[0084] As just noted, lighting estimation system 110 identifies a request to draw virtual object 432 at a specified location within digital scene 428. For example, lighting estimation system 110 may identify a digital request from a computing device executing an augmented reality application to draw a virtual head accessory (or other virtual item) at a specific location on a person (or another real item) depicted in digital scene 428. Figure 4B As shown, the request to draw the virtual object 432 includes a local position indicator 430 for specifying a location at which to draw the virtual object 432 .

[0085] Based on the received Figure 4B At the request shown, the lighting estimation system 110 inputs the digital scene 428 into the local lighting estimation neural network 406. In particular, the lighting estimation system 110 feeds the digital scene 428 to the first set of network layers 408 for the first set of network layers 408 to extract a global feature map 434 from the digital scene 428. As described above, the first set of network layers 408 optionally includes layers from DenseNet and outputs a dense feature map as the global feature map 434. Alternatively, the first set of network layers 408 optionally includes an encoder and outputs an encoded feature map as the global feature map 434.

[0086] like Figure 4B As further shown, the lighting estimation system 110 generates a local position indicator 430 for a specified position within the digital scene 428. For example, in some embodiments, the lighting estimation system 110 identifies the local position indicator 430 from a request to draw a virtual object 432 (e.g., using the local position indicator as a two-dimensional or three-dimensional coordinate in the digital scene 428). The lighting estimation system 110 then modifies the global feature map 434 based on the local position indicator 430.

[0087] To make such modifications, the illumination estimation system 110 can mask the global feature map 434 using the local position indicators 430. In some implementations, for example, the illumination estimation system 110 generates the masked feature map from the local position indicators 430, such as by applying a vector encoder to the local position indicators 430 (e.g., by one-hot encoding). As described above, the masked feature map can include an array of values ​​(e.g., 1s and 0s) of the local position indicators 430 that indicate specified locations within the digital scene 428.

[0088] like Figure 4B As further indicated, the lighting estimation system 110 multiplies the global feature map 434 and the masked feature map for the local position indicator 430 to generate a masked dense feature map 436. Assuming that the masked feature map and the global feature map 434 include the same spatial resolution, in some embodiments, the masked feature map effectively masks the global feature map 434 to create the masked dense feature map 436. The masked dense feature map 436 accordingly represents the local feature map for the specified location within the digital scene 428.

[0089] In some embodiments, when generating the masked dense feature map 436, the lighting estimation system 110 concatenates the global feature map 434 and the masked dense feature map 436 to form a combined feature map 438. To form the combined feature map 438, the lighting estimation system 110 may use any of the concatenation methods described above. The lighting estimation system 110 then feeds the combined feature map 438 to the second set of network layers 416.

[0090] By passing the combined feature map 438 through the second set of network layers 416, the local illumination estimation neural network 406 outputs position-specific spherical harmonic coefficients 440. Consistent with the above disclosure, the position-specific spherical harmonic coefficients 440 indicate the lighting conditions of a specified location within the digital scene 428, such as the specified location identified by the local location indicator 430.

[0091] After generating such lighting parameters, the lighting estimation system 110 renders the modified digital scene 442 including the virtual object 432, which is illuminated at the specified location according to the location-specific spherical harmonic coefficients 440. For example, in some embodiments, the lighting estimation system 110 overlays or otherwise integrates a computer-generated image of the virtual object 432 within the digital scene 428. As part of the rendering, the lighting estimation system 110 selects and renders pixels of the virtual object 432 that reflect the lighting, shading, or appropriate hue indicated by the location-specific spherical harmonic coefficients 440.

[0092] As mentioned above, Figure 5A and Figure 5BAnother embodiment of the lighting estimation system 110 is depicted, which trains and applies a local lighting estimation neural network to generate location-specific lighting parameters. Figure 5A As shown, for example, the illumination estimation system 110 iteratively trains the local illumination estimation neural network 504. As an overview of the training iterations, the illumination estimation system 110 extracts a global feature training map from the digital training scene using a first set of network layers 506 of the local illumination estimation neural network 504. The illumination estimation system 110 also generates a plurality of local position training indicators for specified locations within the digital training scene based on the feature maps output by each layer of the first set of network layers. The illumination estimation system 110 then modifies the global feature training map based on the local position training indicators.

[0093] Based on the combined feature training map obtained from modifying the global feature training map, the illumination estimation system 110 generates position-specific spherical harmonic coefficients for the specified position using the second set of network layers 516 of the local illumination estimation neural network 504. The illumination estimation system 110 then modifies the network parameters of the local illumination estimation neural network 504 based on a comparison of the position-specific spherical harmonic coefficients for the specified position within the digital training scene with the underlying true spherical harmonic coefficients.

[0094] For example, Figure 5A As shown, the illumination estimation system 110 feeds the digital training scene 502 to the local illumination estimation neural network 504. After receiving the digital training scene 504 as training input, the first set of network layers 506 extracts a global feature training map 508 from the digital training scene 502. In some embodiments, the first set of network layers 506 includes layers from DenseNet and outputs the global feature training map 508 in the form of a dense feature map, such as the output of various dense blocks described by Huang. In some such implementations, the first set of network layers 506 also includes several convolutional layers and maximum pooling layers, which reduce the resolution of the global feature training map 508. In an alternative to the DenseNet layers, the first set of network layers 506 may include an encoder and output the global feature training map 508 in the form of an encoded feature map, such as the output of the encoder described by Gardner. Regardless of its form, the global feature training map 508 can represent visual features of the digital training scene 502, such as the location of the light source and the global geometry of the digital training scene 502.

[0095] In addition to generating the global feature training map 508, the lighting estimation system 110 also identifies feature training maps from each of the individual layers of the first set of network layers 506. Figure 5AAs shown, for example, the various layers from the first set of network layers 506 collectively output a feature training map 510 that includes each such feature training map. The lighting estimation system 110 selects training pixels corresponding to a specified location within the digital training scene 502 from each feature training map of the feature training maps 510. For example, the lighting estimation system 110 optionally selects the coordinates of each training pixel corresponding to the specified location from each feature training map. Each training pixel (or corresponding coordinates) represents a local position training indicator for the specified location.

[0096] like Figure 5A As further shown, the lighting estimation system 110 generates a supercolumn training map 512 based on the selected training pixels from the feature training map 510. For example, in some embodiments, the lighting estimation system 110 combines or concatenates each selected training pixel, such as by concatenating the features of each selected pixel, to form the supercolumn training map 512. Thus, the supercolumn training map 512 represents the local features of the specified location within the digital training scene 502. In some such embodiments, the lighting estimation system 110 generates a hypercolumn map, as described in Aayush Bansal et al., “Marr Revisited: 2D-3D Model Alignment via Surface Normal Prediction,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2016), or in Bharath Hariharan et al., “Hypercolumns for Object Segmentation and Fine-Grained Localization,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2015), both of which are incorporated herein by reference in their entirety.

[0097] In some embodiments, when generating the hypercolumn training map 512, the lighting estimation system 110 concatenates the global feature training map 508 and the hypercolumn training map 512 to form a combined feature training map 514. To form the combined feature training map 514, the lighting estimation system 110 optionally (i) couples the global feature training map 508 and the hypercolumn training map 512 together to form a dual or stacked feature map, or (ii) combines rows of values ​​from the global feature training map 508 with rows of values ​​from the hypercolumn training map 512. However, any suitable concatenation method may be used.

[0098] like Figure 5AAs further shown, the lighting estimation system 110 feeds the combined feature training map 514 to a second set of network layers 516 of the local lighting estimation neural network 504. In some implementations, for example, the second set of network layers 516 includes fully connected layers. As part of a training iteration with such layers, the lighting estimation system 110 optionally passes the combined feature training map 514 through several fully connected layers with batch normalization and ELU learning functions.

[0099] After passing the combined feature training map 514 through the second set of network layers 516, the local illumination estimation neural network 504 outputs position-specific spherical harmonic training coefficients 518. Consistent with the above disclosure, the position-specific spherical harmonic training coefficients 518 indicate the illumination conditions of a specified position within the digital training scene 502. For example, the position-specific spherical harmonic training coefficients 518 indicate the illumination conditions of a specified position within the digital training scene 502 identified by the local position training indicator from the feature training map 510.

[0100] After generating the position-specific spherical harmonic training coefficients 518, the illumination estimation system 110 compares the position-specific spherical harmonic training coefficients 518 with the base real spherical harmonic coefficients 522. The base real spherical harmonic coefficients 522 represent the spherical harmonic coefficients projected from the cube map corresponding to the sample position in the digital training scene 502, i.e., the same specified position indicated by the local position training indicator from the feature training map 510.

[0101] like Figure 5A As further shown, the illumination estimation system 110 uses a loss function 520 to compare the position-specific spherical harmonic training coefficients 518 with the base true spherical harmonic coefficients 522. In some embodiments, the illumination estimation system 110 uses an MSE function as the loss function 520. Alternatively, in some implementations, the illumination estimation system 110 uses an L2 loss function, a mean absolute error function, a mean absolute percentage error function, a drawing loss function, a root mean square error function, or other suitable loss functions as the loss function 520.

[0102] After determining the loss from the loss function 520, the lighting estimation system 110 modifies network parameters (e.g., weights or values) of the local lighting estimation neural network 504 to reduce the loss of the loss function 520 in subsequent training iterations. For example, the lighting estimation system 110 may increase or decrease weights or values ​​from some (or all) of the first set of network layers 506 or the second set of network layers 516 within the local lighting estimation neural network 504 to reduce or minimize the loss in subsequent training iterations.

[0103] After modifying the network parameters of the local illumination estimation neural network 504 for the initial training iteration, the illumination estimation system 110 can perform additional training iterations. In subsequent training iterations, for example, the illumination estimation system 110 extracts additional global feature training maps for additional digital training scenes, generates additional local position training indicators for specified locations within the additional digital training scenes, and modifies the additional global feature training maps based on the additional local position training indicators. Based on the additional combined feature training maps, the illumination estimation system 110 generates additional position-specific spherical harmonic training coefficients for the specified locations.

[0104] Illumination estimation system 110 then modifies network parameters of local illumination estimation neural network 504 based on a loss from loss function 520 that compares the additional location-specific spherical harmonic training coefficients to the additional basis real spherical harmonic coefficients for the specified location within the additional digital training scene. In some cases, illumination estimation system 110 performs training iterations until the values ​​or weights of local illumination estimation neural network 504 do not change significantly during the training iterations or a convergence criterion is satisfied.

[0105] Figure 5B An example of using a local illumination estimation neural network 504 to generate location-specific illumination parameters is depicted. In general, as Figure 5B As shown, the lighting estimation system 110 identifies a request to draw a virtual object 526 at a specified location within the digital scene 524. The lighting estimation system 110 uses the local lighting estimation neural network 504 to analyze both the digital scene 524 and the plurality of local position indicators to generate position-specific spherical harmonic coefficients 536 for the specified location. Based on the request, the lighting estimation system 110 draws a modified digital scene 538 that includes the virtual object 526 illuminated at the specified location according to the position-specific spherical harmonic coefficients 536.

[0106] As just noted, the lighting estimation system 110 identifies a request to draw a virtual object 526 at a specified location within the digital scene 524. For example, the lighting estimation system 110 may identify a digital request from a computing device to draw a virtual animal or character (or other virtual item) at a specific location on a landscape (or another real item) depicted in the digital scene 524. Figure 5B Not shown, but the request to draw virtual object 526 includes an indication of a specified location within digital scene 524 at which to draw virtual object 526.

[0107] Based on the received Figure 5BAt the request shown, the lighting estimation system 110 inputs the digital scene 524 into the local lighting estimation neural network 504. In particular, the lighting estimation system 110 feeds the digital scene 524 to the first set of network layers 506 to extract a global feature map 528 from the digital scene 524. As described above, the first set of network layers 506 optionally includes layers from DenseNet and outputs a dense feature map as the global feature map 528. Alternatively, the first set of network layers 506 optionally includes an encoder and outputs an encoded feature map as the global feature map 528.

[0108] like Figure 5B As further shown, the lighting estimation system 110 identifies a feature map from each of the layers of the first set of network layers 506. Figure 5B As shown, for example, the various layers from the first set of network layers 506 collectively output feature maps 530, including each such feature map. The lighting estimation system 110 selects pixels corresponding to a specified location within the digital scene 524 from each feature map of the feature maps 510. For example, the lighting estimation system 110 selects coordinates for each pixel corresponding to the specified location from each feature map.

[0109] like Figure 5B As further shown, the lighting estimation system 110 generates a hypermap 532 based on the pixels selected from the feature map 530. For example, in some embodiments, the lighting estimation system 110 combines or concatenates each selected pixel, such as by concatenating the features from each selected pixel, to form the hypermap 532. The hypermap 532 accordingly represents the local features of the specified location within the digital scene 524.

[0110] In some embodiments, when generating the hypercolumn graph 532, the lighting estimation system 110 concatenates the global feature map 528 and the hypercolumn graph 532 to form a combined feature map 534. To form the combined feature map 534, the lighting estimation system 110 may use any of the concatenation methods described above. The lighting estimation system 110 then feeds the combined feature map 534 to the second set of network layers 516.

[0111] By passing the combined feature map 534 through the second set of network layers 516, the local illumination estimation neural network 504 outputs position-specific spherical harmonic coefficients 536. Consistent with the above disclosure, the position-specific spherical harmonic coefficients 536 indicate the lighting conditions at a specified location within the digital scene 524, such as the specified location identified by the local location indicator from the feature map 530.

[0112] After generating such lighting parameters, the lighting estimation system 110 renders the modified digital scene 538 including the virtual object 526, which is illuminated at the specified location according to the location-specific spherical harmonic coefficients 536. For example, in some embodiments, the lighting estimation system 110 overlays or otherwise integrates a computer-generated image of the virtual object 526 within the digital scene 524. As part of the rendering, the lighting estimation system 110 selects and renders pixels of the virtual object 526 that reflect the lighting, shadows, or appropriate hue indicated by the location-specific spherical harmonic coefficients 536.

[0113] In addition to accurately depicting the lighting conditions at a given location, Figure 4B and Figure 5B The position-specific spherical harmonic coefficients generated in can dynamically capture the lighting from different perspectives of a specified location within a digital scene. In some embodiments, when the perspective of the digital scene changes in camera viewpoint, model orientation, or other perspective adjustments, the position-specific spherical harmonic coefficients can accurately indicate the lighting conditions at the specified location despite such changes in perspective. For example, in some embodiments, the lighting estimation system 110 identifies a perspective adjustment request to draw the digital scene from a different viewpoint, such as by detecting movement of a mobile device that redirects the digital scene or identifying user input that modifies the perspective of the digital scene (e.g., camera movement that adjusts the perspective). Based on the perspective adjustment request, the lighting estimation system 110 can draw a modified digital scene from a different viewpoint, the scene including the scene according to Figure 4B or Figure 5B The position-specific spherical harmonic coefficients generated in the virtual object that is illuminated at the specified position.

[0114] Alternatively, in some implementations, the lighting estimation system 110 adjusts or generates new location-specific lighting parameters in response to a perspective adjustment of the digital scene and a corresponding viewpoint change (e.g., a camera movement to adjust the perspective). For example, in some embodiments, the lighting estimation system 110 identifies a perspective adjustment request to draw a virtual object at a specified location within the digital scene from a new or different viewpoint. Based on such a perspective adjustment request, the lighting estimation system 110 may generate the above-described Figure 4B or Figure 5B New consistent location specific lighting parameters.

[0115] In some cases, for example, the lighting estimation system 110 generates a new local position indicator for a specified position within the digital scene from a different viewpoint (e.g., Figure 4B The new coordinates for the newly specified location are as shown, or as Figure 5BThe illumination estimation system 110 then modifies the global feature map of the digital scene based on the new local position indicators for the designated positions from the different viewpoints to form a new modified global feature map (e.g., concatenating the new global feature map and the new masked dense feature map to form a new combined feature map, such as Figure 4B As shown, or concatenate the new global feature map and the new hypercolumn map to form a new combined feature map, such as Figure 5B Based on the new modified global feature map, the lighting estimation system 110 uses a second set of network layers (e.g., Figure 4B The second network layer 416 or Figure 5B The second set of network layers 516 in the ) generates new location-specific lighting parameters for the specified location from different viewpoints. In response to the viewpoint adjustment request, the lighting estimation system 110 accordingly draws the adjusted digital scene, which includes virtual objects illuminated at the specified location according to the new location-specific lighting parameters.

[0116] Apart from Figure 4B and Figure 5B In addition to the position-specific spherical harmonic coefficients and modified digital scene shown in , the lighting estimation system 110 can generate new position-specific spherical harmonic coefficients and adjusted digital scene in response to the position adjustment request to draw the virtual object at the new specified position. Figure 4B The local illumination estimation neural network 406 of the illumination estimation system 110 can identify a new local position indicator from the position adjustment request and modify the global feature map based on the new local position indicator to generate a new masked dense feature map. Figure 5B The local illumination estimation neural network 504 of the present invention may be used to identify a plurality of new local position indicators from the feature maps output by the first set of network layers 506 and generate a hypercolumn graph based on pixels selected from the feature maps.

[0117] Use from Figure 4B or Figure 5B The lighting estimation system 110 can generate a new combined feature map and use a second set of network layers to generate new position-specific spherical harmonic coefficients based on the neural network architecture. Based on the position adjustment request, the lighting estimation system 110 can thus draw an adjusted digital scene including a virtual object, which is illuminated at the new specified position according to the new position-specific spherical harmonic coefficients. In response to the position adjustment request, the lighting estimation system 110 can accordingly adapt the lighting conditions to the different position to which the computing device moves the virtual object.

[0118] In addition to updating the position-specific spherical harmonic coefficients and the digital scene in response to a position adjustment, the lighting estimation system 110 can generate new position-specific spherical harmonic coefficients and an adjusted scene in response to a change or adjustment in lighting conditions, movement of objects in the scene, or other changes to the scene. Figure 4B The lighting estimation system 110 can extract a new global feature map from the scene and modify the new global feature map based on the local position indicators to generate a new masked dense feature map. Figure 5B By using the first set of network layers 506 , the lighting estimation system 110 may generate a new global feature map and a new hypercolumn map from pixels selected from the new global feature map.

[0119] Use from Figure 4B or Figure 5B The lighting estimation system 110 can generate a new combined feature map and use a second set of network layers to generate new position-specific spherical harmonic coefficients that reflect the adjusted lighting conditions. Therefore, based on the adjustment of the lighting conditions for the specified location within the digital scene, the lighting estimation system 110 can draw the adjusted digital scene including the virtual object, which is illuminated at the specified location according to the new position-specific spherical harmonic coefficients. In response to the adjustment of the lighting conditions, the lighting estimation system 110 can adapt the position-specific lighting parameters accordingly to reflect the adjustment of the lighting conditions.

[0120] Figure 6A-6C An embodiment of a lighting estimation system 110 is depicted in which a computing device, in response to different rendering requests, renders a digital scene including a virtual object at a specified location according to location-specific lighting parameters and new location-specific lighting parameters, respectively, and then renders a digital scene including a virtual object at the new specified location. As an overview, Figure 6A-6C Each depicts a computing device 600 including an augmented reality application for the augmented reality system 108 and the lighting estimation system 110. The augmented reality application includes causing the computing device 600 to execute Figure 6A-6C Computer executable instructions for certain actions depicted in .

[0121] Rather than repeatedly describing computer executable instructions in an augmented reality application as causing the computing device 600 to perform such actions, the present disclosure primarily describes the computing device 600 or the lighting estimation system 110 as shorthand for performing an action. Figure 6A-6C Various user interactions indicated, such as when computing device 600 detects user selection of a virtual object. Figure 6A-6C6 is shown as a mobile device (e.g., a smartphone), but computing device 600 may alternatively be any type of computing device, such as a desktop, laptop, or tablet, and may also detect any appropriate user interaction, including but not limited to audio input from a microphone, gaming device button input, keyboard input, mouse clicks, stylus interaction with a touch screen, or touch gestures on a touch screen.

[0122] Now back to Fig. 6A , which depicts a computing device 600 presenting a graphical user interface 606a that includes a digital scene 610 within a screen 602. As shown, the digital scene 610 includes real objects 612a and 612b. The graphical user interface 606a also includes a selectable option bar 604 for virtual objects 608a-608c. By presenting the graphical user interface 606a, the computing device 600 provides the user with an option to request the lighting estimation system 110 to draw one or more virtual objects 608a-608c at a specified location within the digital scene 610.

[0123] For example, Fig. 6A As shown, the computing device 600 detects a user interaction requesting the lighting estimation system 110 to draw a virtual object 608a at a specified location 614 within the digital scene 610. In particular, Fig. 6A The computing device 600 is depicted detecting a drag-and-drop gesture to move the virtual object 608a to a specified location 614. Fig. 6A A user using a drag-and-drop gesture is illustrated, but computing device 600 may detect any suitable user interaction requesting lighting estimation system 110 to draw a virtual object within a digital scene.

[0124] Based on receiving the request for the lighting estimation system 110 to draw the virtual object 608 a within the digital scene 610 , the augmented reality system 108 , in conjunction with the lighting estimation system 110 , draws the virtual object 608 a at the specified location 614 . Figure 6B An example of such a rendering is depicted. As shown, computing device 600 presents a graphical user interface 606b including a modified digital scene 616 within screen 602. Consistent with the above disclosure, computing device 600 renders a modified digital scene 616 including a virtual object 608a that is illuminated at a specified location 614 according to location-specific lighting parameters generated by lighting estimation system 110.

[0125] To generate such location-specific lighting parameters, the lighting estimation system 110 optionally performs Figure 4B or Figure 5B The action shown. Figure 6B, the location-specific lighting parameters indicate realistic lighting conditions for the virtual object 608a, where the lighting and shadows are consistent with the real objects 612a and 612b. The shadows of the virtual object 608a and the real objects 612a and 612b always reflect light from light sources outside the perspective shown in the modified digital scene 616.

[0126] As described above, the lighting estimation system 110 may generate new location-specific lighting parameters and an adjusted digital scene in response to a location adjustment request to render a virtual object at the new specified location. Figure 6C An example of such an adjusted digital scene is depicted, which reflects the new location-specific lighting parameters. Figure 6C As shown, computing device 600 presents graphical user interface 606c including adjusted digital scene 618 within screen 602. Computing device 600 detects a user interaction including a position adjustment request to move virtual object 608a from designated position 614 to a new designated position 620.

[0127] Based on receiving the request from the lighting estimation system 110 to move the virtual object 608a, the augmented reality system 108, together with the lighting estimation system 110, draws the virtual object 608a at the new designated position 620. Figure 6C The computing device 600 is depicted rendering an adjusted digital scene 618 including a virtual object 608 a illuminated at a new designated location 620 according to new location-specific lighting parameters generated by the lighting estimation system 110 .

[0128] To generate such new location-specific lighting parameters, the lighting estimation system 110 optionally modifies the global feature map and uses a lighting estimation neural network to generate the following: Figure 4B or Figure 5B The position-specific spherical harmonic coefficients are shown in Figure 6C As shown, the location-specific lighting parameters indicate realistic lighting conditions for the virtual object 608a, where the virtual object 608a has adjusted lighting and shadows consistent with the real objects 612a and 612b. FIG. 6B to FIG. 6C As shown in the transition of , the lighting estimation system 110 can adapt the lighting conditions to different positions in real time (or near real time) in response to a position adjustment request to move a virtual object in the virtual scene.

[0129] As described above, the lighting estimation system 110 can generate location-specific lighting parameters that indicate accurate and realistic lighting conditions for locations within a digital scene. To test the accuracy and realism of the lighting estimation system 110, the researchers modified digital scenes from the SUNCG dataset (described above) and applied a local lighting estimation neural network to generate location-specific lighting parameters for various locations in such digital scenes. Fig. 7A and Figure 7B Examples of such accuracy and realism in various renderings of a digital scene with locations illuminated according to base real lighting parameters from the lighting estimation system 110 and locations illuminated according to location specific lighting parameters from the lighting estimation system 110 are illustrated.

[0130] for Fig. 7A and Figure 7B In both cases, the researchers trained Figure 4A A local illumination estimation neural network is depicted, wherein the local illumination estimation neural network includes DenseNet blocks from the first set of network layers of DenseNet 120 and initializes network parameters using weights trained on ImageNet. Fig. 7A and Figure 7B In both cases, the researchers based Figure 4B The actions shown further generate position-specific spherical harmonic coefficients by applying a trained local illumination estimation neural network.

[0131] For example, Fig. 7A As shown, the researchers modified a digital scene 702 from the SUNCG dataset. The digital scene 702 includes a designated location 704 indicating a designated location that serves as a target for estimating lighting conditions. For comparison purposes, the researchers plotted a red-green-blue ("RBG") representation 706 and a light intensity representation 710 of the designated location based on the underlying real spherical harmonic coefficients. Consistent with the disclosure above, the researchers projected the underlying real spherical harmonic coefficients of the designated location 704 from a cubic map. The researchers further used the lighting estimation system 110 to generate position-specific spherical harmonic coefficients for the designated location 704. After generating such lighting parameters, the augmented reality system 108, together with the lighting estimation system 110, plotted a RBG representation 708 and a light intensity representation 712 of the designated location 704 based on the position-specific spherical harmonic coefficients.

[0132] As a pair Fig. 7A As indicated by the comparison of the RGB representation and the light intensity representation shown, the lighting estimation system 110 generates location-specific spherical harmonic coefficients that accurately and realistically estimate the lighting conditions of an object located at a specified location 704. Unlike some conventional augmented reality systems, the lighting estimation system 110 accurately estimates the lighting emitted from light sources outside the viewpoint of the digital scene 702. Although the strongest light source is from behind the camera that captured the digital scene 702, the lighting estimation system 110 estimates the lighting conditions shown in the RBG representation 708 and the light intensity representation 712 with similar accuracy and realism to those shown in the RBG representation 706 and the light intensity representation 710 that reflect the ground truth.

[0133] Figure 7B A modified digital scene 714 is illustrated, the scene including virtual objects 718a-718d at specified locations illuminated according to the underlying real spherical harmonic coefficients. Consistent with the above disclosure, the researchers projected the underlying real spherical harmonic coefficients from the cube map for the specified locations in the modified digital scene 714. Figure 7B Further illustrated is a modified digital scene 716 that includes virtual objects 720a-720d illuminated at specified locations according to location-specific spherical harmonic coefficients generated by the lighting estimation system 110. To determine estimated lighting conditions for the virtual objects 720a-720d at the specified locations, the lighting estimation system 110 generates location-specific spherical harmonic coefficients for each specified location in the modified digital scene 716.

[0134] For comparison purposes, the researchers used the lighting estimation system 110 to draw the metal spheres for the virtual objects 718a-718d in the modified digital scene 714 and the virtual objects 720a-720d in the modified digital scene 716. Figure 7B As shown, both modified digital scene 714 and modified digital scene 716 include virtual objects 718a-718d and virtual objects 720a-720d, respectively, at the same designated locations.

[0135] As indicated by a comparison of the illumination of the virtual objects in the modified digital scenes 714 and 716, the illumination estimation system 110 generates position-specific spherical harmonic coefficients that accurately and realistically estimate the illumination conditions of the virtual objects 720a-720d at the respective specified locations of each object in the modified digital scene 716. Although the light intensities of the virtual objects 720a-720d differ slightly from the light intensities of the virtual objects 718a-720d, the trained local illumination estimation neural network detects sufficient geometric context from the underlying scene of the modified digital scene 716 to generate coefficients that both (i) darken the occluded metal sphere and (ii) reflect strong directional light on the metal sphere when exposed to light from sources outside the viewing angle of the modified digital scene 716.

[0136] Now go to Figure 8 and Fig. 9 ,These figures provide an overview of the environments in which lighting estimation systems can operate and architectural examples of lighting estimation systems. In particular, Figure 8 A block diagram illustrating an exemplary system environment ("environment") 800 in which a lighting estimation system 806 may operate is depicted according to one or more embodiments. Specifically, Figure 8An environment 800 is illustrated, which includes server(s) 802, third party server(s) 810, a network 812, a client device 814, and a user 818 associated with the client device 814. Figure 8 One client device and one user are illustrated, but in alternative embodiments, environment 800 may include any number of computing devices and associated users. Figure 8 A particular arrangement of server(s) 802, third party server(s) 810, network 812, client devices 814, and users 818 is illustrated, but various additional arrangements are possible.

[0137] like Figure 8 As shown, the server(s) 802, the third party server(s) 810, the network 812, and the client device 814 may be directly or indirectly communicatively coupled to each other, such as via the network 812, which will be discussed below with respect to Fig.12 The server(s) 802 and client device 814 may include any type of computing device, including one or more computing devices, as described below with respect to Fig.12 Further discussion.

[0138] like Figure 8 As shown, the server(s) 802 may generate, store, receive and / or transmit any type of data, including user input to input a digital scene into a neural network or requesting drawing of a virtual object to create an augmented reality scene. For example, the server(s) 802 may receive user input from a client device 814 requesting drawing of a virtual object at a specified location within a digital scene, and then utilize a local illumination estimation neural network to generate location-specific lighting parameters for the specified location. After generating such parameters, the server(s) 802 may further draw a modified digital scene including a virtual object, the virtual object being illuminated at the specified location according to the location-specific lighting parameters. In some embodiments, the server(s) 802 include a data server, a communication server, or a web hosting server.

[0139] like Figure 8As further shown, the server(s) 802 may include an augmented reality system 804. In general, the augmented reality system 804 facilitates the generation, modification, sharing, access, storage, and / or deletion of digital content (e.g., a two-dimensional digital image of a scene or a three-dimensional digital model of a scene) in an augmented reality-based image. For example, the augmented reality system 804 may use the server(s) 802 to generate a modified digital image or model including a virtual object or to modify an existing digital scene. In some implementations, the augmented reality system 804 uses the server(s) 802 to receive user input identifying a digital scene, a virtual object, or a specified location within the digital scene from a client device 814, or transmit data representing the digital scene, the virtual object, or the specified location to the client device 814.

[0140] In addition to the augmented reality system 804, the server(s) 802 also include an illumination estimation system 806. The illumination estimation system 806 is an embodiment of the illumination estimation system 110 described above (and can perform functions, methods, and processes). For example, in some embodiments, the illumination estimation system 806 uses the server(s) 802 to identify a request to draw a virtual object at a specified location within a digital scene. The illumination estimation system 806 further uses the server(s) 802 to extract a global feature map from the digital scene using a first set of network layers of a local illumination estimation neural network. In some implementations, the illumination estimation system 806 also uses the server(s) 802 to generate a local position indicator for the specified location and modify the global feature map of the digital scene based on the local position indicator. Based on the modified global feature map, the illumination estimation system 806 further uses the server(s) 802 to (i) generate location-specific illumination parameters for the specified location using a second set of layers of the local illumination estimation neural network, and (ii) draw the modified digital scene, which includes a virtual object illuminated at the specified location according to the location-specific illumination parameters.

[0141] As suggested by the previous embodiments, lighting estimation system 806 may be implemented in whole or in part by various elements of environment 800. Figure 8 The lighting estimation system 806 is illustrated as being implemented within the server(s) 802, but components of the lighting estimation system 806 may be implemented in other components of the environment 800. For example, in some embodiments, the client device 814 includes the lighting estimation system 806 and performs all of the functions, methods, and processes of the lighting estimation system 806 described above and below. Fig. 9 Components of the lighting estimation system 806 are further described.

[0142] like Figure 8As further shown, in some embodiments, the client device 814 includes a computing device that allows a user 818 to send and receive digital communications. For example, the client device 814 may include a desktop computer, a laptop computer, a smart phone, a tablet computer, or other electronic device. In some embodiments, the client device 814 also includes one or more software applications (e.g., an augmented reality application 816) that allow the user 818 to send and receive digital communications. For example, the augmented reality application 816 can be a software application installed on the client device 814, or a software application hosted on (multiple) servers 802. When hosted on (multiple) servers 802, the augmented reality application 816 can be accessed by the client device 814 through another application, such as a web browser. In some implementations, the augmented reality application 816 includes instructions that, when executed by a processor, cause the client device 814 to present one or more graphical user interfaces, such as a user interface that includes digital scenes and / or virtual objects, which are used for the user 818 to select as input when generating location-specific lighting parameters or modified digital scenes, or for the lighting selection system 806 to include as input.

[0143] Also like Figure 8 As shown, augmented reality system 804 is communicatively coupled to augmented reality database 808. In one or more embodiments, augmented reality system 804 accesses and queries data from augmented reality database 808 associated with a request from lighting estimation system 806. For example, augmented reality system 804 can access a digital scene, a virtual object, a specified location within a digital scene, or location-specific lighting parameters of lighting estimation system 806. Figure 8 As shown, augmented reality database 808 is maintained separately from server(s) 802. Alternatively, in one or more embodiments, augmented reality system 804 and augmented reality database 808 comprise a single combined system or subsystem within server(s) 802.

[0144] Now go to Fig. 9 , which provides additional details about the components and features of the lighting estimation system 806. In particular, Fig. 9 A computing device 900 is illustrated that implements an augmented reality system 804 and a lighting estimation system 806. In some embodiments, the computing device 900 includes one or more servers (e.g., server(s) 802). In other embodiments, the computing device 900 includes one or more client devices (e.g., client device 814).

[0145] like Fig. 9As shown, computing device 900 includes augmented reality system 804. In some embodiments, augmented reality system 804 uses its components to provide tools for generating digital scenes or other augmented reality-based images or modifying existing digital scenes or other augmented reality-based images in a user interface of augmented reality application 816. Additionally, in some cases, augmented reality system 804 facilitates the generation, modification, sharing, access, storage, and / or deletion of digital content in augmented reality-based images.

[0146] like Fig. 9 As further shown, computing device 900 includes lighting estimation system 806. Lighting estimation system 806 includes, but is not limited to, digital scene manager 902, virtual object manager 904, neural network trainer 906, neural network operator 908, augmented reality renderer 910, and / or storage manager 912. The following paragraphs describe each of these components in turn.

[0147] As just mentioned, the lighting estimation system 806 includes a digital scene manager 902. The digital scene manager 902 receives input regarding digital scenes, identifies and analyzes the digital scenes. For example, in some embodiments, the digital scene manager 902 receives user input identifying digital scenes, and presents the digital scenes from an augmented reality application. Additionally, in some embodiments, the digital scene manager 902 identifies multiple digital scenes for presentation as part of an image sequence (e.g., an augmented reality sequence).

[0148] like Fig. 9 As further shown, the virtual object manager 904 receives input regarding virtual objects, identifies and analyzes virtual objects. For example, in some embodiments, the virtual object manager 904 receives user input that identifies a virtual object and requests the lighting estimation system 110 to draw the virtual object at a specified location within the digital scene. Additionally, in some embodiments, the virtual object manager 904 provides selectable options for the virtual object, such as selectable options shown in a user interface of an augmented reality application.

[0149] like Fig. 9As further shown, the neural network trainer 906 trains the local illumination estimation neural network 918. For example, in some embodiments, the neural network trainer 906 extracts a global feature training map from the digital training scene using a first set of network layers of the local illumination estimation neural network 918. Additionally, in some embodiments, the neural network trainer 906 generates a local location training indicator for a specified location within the digital training scene and modifies the global feature training map based on the local location training indicator for the specified location. Based on the modified global feature training map, the neural network trainer 906 (i) generates location-specific illumination training parameters for the specified location using a second set of network layers of the local illumination estimation neural network 918, and (ii) modifies network parameters of the local illumination estimation neural network based on a comparison of the location-specific illumination training parameters for the specified location within the digital training scene with the ground truth illumination parameters.

[0150] In some such embodiments, the neural network trainer 906 trains Figure 4A and Figure 5A Local lighting estimation neural network 918 is shown. In some embodiments, neural network trainer 906 also communicates with storage manager 912 to apply and / or access digital training scenes from digital scenes 914, ground truth lighting parameters from location-specific lighting parameters 920, and / or local lighting estimation neural network 918.

[0151] like Fig. 9 As further shown, the neural network operator 908 applies a trained version of the local illumination estimation neural network 918. For example, in some embodiments, the neural network operator 908 extracts a global feature map from the digital scene using a first set of network layers of the local illumination estimation neural network 918. The neural network operator 908 also generates a local position indicator for the specified location and modifies the global feature map of the digital scene based on the local position indicator. Based on the modified global feature map, the neural network operator 908 also generates location-specific lighting parameters for the specified location using a second set of layers of the local illumination estimation neural network. In some such embodiments, the neural network operator 908 applies the following respectively: Figure 4B and Figure 5B Local lighting estimation neural network 918 is shown. In some embodiments, neural network operator 908 also communicates with storage manager 912 to apply and / or access digital scenes from digital scenes 914, virtual objects from virtual objects 916, location-specific lighting parameters from location-specific lighting parameters 920, and / or local lighting estimation neural network 918.

[0152] In addition to the neural network operator 908, in some embodiments, the lighting estimation system 806 also includes an augmented reality renderer 910. The augmented reality renderer 910 renders a modified digital scene including a virtual object. For example, in some embodiments, based on a request to draw a virtual object at a specified location within the digital scene, the augmented reality renderer 910 draws a modified digital scene including the virtual object at the specified location illuminated according to the location-specific lighting parameters from the neural network operator 908.

[0153] In one or more embodiments, each component of the lighting estimation system 806 communicates with each other using any suitable communication technology. Additionally, the components of the lighting estimation system 806 can communicate with one or more other devices including one or more client devices described above. Fig. 9 Components of the lighting estimation system 806 are shown as separate, but any subcomponents may be combined into fewer components, such as into a single component, or may be divided into more components for a particular implementation. Fig. 9 , but at least some of the components used to perform operations in conjunction with the lighting estimation system 806 described herein can be implemented on other devices within the environment 800.

[0154] Each component 902-920 of the lighting estimation system 806 may include software, hardware, or both. For example, the components 902-920 may include one or more instructions stored on a computer-readable storage medium, and the one or more instructions are executable by a processor of one or more computing devices (such as a client device or a server device). When executed by one or more processors, the computer-executable instructions of the lighting estimation system 806 may cause the (multiple) computing devices to perform the methods described herein. Alternatively, the components 902-920 may include hardware, such as a dedicated processing device that performs a specific function or group of functions. Alternatively, the components 902-920 of the lighting estimation system 806 may include a combination of computer-executable instructions and hardware.

[0155] In addition, the components 902-920 of the lighting estimation system 806 can be implemented, for example, as one or more operating systems, as one or more independent applications, as one or more generators of applications, as one or more plug-ins, as one or more library functions or functions that other applications can call, and / or as a cloud computing model. Therefore, the components 902-920 can be implemented as independent applications, such as desktop or mobile applications. In addition, the components 902-920 can be implemented as one or more web-based applications hosted on a remote server. The components 902-920 can also be implemented in a set of mobile device applications or "apps". For illustration, the components 902-920 can be implemented in software applications, including but not limited to ADOBE ILLUSTRATOR, ADOBE EXPERIENCE DESIGN, ADOBE CREATIVE CLOUD, ADOBE PHOTOSHOP, PROJECT AERO, or ADOBE LIGHTROOM. “ADOBE”, “ILLUSTRATOR”, “EXPERIENCE DESIGN”, “CREATIVE CLOUD”, “PHOTOSHOP”, “PROJECT AERO” and “LIGHTROOM” are registered trademarks or trademarks of Adobe Corporation in the United States and / or other countries.

[0156] Now go to Fig.10 , which illustrates a flow chart of a series of actions 1000 for training a local lighting estimation neural network to generate location-specific lighting parameters in accordance with one or more embodiments. Fig.10 The actions according to one embodiment are illustrated, but alternative embodiments may omit, add, reorder and / or modify Fig.10 Any action shown. Fig.10 The actions of may be performed as part of a method. Alternatively, a non-transitory computer-readable storage medium may include instructions that, when executed by one or more processors, cause a computing device to perform Fig.10 In other embodiments, the system may perform Fig.10 action.

[0157] like Fig.10 As shown, action 1000 includes action 1010, which extracts a global feature training map from a digital training scene using a local illumination estimation neural network. In particular, in some embodiments, action 1010 includes extracting a global feature training map from a digital training scene of the digital training scene using a first set of network layers of the local illumination estimation neural network. In some embodiments, the digital training scene includes a three-dimensional digital model or a digital viewpoint image of a real scene.

[0158] like Fig.10 As further shown, act 1000 includes act 1020 of generating a local position training indicator for a specified position within the digital training scene. For example, in some embodiments, generating the local position training indicator includes identifying local position training coordinates representing the specified position within the digital training scene. In contrast, in some implementations, generating the local position training indicator includes: selecting a first training pixel corresponding to the specified position from a first feature training map corresponding to a first layer of the local illumination estimation neural network; and selecting a second training pixel corresponding to the specified position from a second feature training map corresponding to a second layer of the local illumination estimation neural network.

[0159] like Fig.10 As further shown, action 1000 includes action 1030, which modifies the global feature training map of the digital training scene based on the local position training indicator. For example, in some implementations, action 1030 includes modifying the global feature training map to generate a modified global feature training map by: generating a masked feature training map from the local position training coordinates; multiplying the global feature training map and the masked feature training map for the local position training coordinates to generate a masked dense feature training map; and concatenating the global feature training map and the masked dense feature training map to form a combined feature training map.

[0160] like Fig.10 As further shown in , action 1000 includes action 1040, which generates location-specific lighting training parameters for the specified location using the local lighting estimation neural network based on the modified global feature training map. In particular, in some implementations, action 1040 includes generating location-specific lighting training parameters for the specified location using a second set of network layers of the local lighting estimation neural network based on the modified global feature training map. In some embodiments, the first set of network layers of the local lighting estimation neural network includes lower layers of a densely connected convolutional network, and the second set of network layers of the local lighting estimation neural network includes convolutional layers and fully connected layers.

[0161] As described above, in some implementations, generating location-specific lighting training parameters for a specified location includes generating location-specific spherical harmonic training coefficients that indicate lighting conditions at the specific location. In some such embodiments, generating location-specific spherical harmonic training coefficients includes generating location-specific spherical harmonic training coefficients five times for each color channel.

[0162] like Fig.10As further shown, act 1000 includes act 1050 of modifying network parameters of the local lighting estimation neural network based on the comparison of the location-specific training parameters with the ground truth lighting parameters. In particular, in some embodiments, act 1050 includes modifying network parameters of the local lighting estimation neural network based on the comparison of the location-specific lighting training parameters with a set of ground truth lighting parameters for a specified location within the digital training scene.

[0163] In addition to actions 1010-1050, in some cases, action 1000 also includes determining a set of ground truth lighting parameters for the specified location by determining a set of ground truth location-specific spherical harmonic coefficients indicating lighting conditions at the specified location. Additionally, in one or more embodiments, action 1000 also includes generating location-specific lighting training parameters by providing the combined feature training map to a second set of network layers.

[0164] As described above, in some embodiments, act 1000 also includes determining a set of ground-truth lighting parameters for the specified location by determining a set of ground-truth location-specific spherical harmonic coefficients indicating lighting conditions at the specified location. In some such implementations, determining the set of ground-truth location-specific spherical harmonic coefficients includes: identifying locations within a digital training scene; generating a cube map for each location in the digital training scene; and projecting the cube map for each location in the digital training scene to the set of ground-truth location-specific spherical harmonic coefficients.

[0165] Now go to Fig.11 , which illustrates a flow chart of a series of actions 1100 for applying a trained local illumination estimation neural network to generate location-specific illumination parameters in accordance with one or more embodiments. Fig.11 The actions according to one embodiment are illustrated, but alternative embodiments may omit, add, reorder and / or modify Fig.11 Any action shown. Fig.11 The actions of may be performed as part of a method. Alternatively, a non-transitory computer-readable storage medium may include instructions that, when executed by one or more processors, cause a computing device to perform Fig.11 In other embodiments, the system may perform Fig.11 action.

[0166] like Fig.11 As shown, act 1100 includes an act 1110 of identifying a request to draw a virtual object at a specified location within a digital scene. For example, in some embodiments, identifying the request includes receiving a request to draw a virtual object at a specified location from a mobile device.

[0167] like Fig.11As further shown, act 1100 includes extracting a global feature map from the digital scene using a local illumination estimation neural network 1120. In particular, in some embodiments, act 1120 includes extracting a global feature map from the digital scene using a first set of network layers of the local illumination estimation neural network.

[0168] like Fig.11 As further shown in , act 1100 includes act 1130 of generating a local position indicator for a specified position within a digital scene. In particular, in some embodiments, act 1130 includes generating a local position indicator for a specified position by identifying local position coordinates representing the specified position within the digital scene.

[0169] In contrast, in some implementations, action 1130 includes generating a local position indicator for the specified position by selecting a first pixel corresponding to the specified position from a first feature map corresponding to a first layer of the first set of network layers; and selecting a second pixel corresponding to the specified position from a second feature map corresponding to a second layer of the first set of network layers.

[0170] like Fig.11 As further shown, action 1100 includes action 1140, which modifies the global feature map of the digital scene based on the local position indicator for the specified position. In particular, in some embodiments, action 1140 includes modifying the global feature map to generate a modified global feature map by the following steps: combining the features of the first pixel and the second pixel corresponding to the specified position to generate a hypercolumn map; and concatenating the global feature map and the hypercolumn map to form a combined feature map.

[0171] like Fig.11 As further shown in , action 1100 includes action 1150, which generates location-specific lighting parameters for the specified location using a local illumination estimation neural network based on the modified global feature map. In particular, in some embodiments, action 1140 includes: generating location-specific lighting parameters for the specified location using a second set of network layers of the local illumination estimation neural network based on the modified global feature map. In some embodiments, the first set of network layers of the local illumination estimation neural network includes lower layers of a densely connected convolutional network, and the second set of network layers of the local illumination estimation neural network includes convolutional layers and fully connected layers.

[0172] As an example of action 1150, in some embodiments, generating location-specific lighting parameters for the specified location includes generating location-specific spherical harmonic coefficients indicating lighting conditions of an object at the location-specific location. As another example, in some implementations, generating location-specific lighting parameters includes providing the combined feature map to the second set of network layers.

[0173] like Fig.11 As further shown, action 1100 includes action 1160, which draws a modified digital scene including a virtual object, the virtual object being illuminated at a specified location according to the location-specific lighting parameters. In particular, in some embodiments, action 1160 includes drawing a modified digital scene including a virtual object, the virtual object being illuminated at the specified location according to the location-specific lighting parameters, based on the request. For example, in some cases, drawing the modified digital scene includes: based on receiving the request from the mobile device, drawing the modified digital scene including the virtual object, the virtual object being illuminated at the specified location according to the location-specific lighting parameters, within a graphical user interface of the mobile device.

[0174] In addition to actions 1110-1160, in some implementations, action 1100 also includes identifying a position adjustment request to move the virtual object from a specified position within the digital scene to a new specified position within the digital scene; generating a new local position indicator for the new specified position within the digital scene; modifying a global feature map of the digital scene based on the new local position indicator for the new specified position to form a new modified global feature map; generating new position-specific lighting parameters for the new specified position using a second set of network layers based on the new modified global feature map; and based on the position adjustment request, drawing an adjusted digital scene including the virtual object, which is illuminated at the new specified position according to the new position-specific lighting parameters.

[0175] As suggested above, in one or more embodiments, action 1100 also includes: identifying a perspective adjustment request to draw the digital scene from a different viewpoint; and based on the perspective adjustment request, drawing a modified digital scene from the different viewpoint, the modified digital scene including a virtual object illuminated at a specified location according to location-specific lighting parameters.

[0176] In addition, in some cases, action 1100 also includes identifying a perspective adjustment request from a different viewpoint to draw a virtual object at a specified location within the digital scene; generating a new local position indicator for the specified location within the digital scene from the different viewpoint; modifying a global feature map of the digital scene based on the new local position indicator for the specified location from the different viewpoint to form a new modified global feature map; generating new location-specific lighting parameters for the specified location from the different viewpoint using a second set of network layers based on the new modified global feature map; and based on the adjustment of the lighting conditions, drawing an adjusted digital scene, the scene including the virtual object illuminated at the specified location according to the new location-specific lighting parameters.

[0177] Additionally, in some implementations, action 1100 further includes: identifying an adjustment to a lighting condition for a specified location within a digital scene; extracting a new global feature map from the digital scene using a first set of network layers of a local illumination estimation neural network; modifying the new global feature map of the digital scene based on a local position indicator of the specified location; generating new location-specific lighting parameters for the specified location using a second set of network layers based on the new modified global feature map; and based on the adjustment to the lighting condition, drawing an adjusted digital scene including a virtual object illuminated at the specified location according to the new location-specific lighting parameters.

[0178] In addition to (or as an alternative to) the above actions, in some embodiments, action 1000 (or action 1100) includes a step for training a local illumination estimation neural network using a global feature training map for a digital training scene and a local position training indicator for a specified position within the digital training scene. Figure 4A or Figure 5A The described algorithms and actions may include corresponding actions for performing steps for training a local illumination estimation neural network using a global feature training map for digital training scenes and local location training indicators for specified locations within the digital training scenes.

[0179] Additionally or alternatively, in some embodiments, action 1000 (or action 1100) includes a step for generating location-specific lighting parameters for a specified location by utilizing a trained local lighting estimation neural network. Figure 4B or Figure 5B The described algorithms and actions may include corresponding actions for performing the following steps: generating location-specific lighting parameters for a specified location by utilizing a trained local lighting estimation neural network.

[0180] Embodiments of the present disclosure may include or utilize a special-purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in more detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more processes described herein may be implemented at least in part as instructions embodied in a non-transient computer-readable medium, and the instructions are executable by one or more computing devices (e.g., any media content access device described herein). Typically, a processor (e.g., a microprocessor) receives instructions from a non-transient computer-readable medium (e.g., a memory, etc.) and executes those instructions to perform one or more processes, including one or more of the processes described herein.

[0181] Computer readable media can be any available media that can be accessed by a general or special purpose computer system. A computer readable medium that stores computer executable instructions is a non-transient computer readable storage medium (device). A computer readable medium that carries computer executable instructions is a transmission medium. Therefore, by way of example and not limitation, embodiments of the present disclosure may include at least two distinct types of computer readable media: a non-transient computer readable storage medium (device) and a transmission medium.

[0182] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drives (“SSD”) (e.g., RAM-based), flash memory, phase-change memory (“PCM”), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code components in the form of computer-executable instructions or data structures that can be accessed by a general or special purpose computer.

[0183] "Network" is defined as one or more data links that enable the transport of electronic data between computer systems and / or generators and / or other electronic devices. When information is transmitted or provided to a computer via a network or other communication connection (hardwired, wireless, or a combination of hardwired or wireless), the computer will correctly view the connection as a transmission medium. The transmission medium may include a network and / or data link that can be used to carry the desired program code components in the form of computer executable instructions or data structures, and the desired program code components can be accessed by a general or special purpose computer. The above combinations should also be included in the scope of computer readable media.

[0184] In addition, upon reaching various computer system components, program code components in the form of computer executable instructions or data structures can be automatically transferred from the transmission medium to the non-transient computer readable storage medium (device) (or vice versa). For example, computer executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface generator (e.g., "NIC") and then ultimately transferred to the computer system RAM and / or a less volatile computer storage medium (device) in the computer system. Therefore, it should be understood that non-transient computer readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.

[0185] Computer executable instructions include, for example, instructions and data that, when executed on a processor, cause a general purpose computer, a special purpose computer, or a special purpose processing device to perform a specific function or set of functions. In one or more embodiments, computer executable instructions are executed on a general purpose computer to transform the general purpose computer into a special purpose computer that implements the elements of the present disclosure. Computer executable instructions can be, for example, binary files, intermediate format instructions (such as assembly language), or even source code. Although the subject matter has been described in language specific to structural marketing features and / or methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the marketing features or actions described above. Instead, the described marketing features and actions are disclosed as example forms of implementing the claims.

[0186] Those skilled in the art will appreciate that the present disclosure can be practiced in a network computing environment with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, etc. The present disclosure can also be practiced in a distributed system environment, in which local and remote computer systems linked by a network (by a hardwired data link, a wireless data link, or a combination of hardwired and wireless data links) all perform tasks. In a distributed system environment, the program generator can be located in local and remote memory devices.

[0187] Embodiments of the present disclosure may also be implemented in a cloud computing environment. In this specification, "cloud computing" is defined as a subscription model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing may be adopted in the market to provide universal and convenient on-demand access to a shared pool of configurable computing resources. A shared pool of configurable computing resources may be quickly provided via virtualization and released with less management effort or service provider interaction, and then scaled accordingly.

[0188] The cloud computing subscription model may consist of various features, such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured services, and the like. The cloud computing subscription model may also expose various service subscription models, such as, for example, software as a service ("SaaS"), web services, platform as a service ("PaaS"), and infrastructure as a service ("IaaS"). The cloud computing subscription model may also be deployed using different deployment subscription models, such as private cloud, community cloud, public cloud, hybrid cloud, and the like. In this specification and claims, a "cloud computing environment" is an environment in which cloud computing is employed.

[0189] Fig.12 A block diagram of an exemplary computing device 1200 that may be configured to perform one or more of the above-described processes is shown. Fig.12 As shown, computing device 1200 may include processor 1202, memory 1204, storage device 1206, I / O interface 1208, and communication interface 1210, which may be communicatively coupled via communication infrastructure 1212. Fig.12 The computing device 1200 may include fewer or more components than those shown. Fig.12 Components of computing device 1200 are shown.

[0190] In one or more embodiments, the processor 1202 includes hardware for executing instructions, such as those that constitute a computer program. By way of example and not limitation, to execute instructions for digitizing a real-world object, the processor 1202 may retrieve (or fetch) instructions from an internal register, an internal cache, a memory 1204, or a storage device 1206, and decode and execute them. The memory 1204 may be a volatile or non-volatile memory for storing data, metadata, and programs for execution by the processor(s). The storage device 1206 includes a storage device such as a hard disk, a flash disk drive, or other digital storage device for storing data or instructions related to an object digitization process (e.g., a digital scan, a digital model).

[0191] The I / O interface 1208 allows a user to provide input to the computing device 1200, receive output from the computing device 1200, and otherwise transmit data to and receive data from the computing device 1200. The I / O interface 1208 may include a mouse, a keypad or keyboard, a touch screen, a camera, an optical scanner, a network interface, a modem, other known I / O devices, or a combination of such I / O interfaces. The I / O interface 1208 may include one or more devices for presenting output to a user, including but not limited to a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In some embodiments, the I / O interface 1208 is configured to provide graphics data to a display for presentation to a user. The graphics data may represent one or more graphical user interfaces and / or any other graphical content that may serve a particular implementation.

[0192] The communication interface 1210 may include hardware, software, or both. In any case, the communication interface 1210 may provide one or more interfaces for communication between the computing device 1200 and one or more other computing devices or networks (such as, for example, packet-based communication). By way of example and not limitation, the communication interface 1210 may include a network interface controller (“NIC”) or network adapter for communicating with an Ethernet or other wired-based network, or a wireless NIC (“WNIC”) or wireless adapter for communicating with a wireless network (such as WI-FI).

[0193] In addition, the communication interface 1210 can facilitate communication with various types of wired or wireless networks. The communication interface 1210 can also facilitate communication using various communication protocols. The communication infrastructure 1212 can also include hardware, software, or both that couple the components of the computing device 1200 to each other. For example, the communication interface 1210 can use one or more networks and / or protocols to enable multiple computing devices connected by a specific infrastructure to communicate with each other to perform one or more aspects of the digitization process described herein. For illustration, the image compression process can allow multiple devices (e.g., a server device for performing image processing tasks for a large number of images) to exchange information using various communication networks and protocols for exchanging information about selected workflows and image data for multiple images.

[0194] In the foregoing description, the present disclosure has been described with reference to specific exemplary embodiments of the present disclosure. Various embodiments and aspects of the present disclosure are described with reference to the details discussed herein, and the accompanying drawings illustrate various embodiments. The above description and drawings are illustrative of the present disclosure and should not be construed as limiting the present disclosure. Many specific details are described to provide a thorough understanding of the various embodiments of the present disclosure.

[0195] Without departing from the spirit or essential features of the present invention, the present invention may be embodied in other specific forms. The described embodiments should be considered in all respects only as illustrative and not restrictive. For example, the method described herein may be performed with fewer or more steps / actions, or the steps / actions may be performed in different orders. In addition, the steps / actions described herein may be repeated or performed in parallel with each other or with different instances of the same or similar steps / actions. Therefore, the scope of the present application is indicated by the appended claims rather than the foregoing description. All changes falling within the equivalent meaning and scope of the claims should be included within their scope.

Claims

1. A non-transitory computer readable medium storing thereon instructions which, when executed by at least one processor, cause a computer system to: Identifying a request to draw a virtual object at a specified location within a digital scene; extracting a global feature map from the digital scene using a first set of network layers of a local illumination estimation neural network; generating a local position indicator for the specified position within the digital scene; modifying the global feature map for the digital scene based on the local position indicator for the specified position; generating location-specific lighting parameters for the designated location using a second set of network layers of the local lighting estimation neural network based on the modified global feature map; as well as Based on the request, a modified digital scene is rendered, the modified digital scene including the virtual object illuminated according to the location-specific lighting parameters at the specified location.

2. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the computer system to: generate the location-specific lighting parameters for the specified location by generating location-specific spherical harmonic coefficients indicating lighting conditions for the object at the specified location.

3. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the computer system to: identifying a position adjustment request to move the virtual object from the specified position within the digital scene to a new specified position within the digital scene; generating a new local position indicator for the new specified position within the digital scene; modifying the global feature map for the digital scene based on the new local position indicator for the new specified position to form a new modified global feature map; generating new location-specific lighting parameters for the new specified location using the second set of network layers based on the new modified global feature map; as well as Based on the position adjustment request, an adjusted digital scene including the virtual object is rendered, the virtual object being illuminated at the new specified position according to the new position-specific lighting parameters.

4. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the computer system to: identifying a view adjustment request for rendering the digital scene from a different viewpoint; and Based on the perspective adjustment request, the modified digital scene is rendered from the different viewpoint, the modified digital scene including the virtual object illuminated at the specified location according to the location-specific lighting parameters.

5. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the computer system to: generate the local position indicator for the specified position by identifying local position coordinates representing the specified position within the digital scene.

6. The non-transitory computer-readable medium of claim 5, further comprising instructions that, when executed by the at least one processor, cause the computer system to: The global feature map is modified to generate a modified global feature map in the following manner: generating a masked feature map from the local position coordinates; Multiplying the global feature map and the masked feature map for the local position coordinates to generate a masked dense feature map; as well as Concatenating the global feature map and the masked dense feature map to form a combined feature map; as well as The location-specific lighting parameters are generated by providing the combined feature map to the second set of network layers.

7. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the computer system to: identifying an adjustment to a lighting condition for the specified location within the digital scene; extracting a new global feature map from the digital scene using the first set of network layers of the local illumination estimation neural network; modifying the new global feature map for the digital scene based on the local position indicator for the specified position; generating new location-specific lighting parameters for the specified location using the second set of network layers based on the new modified global feature map; as well as Based on the adjustment of the lighting condition, an adjusted digital scene is rendered, the adjusted digital scene including the virtual object illuminated at the designated location according to the new location-specific lighting parameters.

8. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the computer system to: Identify a view adjustment request for rendering the virtual object at the specified position within the digital scene from a different viewpoint; generating a new local position indicator for the specified position within the digital scene from the different viewpoint; modifying the global feature map for the digital scene based on the new local position indicator for the specified position from the different viewpoint to form a new modified global feature map; generating new location-specific lighting parameters for the specified location from the different viewpoints using the second set of network layers based on the new modified global feature map; as well as Based on the viewing angle adjustment request, an adjusted digital scene is rendered, the adjusted digital scene including the virtual object illuminated at the designated position according to the new position-specific lighting parameters.

9. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the computer system to: The local position indicator for the specified position is generated by: Selecting a first pixel corresponding to the specified position from a first feature map corresponding to a first layer of the first set of network layers; and Selecting a second pixel corresponding to the specified position from a second feature map corresponding to a second layer of the first set of network layers; The global feature map is modified to generate a modified global feature map in the following manner: combining features for the first pixel and the second pixel corresponding to the specified position to generate a hypercolumn graph; and Concatenating the global feature map and the hypercolumn map to form a combined feature map; as well as The location-specific lighting parameters are generated by providing the combined feature map to the second set of network layers.

10. A system for local illumination estimation, comprising: at least one processor; at least one non-transitory computer-readable medium comprising respective digital training scenes and ground truth lighting parameters for respective locations within said respective digital training scenes; Local illumination estimation neural network; as well as instructions that, when executed by at least one processor, cause the system to train the local illumination estimation neural network by: extracting a global feature training graph from the digital training scenes of each of the digital training scenes using a first group of network layers of the local illumination estimation neural network; generating a local position training indicator for a specified position within the digital training scene; modifying the global feature training map for the digital training scene based on the local location training indicator for the specified location; generating location-specific lighting training parameters for the designated location using a second set of network layers of the local lighting estimation neural network based on the modified global feature training map; as well as Network parameters of the local illumination estimation neural network are modified based on a comparison of the location-specific lighting training parameters with a set of ground truth lighting parameters for the specified location within the digital training scene.

11. The system of claim 10, wherein the digital training scene comprises a three-dimensional digital model or a digital viewpoint image of a real scene.

12. The system of claim 10, further comprising instructions that, when executed by the at least one processor, cause the system to: generating the location-specific lighting training parameters for the specified location by generating location-specific spherical harmonic training coefficients indicative of lighting conditions at the specified location; and The set of ground truth lighting parameters for the specified position is determined by determining a set of ground truth position-specific spherical harmonic coefficients indicative of lighting conditions at the specified position.

13. The system of claim 12, further comprising instructions that, when executed by the at least one processor, cause the system to: generate a number of five of the position-specific spherical harmonic training coefficients for each color channel.

14. The system of claim 12, further comprising instructions that, when executed by the at least one processor, cause the system to: determine the set of base true position-specific spherical harmonic coefficients by: identifying locations within the digital training scene; generating a cube map for each location within the digital training scene; and The cube map for each location within the digital training scene is projected to the set of ground truth location-specific spherical harmonic coefficients.

15. The system of claim 10, further comprising instructions that, when executed by the at least one processor, cause the system to: generate the local position training indicator for the specified position by identifying local position training coordinates representing the specified position within the digital training scene.

16. The system of claim 15, further comprising instructions that, when executed by the at least one processor, cause the system to: The global feature training graph is modified to generate a modified global feature training graph in the following manner: generating a masked feature training map from the local position training coordinates; multiplying the global feature training map and the masked feature training map for the local position training coordinates to generate a masked dense feature training map; as well as splicing the global feature training graph and the masked dense feature training graph to form a combined feature training graph; as well as The location-specific lighting training parameters are generated by providing the combined feature training map to the second set of network layers.

17. The system of claim 10, further comprising instructions that, when executed by the at least one processor, cause the system to: generate the local location training indicator for the specified location by: selecting a first training pixel corresponding to the specified position from a first feature training map corresponding to a first layer of the local illumination estimation neural network; and A second training pixel corresponding to the specified position is selected from a second feature training map corresponding to a second layer of the local illumination estimation neural network.

18. The system of claim 10, wherein the first set of network layers of the local illumination estimation neural network comprises lower layers of a densely connected convolutional network, and the second set of network layers of the local illumination estimation neural network comprises convolutional layers and fully connected layers.

19. A computer-implemented method for estimating lighting conditions for virtual objects in a digital media environment for rendering an augmented reality scene, comprising: accessing each digital training scene and ground truth lighting parameters for each location within said each digital training scene; performing a step for training a local illumination estimation neural network using a global feature training map for each digital training scene and a local position training indicator for each specified position within each digital training scene; Identifying a request to draw a virtual object at a specified location within a digital scene; performing steps for generating location-specific lighting parameters for said specified location by utilizing said trained local lighting estimation neural network; as well as Based on the request, a modified digital scene is rendered, the modified digital scene including the virtual object illuminated at the specified location according to the location-specific lighting parameters.

20. The computer-implemented method of claim 19, further comprising: receiving, from a mobile device, the request to draw the virtual object at the specified location; as well as Based on receiving the request from the mobile device, the modified digital scene including the virtual object illuminated according to the location-specific lighting parameters at the specified location is drawn in a graphical user interface of the mobile device.

Citation Information

Patent Citations

  • A method and an apparatus for rendering an augmented reality scene

    CN109166170A

  • An efficient global illumination rendering method based on depth learning

    CN109389667A