Dynamically estimating lighting parameters for locations in augmented reality scenes using neural networks
The local lighting estimation neural network generates location-specific lighting parameters, which solves the problem of unreal-time adjustment of lighting conditions in augmented reality systems, realizes the realism and efficiency of virtual objects, and adapts to three-dimensional scenes of complex light changes.
Patent Information
- Application Number
- CN202510507766.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-21
- Filing Date
- 2019-11-22
- Publication Date
- 2025-08-08
AI Technical Summary
Existing augmented reality systems are difficult to adjust lighting conditions in real time or near real time to reflect changes in the real world, resulting in virtual objects appearing unreal or unnatural in digital augmented scenarios, especially in cases where light direction dependence and complex changes are found in three-dimensional scenarios.
The local lighting estimation neural network is used to generate position-specific lighting parameters, adjust the lighting conditions of virtual objects in real time, and use the first set of network layers of the local lighting estimation neural network to extract the global feature map, and modify the global feature map based on the local position indicator to generate position-specific lighting parameters to realize flexible lighting adjustments for different locations in the digital scene.
Accurate and rapid lighting adjustments to virtual objects are realized, and the lighting conditions can be flexibly updated from different perspectives and positions, improving the authenticity and efficiency of the augmented reality system, and reducing the consumption of computing resources.
Smart Images

Figure CN120451456A_ABST
Abstract
Description
[0001] Divisional Application Instructions
[0002] This application is a divisional application of the Chinese invention patent application with the application date of November 22, 2019, application number 201911159261.7, and name “Dynamic estimation of lighting parameters of positions in augmented reality scenes using neural networks”. Technical Field
[0003] Embodiments of the present disclosure relate to dynamically estimating lighting parameters of locations in an augmented reality scene using a neural network. Background Art
[0004] Augmented reality systems typically depict digitally enhanced images or other scenes with computer-simulated objects. To depict these scenes, augmented reality systems sometimes render real and computer-simulated objects with shadows and other lighting conditions. Many augmented reality systems attempt to seamlessly render virtual objects composited with real-world objects. To achieve convincing synthesis, augmented reality systems must illuminate virtual objects with consistent lighting that matches the physical scene. Because the real world is constantly changing (e.g., objects move, lighting changes), augmented reality systems that pre-capture lighting conditions are typically unable to adjust lighting conditions to reflect real-world changes.
[0005] Despite advances in estimating lighting conditions for digitally augmented scenes, several technical limitations still prevent conventional augmented reality systems from faithfully depicting lighting conditions on a computing device. These limitations include changing lighting conditions as the digitally augmented scene changes, rapidly rendering or adjusting lighting conditions in real time (or near real time), and faithfully capturing lighting variations across the entire scene. These limitations are exacerbated in three-dimensional scenes, where every location at a given moment can receive varying amounts of light from a full 360-degree range of directions. Both directional dependence and lighting variations across the entire scene play a crucial role when trying to faithfully and convincingly render synthetic objects into a scene.
[0006] For example, some conventional augmented reality systems are unable to depict the lighting conditions of computer-simulated objects in real time (or near real time). In some cases, conventional augmented reality systems use an ambient light model (i.e., only a single constant term and no directional information) to estimate the light that an object receives from its environment. For example, conventional augmented reality systems often use simple heuristics to create lighting conditions, such as by relying on the average brightness value of pixels of the object (or surrounding the object) to create lighting conditions in the ambient light model. Such an approximation does not capture directional variations in lighting and may not produce a reasonable approximation of ambient lighting under many conditions - resulting in unrealistic and unnatural lighting. Such lighting makes computer-simulated objects appear unrealistic or out of place in the digitally enhanced scene. For example, in some cases, when the light for the computer-simulated object comes from outside the perspective (or viewpoint) shown in the digitally enhanced image, conventional systems are unable to accurately depict lighting on the object.
[0007] In addition to the challenges of depicting realistic lighting, in some cases, conventional augmented reality systems lack the flexibility to adjust or change the lighting conditions for specific computer-simulated objects in a scene. For example, some augmented reality systems determine the lighting conditions of a digitally augmented image for a group of objects or the entire image, rather than for a specific object or location within the digitally augmented image. Because such lighting conditions typically apply to a group of objects or the entire image, conventional systems cannot adjust the lighting conditions for specific objects, or can only do so by re-determining the lighting conditions for the entire digitally augmented image, which inefficiently uses computing resources.
[0008] Independent of technical limitations that affect the realism or flexibility of lighting in augmented reality, conventional augmented reality systems are sometimes unable to quickly estimate the lighting conditions of objects in a digitally augmented scene. For example, some conventional augmented reality systems receive user input defining baseline parameters, such as image geometry or material properties, and estimate parametric lighting for the digitally augmented scene based on the baseline parameters. While some conventional systems can apply such user-defined parameters to accurately estimate lighting conditions, such systems are unable to quickly estimate parametric lighting nor apply image geometry-specific lighting models to other scenes with different light sources and geometries. Summary of the Invention
[0009] The present disclosure describes embodiments of methods, non-transitory computer-readable media, and systems that, among other benefits, address the aforementioned problems. For example, based on a request to draw a virtual object in a digital scene, the disclosed system generates location-specific lighting parameters for a specified location within the digital scene using a local lighting estimation neural network. In certain implementations, the disclosed system draws a modified digital scene that includes a virtual object at a specified location illuminated according to the location-specific lighting parameters. As described below, the disclosed system can generate such location-specific lighting parameters to spatially vary lighting for different locations within the digital scene. Therefore, because the request to draw the virtual object is in real-time (or near real-time), the disclosed system can quickly generate different location-specific lighting parameters based on such drawing requests that accurately reflect the lighting conditions at different locations of the digital scene.
[0010] For example, in some embodiments, the disclosed system identifies a request to draw a virtual object at a specified location within a digital scene. The disclosed system extracts a global feature map from the digital scene using a first set of network layers of a local illumination estimation neural network. The system also generates a local position indicator for the specified location and modifies the global feature map of the digital scene based on the local position indicator. Based on the modified global feature map, the system generates location-specific lighting parameters for the specified location using a second set of layers of the local illumination estimation neural network. In response to the drawing request, the system draws the modified digital scene, which includes the virtual object at the specified location illuminated according to the location-specific lighting parameters.
[0011] Additional features and advantages of the disclosed methods, non-transitory computer-readable media, and systems are set forth in the description that follows and may be obvious from, or may be disclosed by, the practice of the exemplary embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The detailed description refers to the accompanying drawings, which are briefly described below.
[0013] Figure 1 An augmented reality system and a lighting estimation system are illustrated that use a local lighting estimation neural network to generate location-specific lighting parameters for a specified location within a digital scene, and draw a modified digital scene including a virtual object at the specified location according to the parameters of one or more embodiments.
[0014] Figure 2 Illustrated are digital training scenes and corresponding cubemaps in accordance with one or more embodiments.
[0015] Figure 3 Illustrated are degrees of position-specific spherical harmonic training coefficients according to one or more embodiments.
[0016] Figure 4A The diagram illustrates an illumination estimation system that trains a local illumination estimation neural network to generate location-specific spherical harmonic coefficients for each specified location within a digital scene, in accordance with one or more embodiments.
[0017] Figure 4B The diagram illustrates an illumination estimation system that uses a trained local illumination estimation neural network to generate location-specific spherical harmonic coefficients for a specified location within a digital scene, in accordance with one or more embodiments.
[0018] Figure 5A The diagram illustrates an illumination estimation system that trains a local illumination estimation neural network to generate location-specific spherical harmonic coefficients for each specified location within a digital scene using coordinates of feature maps from layers of the neural network, in accordance with one or more embodiments.
[0019] Figure 5B The diagram illustrates an illumination estimation system that uses a trained local illumination estimation neural network to generate location-specific spherical harmonic coefficients for a specified location in a digital scene using coordinates of feature maps from layers of the neural network, in accordance with one or more embodiments.
[0020] Figures 6A-6C The diagram illustrates a computing device that renders a digital scene including virtual objects at a specified location and a new specified location according to location-specific lighting parameters and new location-specific lighting parameters, respectively, in response to different rendering requests, according to one or more embodiments.
[0021] Figure 7A 1 illustrates a digital scene and corresponding red, green, and blue representations and light intensity representations rendered according to ground truth lighting parameters and location-specific lighting parameters generated by a lighting estimation system in accordance with one or more embodiments.
[0022] Figure 7B The diagram illustrates a digital scene including a virtual object at a specified location illuminated according to base real lighting parameters, and a digital scene including a virtual object at a specified location illuminated according to location-specific lighting parameters generated by a lighting estimation system, according to one or more embodiments.
[0023] Figure 8 FIG. 1 illustrates a block diagram of an environment in which a lighting estimation system may operate in accordance with one or more embodiments.
[0024] Figure 9 FIGURES illustrate an embodiment of the present invention according to one or more embodiments. Figure 8 Schematic diagram of the lighting estimation system.
[0025] Figure 10
[0046] Figure 1 is a flow diagram illustrating a series of actions for training a local lighting estimation neural network to generate location-specific lighting parameters in accordance with one or more embodiments.
[0026] Figure 11
[0014] Figure 1 is a flow diagram illustrating a series of actions for applying a trained local illumination estimation neural network to generate location-specific illumination parameters in accordance with one or more embodiments.
[0027] Figure 12 The figure illustrates a block diagram of an exemplary computing device for implementing one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0028] This disclosure describes one or more embodiments of a lighting estimation system that uses a local lighting estimation neural network to estimate lighting parameters for specific locations within a digital scene for augmented reality. For example, based on a request to draw a virtual object in the digital scene, the lighting estimation system uses the local lighting estimation neural network to generate location-specific lighting parameters for the specified location within the digital scene. In certain implementations, the lighting estimation system also draws a modified digital scene including the virtual object at the specified location based on the parameters. In some embodiments, the lighting estimation system generates such location-specific lighting parameters to spatially vary and adapt lighting conditions for different locations within the digital scene. Because requests to draw virtual objects are processed in real time (or near real time), the lighting estimation system can rapidly generate different location-specific lighting parameters in response to the drawing requests that accurately reflect lighting conditions at different locations within the digital scene or from different perspectives of the digital scene. The lighting estimation system can also rapidly generate different location-specific lighting parameters to reflect changes in lighting or other conditions.
[0029] For example, in some embodiments, an illumination estimation system identifies a request to draw a virtual object at a specified location within a digital scene. To draw such a scene, the illumination estimation system extracts a global feature map from the digital scene using a first set of network layers of a local illumination estimation neural network. The illumination estimation system also generates a local position indicator for the specified location and modifies the global feature map of the digital scene based on the local position indicator. Based on the modified global feature map, the illumination estimation system generates location-specific lighting parameters for the specified location using a second set of layers of the local illumination estimation neural network. In response to the drawing request, the illumination estimation system draws the modified digital scene, which includes the virtual object illuminated at the specified location according to the location-specific lighting parameters.
[0030] By using location-specific lighting parameters, in some embodiments, the lighting estimation system can illuminate virtual objects from different viewpoints of the scene and rapidly update lighting conditions for different locations, different viewpoints, changes in lighting applied to the scene or virtual objects within the scene, or other environmental changes. For example, in some cases, the lighting estimation system generates location-specific lighting parameters that capture the lighting conditions at the location of the virtual object from various viewpoints within the digital scene. Upon identifying a position adjustment request to move the virtual object to a new designated location, the lighting estimation system can also generate a new local position indicator for the new designated location and modify the global feature map for the digital scene using a neural network to output new lighting parameters for the new designated location. Upon identifying or otherwise responding to a change in the lighting conditions of the digital scene, the lighting estimation system can similarly update the global feature map of the digital scene to output new lighting parameters for the new lighting conditions. For example, the lighting estimation system can dynamically determine or update lighting parameters when the viewer's viewpoint changes (e.g., the camera moves within the scene), when the lighting in the scene changes (e.g., lights are added, dimmed, obscured, or exposed), or when objects in the scene or the scene itself changes.
[0031] To generate location-specific lighting parameters, the lighting estimation system can use different types of local position indicators. For example, in some embodiments, the lighting estimation system identifies (as a local position indicator) local position coordinates representing a specified location within the digital scene. In contrast, in some implementations, the lighting estimation system identifies one or more local position indicators from features extracted by different layers of the local lighting estimation neural network, such as pixels corresponding to the specified location from different feature maps extracted by the neural network layers.
[0032] When generating location-specific lighting parameters, the lighting estimation system can generate spherical harmonic coefficients that indicate the lighting conditions for a specified location of a virtual object in a digital scene. When the digital scene is represented with low dynamic range ("LDR") lighting, such location-specific spherical harmonic coefficients can capture the high dynamic range ("HDR") lighting at a location within the digital scene. When the virtual object changes position in the digital scene, the lighting estimation system can use a local lighting estimation neural network to generate new location-specific spherical harmonic coefficients by requesting that the change in lighting be realistically depicted at the changed location of the virtual object.
[0033] As described above, in some embodiments, the lighting estimation system not only applies a local lighting estimation neural network but can also optionally train such a network to generate location-specific lighting parameters. When training the neural network, in certain implementations, the lighting estimation system uses a first set of layers of the local lighting estimation neural network to extract a global feature training map from the digital training scene. The lighting estimation system also generates a local location training indicator for a specified location within the digital training scene and modifies the global feature training map based on the local location training indicator for the specified location.
[0034] Based on the modified global feature training graph, the illumination estimation system uses a second set of network layers of the local illumination estimation neural network to generate location-specific illumination training parameters for the specified location. The illumination estimation system then modifies the network parameters of the local illumination estimation neural network based on a comparison of the location-specific illumination training parameters for the specified location within the digital training scene with the underlying real-world illumination parameters. By iteratively generating these location-specific illumination training parameters and adjusting the network parameters of the neural network, the illumination estimation system can train the local illumination estimation neural network to a point of convergence.
[0035] As previously described, the lighting estimation system can specify locations using ground-truth lighting parameters to facilitate training. To create such ground-truth lighting parameters, in some embodiments, the lighting estimation system generates cubemaps for various locations within the digital training scene. The lighting estimation system then projects the cubemaps of the digital training scene onto ground-truth spherical harmonic coefficients. These ground-truth spherical harmonic coefficients can be used for comparison when iteratively training the local lighting estimation neural network.
[0036] As described above, the disclosed lighting estimation system overcomes several technical deficiencies that have hindered conventional augmented reality systems. For example, the lighting estimation system improves the accuracy and realism with which existing augmented reality systems generate lighting conditions for specific locations within a digital scene. As described above, the lighting estimation system can create such realistic lighting in part by using a local lighting estimation neural network trained to generate location-specific spherical lighting parameters based on local position indicators for specified locations within the digital scene.
[0037] Unlike some conventional systems that use average luminance values, which can result in unnatural brightness, the disclosed lighting estimation system can create lighting parameters with coordinate-level accuracy corresponding to local position indicators. Furthermore, unlike some conventional systems that cannot depict lighting from outside the digital scene's viewing angle, the disclosed lighting estimation system can create lighting parameters that capture lighting conditions emanating from light sources outside the digital scene's viewing angle. To achieve this accuracy, in some embodiments, the lighting estimation system generates location-specific spherical harmonic coefficients that effectively capture location-specific, realistic, and natural-looking lighting conditions from multiple viewpoints within the digital scene.
[0038] In addition to more realistically depicting lighting, in some embodiments, the lighting estimation system exhibits greater flexibility in rendering different lighting conditions for different locations relative to existing augmented reality systems. Unlike some conventional augmented reality systems, which are limited to redefining lighting for a group of objects or an entire image, the lighting estimation system can flexibly adapt lighting conditions to different locations to which a virtual object is moved. For example, upon identifying a position adjustment request for a moving virtual object, the disclosed lighting estimation system can modify an existing global feature map of the digital scene using a new local position indicator. By modifying the global feature map to reflect the newly specified location, the lighting estimation system can generate new location-specific lighting parameters for the newly specified location without having to redefine lighting conditions for other objects or the entire image. This flexibility enables users to manipulate objects in augmented reality applications on mobile devices or other computing devices.
[0039] Being both reality-independent and flexible, the disclosed lighting estimation system can also increase the speed at which augmented reality systems can render digital scenes with location-specific lighting for virtual objects. Unlike illumination models that rely on examining the geometry of an image or similar baseline parameters, the disclosed lighting estimation system estimates lighting using a neural network that requires relatively few inputs: indicators of the digital scene and the locations of virtual objects. By training a localized lighting estimation neural network to analyze these inputs, the lighting estimation system reduces the computational resources required to quickly generate location-specific lighting parameters within a digital scene.
[0040] Now go to Figure 1 , which illustrates an augmented reality system 108 and a lighting estimation system 110 that use a neural network to estimate location-specific lighting parameters. In general, as Figure 1 As shown, the lighting estimation system 110 identifies a request to render a virtual object 106 at a specified location in the digital scene 102 and generates location-specific lighting parameters 114 for the specified location using the local lighting estimation neural network 112. Based on the request, the augmented reality system 108, in conjunction with the lighting estimation system 110, renders a modified digital scene 116 that includes the virtual object 106 at the specified location illuminated according to the location-specific lighting parameters 114. Figure 1 Augmented reality system 108 is depicted including lighting estimation system 110 and rendering modified digital scene 116 , but lighting estimation system 110 may alternatively render modified digital scene 116 alone.
[0041] As just noted, the lighting estimation system 110 identifies a request to draw a virtual object 106 at a specified location within the digital scene 102. For example, the lighting estimation system 110 may identify a digital request from a mobile device to draw a virtual pillow (or other virtual item) at a specific location on a piece of furniture (or another real item) depicted in a digital image. Regardless of the type of object or scene from the request, in some embodiments, the request to draw a digital scene includes an indication of a specified location to draw the virtual object.
[0042] As used in this disclosure, the term "digital scene" refers to a digital image, model, or depiction of an object. For example, in some embodiments, a digital scene includes a digital image of a real scene from a particular viewpoint or from multiple viewpoints. As another example, a digital scene may include a three-dimensional digital model of the scene. Regardless of the format, a digital scene may include a depiction of light from a light source. As just one example, a digital scene may include a digital image of a real room containing real walls, carpet, furniture, and people, with light emanating from a lamp or window. As discussed further below, the digital scene may be modified to include virtual objects in the adjusted or modified digital scene depicting augmented reality.
[0043] Relatedly, the term "virtual object" refers to a computer-generated graphical object that does not exist in the physical world. For example, a virtual object may include an object created by a computer for use in an augmented reality application. Such a virtual object may be, but is not limited to, a virtual accessory, an animal, clothing, cosmetics, footwear, fixtures, furniture, furnishings, hair, a person, a human feature, a vehicle, or any other computer-created graphical object. This disclosure often uses the word "virtual" to designate a specific virtual object (e.g., a "virtual pillow" or "virtual shoes"), but generally refers to real objects without the word "real" (e.g., a "bed," a "sofa").
[0044] like Figure 1 As further shown, the lighting estimation system 110 identifies or generates a local position indicator 104 for a specified location in a drawing request. As used herein, the term "local position indicator" refers to a digital identifier for a location within a digital scene. For example, in some implementations, the local position indicator includes digital coordinates, pixels, or other markers that indicate a specified location within the digital scene from a request to draw a virtual object. For illustration, the local position indicator can be a coordinate representing the specified location or a pixel (or coordinates of a pixel) corresponding to the specified location. In other embodiments, the lighting estimation system 110 can generate (and input) the local position indicator into the local lighting estimation neural network 112, or use the local lighting estimation neural network 112 to identify one or more local position indicators from a feature map.
[0045] In addition to generating the local position indicators, illumination estimation system 110 analyzes one or both of digital scene 102 and local position indicators 104 using local illumination estimation neural network 112. For example, in some cases, illumination estimation system 110 extracts a global feature map from digital scene 102 using a first set of layers of local illumination estimation neural network 112. Illumination estimation system 110 can further modify the global feature map of digital scene 102 based on local position indicators 104 for the specified location.
[0046] As used herein, the term "global feature map" refers to a multidimensional array or multidimensional vector representing features of a digital scene (e.g., a digital image or a three-dimensional digital model). For example, a global feature map for a digital scene can represent different visual or latent features of the entire digital scene, such as lighting or geometric features visible or embedded in the digital image or three-dimensional digital model. As described below, one or more layers of a local illumination estimation neural network output a global feature map of the digital scene.
[0047] The term "local illumination estimation neural network" refers to an artificial neural network that generates lighting parameters that indicate lighting conditions at a location within a digital scene. In particular, in certain implementations, the local illumination estimation neural network refers to an artificial neural network that generates a location-specific lighting parameter image that indicates lighting conditions for a specified location corresponding to a virtual object within the digital scene. In some embodiments, the local illumination estimation neural network includes some or all of the following network layers: one or more layers in a densely connected convolutional network ("DenseNet"), convolutional layers, and fully connected layers.
[0048] After modifying the global feature map, the lighting estimation system 110 uses a second set of network layers of the local lighting estimation neural network 112 to generate location-specific lighting parameters 114 based on the modified global feature map. As used in this disclosure, the term "location-specific lighting parameters" refers to parameters that indicate the lighting or illumination of a portion of a digital scene, or the lighting or illumination of a location in the digital scene. For example, in some embodiments, the location-specific lighting parameters define, specify, or otherwise indicate the lighting or shading of pixels corresponding to a specified location of the digital scene. Such location-specific lighting parameters can define the shading or hue of a pixel of a virtual object at the specified location. In some embodiments, the location-specific lighting parameters include spherical harmonic coefficients that indicate lighting conditions at a specified location within the digital scene for the virtual object. Thus, the location-specific lighting parameters can be functions corresponding to the surface of a sphere.
[0049] like Figure 1As further shown, in addition to generating such lighting parameters, the augmented reality system 108 also renders the modified digital scene 116, which includes the virtual object 106 at the specified location illuminated according to the location-specific lighting parameters 114. For example, in some embodiments, the augmented reality system 108 overlays or otherwise integrates a computer-generated image of the virtual object 106 within the digital scene 102. As part of the rendering, the augmented reality system 108 selects and renders pixels of the virtual object 106 that reflect the lighting, shading, or appropriate hue indicated by the location-specific lighting parameters 114.
[0050] As suggested above, in some embodiments, the lighting estimation system 110 uses a cubemap for the digital scene to project ground truth lighting parameters for specified locations of the digital scene. Figure 2 An example of a cubemap corresponding to a digital scene is shown in FIG. Figure 2 As shown, digital training scene 202 includes viewpoints of objects illuminated by light sources. To generate ground truth lighting parameters for training, in some cases, lighting estimation system 110 selects and identifies locations in digital training scene 202. Lighting estimation system 110 also generates cubemaps 204a-204d corresponding to the identified locations, where each cubemap represents the identified location within digital training scene 202. Lighting estimation system 110 then projects cubemaps 204a-204d to ground truth spherical harmonic coefficients for use in training the local lighting estimation neural network.
[0051] The lighting estimation system 110 optionally generates or prepares a digital training scene, such as digital training scene 202, by modifying an image of a real-world scene or a computer-generated scene. For example, in some cases, the lighting estimation system 110 modifies a three-dimensional scene from the Princeton University SUNCG dataset, as described by Shuran Song et al. in "Semantic Scene Completion from a Single Depth Image," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017), the entire contents of which are incorporated herein by reference. Scenes in the SUNCG dataset typically include realistic room and furniture layouts. Based on the SUNCG dataset, the lighting estimation system 110 computes a physically based rendering of the scene image. In some such cases, the lighting estimation system 110 uses the Mitsuba framework to compute physically based rendering, as described by Yinda Zhang et al., “Physically-Based Rendering for Indoor Scene Understanding Using Convolutional Neural Networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017), (hereinafter referred to as “Zhang”), the entire contents of which are incorporated herein by reference.
[0052] To eliminate some of the inaccuracies and biases in this rendering, in some embodiments, the lighting estimation system 110 changes the computational approach of Zhang's algorithm in some ways. First, the lighting estimation system 110 digitally removes lights that appear inconsistent with the indoor scene, such as area lights for floors, walls, and ceilings. Second, instead of using a single panorama for outdoor lighting as in Zhang, the researchers randomly selected a panorama from a dataset of 200 HDR outdoor panoramas and applied a random rotation around the panorama's Y axis. Third, instead of assigning the same intensity to each indoor area light, the lighting estimation system 110 randomly selects light intensities between one hundred and five hundred candelas with a uniform distribution. However, in some implementations of generating digital training scenes, the lighting estimation system 110 uses the same rendering method and spatial resolution as described by Zhang.
[0053] Because indoor scenes can include arbitrary distributions of light sources and light intensities, the lighting estimation system 110 normalizes the rendering of each physically based digital scene. When normalizing the rendering, the lighting estimation system 110 uses the following equation:
[0054]
[0055] In Equation (1), I represents the original image rendering with HDR, I′ represents the re-exposed image rendering, m is set to a value of 0.8, and P 90 represents 90% of the original image drawing I. By re-exposing the original image drawing I, the re-exposed image drawing I′ still contains HDR values. The lighting estimation system 110 further applies a gamma tone mapping operation with random values between 1.8 and 2.2 to the re-exposed image drawing I′ and clips all values greater than 1. By applying the gamma tone mapping operation, the lighting estimation system 110 can produce an image with a saturated bright window and improved scene contrast.
[0056] As described above, the lighting estimation system 110 identifies sample locations within a digital training scene, such as the digital training scene 202. Figure 2 As shown, digital training scene 202 includes four spheres that identify four sample locations identified by illumination estimation system 110. Digital training scene 202 includes the spheres for illustrative purposes. Therefore, the spheres represent visualizations of the sample locations identified by illumination estimation system 110, rather than objects in digital training scene 202.
[0057] To identify such sample locations, in some embodiments, lighting estimation system 110 identifies four different quadrants of digital training scene 202 with a margin of 20% of the image resolution from the image border. In each quadrant, lighting estimation system 110 identifies one sample location.
[0058] As further noted above, the lighting estimation system 110 generates cube maps 204a-204d based on sample positions from the digital training scene 202. Figure 2 As shown, each cubemap 204a-204d includes six visual portions in HDR, representing six different perspectives of the digital training scene 202. Cubemap 204a, for example, includes visual portions 206a-206e. In some cases, each cubemap 204a-204d includes a 64×64×6 resolution. This resolution has minimal impact on spherical harmonics of degree 5 or less and facilitates rapid cubemap rendering.
[0059] To draw cubemaps 204a-204d, the lighting estimation system 110 can apply a two-stage Principal Sample Space Metropolis Lighting Transport ("PSSMLT") with 512 direct samples. When generating the visual portion of a cubemap, such as visual portion 206c, the lighting estimation system 110 translates the surface position by 10 centimeters along the surface normal to minimize the risk of placing a portion of the cubemap within the surface of another object. In some implementations, the lighting estimation system 110 uses the same method to identify sample locations in a digital training scene and generate corresponding cubemaps as a prelude to determining the specific spherical harmonic coefficients of the underlying real-world locations.
[0060] For example, after generating the cubemaps 204a-204d, the lighting estimation system 110 projects the cubemaps 204a-204d onto the ground truth position-specific spherical harmonic coefficients for each identified position in the digital training scene 202. In some cases, the ground truth position-specific spherical harmonic coefficients include coefficients of degree 5. To calculate such spherical harmonics, in some embodiments, the lighting estimation system 110 applies a least squares method to the projected cubemaps.
[0061] For example, the lighting estimation system 110 may project the cubemap using the following equation:
[0062]
[0063] In equation (2), f represents the light intensity in each direction shown by the visible part of the cube map, where the light intensity is weighted corresponding to the solid angle of the pixel position. Denotes the spherical harmonics of degree 1 and order m. In some cases, for each cubemap, the illumination estimation system 110 computes spherical harmonic coefficients of degree 5 (or some other order) for each color channel (eg, third order), resulting in 36×3 spherical harmonic coefficients.
[0064] In some embodiments, in addition to generating ground truth position-specific spherical harmonic coefficients, the lighting estimation system 110 also enhances the digital training scene in specific ways. First, the lighting estimation system 110 randomly scales the exposure to a uniform distribution between 0.2 and 4. Second, the lighting estimation system 110 randomly sets the gamma value used for the tone mapping operator between 1.8 and 2.2. Third, the lighting estimation system 110 inverts the viewpoint of the digital training scene on the X-axis. Similarly, the lighting estimation system 110 inverts the negative order harmonics (e.g., the sign ) to transform the underlying real spherical harmonic coefficients to match the inverted viewpoint.
[0065] As further suggested above, in some implementations, the illumination estimation system 110 can use spherical harmonic coefficients of different degrees. For example, the illumination estimation system 110 can generate five-degree base true spherical harmonic coefficients for each color channel, or five-degree position-specific spherical harmonic coefficients for each color channel. Figure 3 Illustrated are lighting conditions according to spherical harmonics of different orders, visual representations of various orders of spherical harmonics, and a complete environment map for a location within a digital scene.
[0066] like Figure 3 As shown, visual representations 302a, 302b, 302c, 302d, and 302e correspond to spherical harmonics of the first, second, third, fourth, and fifth orders, respectively. In particular, each of visual representations 302a-302e includes one or more spheres in a row that visually represent different orders. Figure 3 As shown, with each increase in order, the spherical harmonic coefficients indicate more detailed lighting conditions.
[0067] To illustrate, Figure 3 Lighting depictions 306a-306e of a complete environment map 304 for a location within a digital scene are included. Lighting depictions 306a, 306b, 306c, 306d, and 306e correspond to first-order, second-order, third-order, fourth-order, and fifth-order spherical harmonic coefficients, respectively. As the order of the spherical harmonics increases, lighting depictions 306a-306e better capture the lighting shown in complete environment map 304. As shown by lighting depictions 306a-306e, spherical harmonic coefficients can implicitly capture occlusion and geometry of the digital scene.
[0068] As suggested above, the lighting estimation system 110 may use various architectures and inputs for the local lighting estimation neural network. Figure 4A and Figure 4B Depicted are examples of an illumination estimation system 110 that trains and applies, respectively, a local illumination estimation neural network to generate location-specific illumination parameters. Figure 5A and 5B Another embodiment of an illumination estimation system 110 is depicted that trains and applies a local illumination estimation neural network to generate location-specific illumination parameters. Both embodiments can use digital training scenes and corresponding spherical harmonic coefficients from a cubemap, such as Figure 2 and Figure 3 described.
[0069] For example, Figure 4AAs shown, the lighting estimation system 110 iteratively trains the local lighting estimation neural network 406. As an overview of the training iterations, the lighting estimation system 110 extracts a global feature training map from the digital training scene using a first set of network layers 408 of the local lighting estimation neural network 406. The lighting estimation system 110 also generates local location training indicators for specified locations within the digital training scene and modifies the global feature training map based on the local location training indicators.
[0070] Based on the modifications to the global feature training map reflected in the combined feature training map, the illumination estimation system 110 generates location-specific spherical harmonic coefficients for the specified location using a second set of network layers 416 of the local illumination estimation neural network 406. The illumination estimation system 110 then modifies the network parameters of the local illumination estimation neural network 406 based on a comparison of the location-specific spherical harmonic coefficients for the specified location in the digital training scene with the underlying true spherical harmonic coefficients.
[0071] For example, Figure 4A As shown, the lighting estimation system 110 feeds the digital training scene 402 to the local lighting estimation neural network 406. After receiving the digital training scene 402 as training input, the first set of network layers 408 extracts a global feature training map 410 from the digital training scene 402. In some such embodiments, the global feature training map 410 represents visual features of the digital training scene 402, such as the position of the light source and the global geometry of the digital training scene 402.
[0072] As described above, in some implementations, the first set of network layers 408 includes layers of a DenseNet, such as the various lower layers of a DenseNet. For example, the first set of network layers 408 may include a convolutional layer followed by a dense block, and (in some cases) one or more groups of convolutional layers, pooling layers, and dense blocks. In some cases, the first set of network layers 408 includes layers from DenseNet 120, as described in "Densely Connected Convolutional Layers" by G. Huang et al., in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017) (hereinafter referred to as "Huang"), the entire contents of which are incorporated herein by reference. The lighting estimation system 110 optionally initializes network parameters for each layer of the DenseNet using weights trained on ImageNet, as described in "ImageNet Large Scale Visual Recognition Challenge," International Journal of Computer Vision, Vol. 30, No. 3, 211-252 (2015) (hereinafter referred to as "Russakovsky"), the entire contents of which are incorporated herein by reference. Regardless of the architecture for initializing network parameters for the first set of network layers 408, the first set of network layers 408 optionally outputs a global feature training map 410 in the form of a dense feature map corresponding to the digital training scene 402.
[0073] In an alternative embodiment of the DenseNet layer, the first set of network layers 408 includes an encoder from a convolutional neural network ("CNN"), which includes several convolutional layers followed by four residual layers. In some such embodiments, the first set of network layers 408 includes an encoder as described by Marc-André Gardner et al., "Learning to Predict Indoor Illumination from a Single Image," ACM Transactions on Graphics, Vol. 36, No. 6, 2017 (hereinafter referred to as "Gardner"), the entire contents of which are incorporated herein by reference. Thus, as an encoder, the first set of network layers 408 optionally outputs a global feature training map 410 in the form of an encoded feature map of the digital training scene 402.
[0074] like Figure 4AAs further shown, illumination estimation system 110 identifies or generates a local position training indicator 404 for a specified location within digital training scene 404. For example, illumination estimation system 110 may identify local position training coordinates representing the specified location within digital training scene 402. Such coordinates may represent two-dimensional coordinates of a sample location within digital training scene 402 that correspond to a cubemap (and underlying real spherical harmonic coefficients) generated by illumination estimation system 110.
[0075] After having identified the local position training indicators 404, the illumination estimation system 110 uses the local position training indicators 404 to modify the global feature training map 410. In some cases, for example, the illumination estimation system 110 uses the local position training indicators 404 to mask the global feature training map 410. For example, the illumination estimation system 110 optionally generates a masked feature training map based on the local position training indicators 404, such as by applying a vector encoder to the local position training indicators 404 (e.g., by one-hot encoding). In some implementations, the masked feature training map includes an array of values that indicates the local position training indicators 404 for specified locations within the digital training scene 402, such as one or more values indicating the coordinates of the specified locations within the digital training scene 402, and other values (e.g., the number zero) indicating the coordinates of other locations within the digital training scene 402.
[0076] like Figure 4A As further shown, the lighting estimation system 110 multiplies the global feature training map 410 and the masked feature training map of the local location training indicator 404 to generate a masked dense feature training map 412. Because the masked feature training map and the global feature training map 410 can include the same spatial resolution, in some embodiments, the masked feature training map effectively masks the global feature training map 410 to create the masked dense feature training map 412. The masked dense feature training map 412 accordingly represents the local feature map for the specified location within the digital training scene 402.
[0077] In some embodiments, when generating the masked dense feature training map 412, the illumination estimation system 110 concatenates the global feature training map 410 and the masked dense feature training map 412 to form a combined feature training map 414. For example, the illumination estimation system 110 couples the global feature training map 410 and the masked dense feature training map 412 together to form a dual feature map or a stacked feature map as the combined feature training map 414. Alternatively, in some implementations, the illumination estimation system 110 combines rows of values from the global feature training map 410 with rows of values from the masked dense feature training map 412 to form the combined feature training map 414. However, any suitable concatenation method may be used.
[0078] like Figure 4A As further shown, the lighting estimation system 110 feeds the combined feature training map 414 to a second set of network layers 416 of the local lighting estimation neural network 406. Figure 4A As shown, the second set of network layers 416 includes one or both of network layers 418 and 420 and may include a regressor. In some implementations, for example, the second set of network layers 416 includes a convolutional layer as network layer 418 and a fully connected layer as network layer 420. The second set of network layers 416 optionally includes a max pooling layer between the convolutional layers constituting network layer 418 and a dropout layer before the fully connected layer constituting network layer 420. As part of a training iteration with such layers, the lighting estimation system 110 optionally passes the combined feature training map 414 through several convolutional layers and several fully connected layers with batch normalization and exponential learning unit (“ELU”) learning functions.
[0079] After passing the combined feature training map 414 through the second set of network layers 416, the local illumination estimation neural network 406 outputs position-specific spherical harmonic training coefficients 422. Consistent with the disclosure above, the position-specific spherical harmonic training coefficients 422 indicate the lighting conditions for a specified position within the digital training scene 402. For example, the position-specific spherical harmonic training coefficients 422 indicate the lighting conditions for a specified position within the digital training scene 402 identified by the local position training indicator 404.
[0080] After generating the position-specific spherical harmonic training coefficients 422, the illumination estimation system 110 compares the position-specific spherical harmonic training coefficients 422 with the base real spherical harmonic coefficients 426. As used in this disclosure, the term "base real spherical harmonic coefficients" refers to spherical harmonic coefficients that are empirically determined based on one or more cubemaps. For example, the base real spherical harmonic coefficients 426 represent spherical harmonic coefficients projected from a cubemap corresponding to a position within the digital training scene 402 identified by the local position training indicator 404.
[0081] As Figure 4A As further indicated, the illumination estimation system 110 uses a loss function 424 to compare the position-specific spherical harmonic training coefficients 422 with the underlying true spherical harmonic coefficients 426. In some embodiments, the illumination estimation system 110 uses a mean squared error ("MSE") function as the loss function 424. Alternatively, in some implementations, the illumination estimation system 110 uses an L2 loss function, a mean absolute error function, a mean absolute percentage error function, a rendering loss function, a root mean square error function, or other suitable loss function as the loss function 424.
[0082] After determining the loss from loss function 424, illumination estimation system 110 modifies network parameters (e.g., weights or values) of local illumination estimation neural network 406 to reduce the loss of loss function 424 in subsequent training iterations using backpropagation indicated by arrows from loss function 434 to local illumination estimation neural network 406. For example, illumination estimation system 110 may increase or decrease weights or values from some (or all) of first set of network layers 408 or second set of network layers 416 within local illumination estimation neural network 406 to reduce or minimize the loss in subsequent training iterations.
[0083] After modifying the network parameters of the local illumination estimation neural network 406 for the initial training iteration, the illumination estimation system 110 can perform additional training iterations. In subsequent training iterations, for example, the illumination estimation system 110 extracts additional global feature training maps for additional digital training scenes, generates additional local position training indicators for specified locations within the additional digital training scenes, and modifies the additional global feature training maps based on the additional local position training indicators. Based on the additional combined feature training maps, the illumination estimation system 110 generates additional position-specific spherical harmonic training coefficients for the specified locations.
[0084] The lighting estimation system 110 then modifies the network parameters of the local lighting estimation neural network 406 based on the loss from the loss function 424 to compare the additional location-specific spherical harmonic training coefficients with the additional basis real spherical harmonic coefficients for the specified location in the additional digital training scene. In some cases, the lighting estimation system 110 performs training iterations until the values or weights of the local lighting estimation neural network 406 do not change significantly between training iterations or a convergence criterion is met.
[0085] To reach a convergence point, the lighting estimation system 110 is optionally trained using mini-batches of 20 digital training scenes and the Adam optimizer. Figure 4A or Figure 5A In the illustrated local illumination estimation neural network 406 (the latter is described below), the optimizer minimizes the loss during training iterations with a learning rate of 0.0001 and a weight decay of 0.0002. During some such training experiments, the illumination estimation system 110 reaches convergence more quickly when generating position-specific spherical harmonic training coefficients with a degree greater than zero (e.g., five).
[0086] The lighting estimation system 110 also uses the trained local lighting estimation neural network to generate location-specific lighting parameters. Figure 4B An example of such an application is depicted in FIG. Figure 4BAs shown, the lighting estimation system 110 identifies a request to render a virtual object 432 at a specified location in a digital scene 428. The lighting estimation system 110 uses the local lighting estimation neural network 406 to analyze both the digital scene 428 and the local position indicator 430 to generate position-specific spherical harmonic coefficients 440 for the specified location. Based on the request, the lighting estimation system 110 renders a modified digital scene 442 that includes the virtual object 432 illuminated at the specified location according to the position-specific spherical harmonic coefficients 440.
[0087] As just noted, the lighting estimation system 110 identifies a request to draw a virtual object 432 at a specified location within the digital scene 428. For example, the lighting estimation system 110 may identify a digital request from a computing device executing an augmented reality application to draw a virtual head accessory (or other virtual item) at a specific location on a person (or another real item) depicted in the digital scene 428. Figure 4B As shown, the request to draw a virtual object 432 includes a local position indicator 430 specifying a location for drawing the virtual object 432 .
[0088] Based on the received Figure 4B 4. As shown in the request, the lighting estimation system 110 inputs the digital scene 428 into the local lighting estimation neural network 406. In particular, the lighting estimation system 110 feeds the digital scene 428 to the first set of network layers 408, for the first set of network layers 408 to extract a global feature map 434 from the digital scene 428. As described above, the first set of network layers 408 optionally includes layers from DenseNet and outputs a dense feature map as the global feature map 434. Alternatively, the first set of network layers 408 optionally includes an encoder and outputs an encoded feature map as the global feature map 434.
[0089] like Figure 4B As further shown, the lighting estimation system 110 generates a local position indicator 430 for a specified position within the digital scene 428. For example, in some embodiments, the lighting estimation system 110 identifies the local position indicator 430 from a request to draw a virtual object 432 (e.g., using the local position indicator as a two-dimensional or three-dimensional coordinate in the digital scene 428). The lighting estimation system 110 then modifies the global feature map 434 based on the local position indicator 430.
[0090] To make such modifications, the lighting estimation system 110 can mask the global feature map 434 using the local position indicators 430. In some implementations, for example, the lighting estimation system 110 generates a masked feature map from the local position indicators 430, such as by applying a vector encoder to the local position indicators 430 (e.g., by one-hot encoding). As described above, the masked feature map can include an array of values (e.g., 1s and 0s) of the local position indicators 430 that indicate a specified location within the digital scene 428.
[0091] like Figure 4B As further indicated, the lighting estimation system 110 multiplies the global feature map 434 and the masked feature map for the local location indicator 430 to generate a masked dense feature map 436. Assuming that the masked feature map and the global feature map 434 include the same spatial resolution, in some embodiments, the masked feature map effectively masks the global feature map 434 to create the masked dense feature map 436. The masked dense feature map 436 accordingly represents the local feature map for the specified location within the digital scene 428.
[0092] In some embodiments, when generating the masked dense feature map 436, the lighting estimation system 110 concatenates the global feature map 434 and the masked dense feature map 436 to form a combined feature map 438. To form the combined feature map 438, the lighting estimation system 110 can use any of the concatenation methods described above. The lighting estimation system 110 then feeds the combined feature map 438 to the second set of network layers 416.
[0093] By passing the combined feature map 438 through the second set of network layers 416, the local illumination estimation neural network 406 outputs position-specific spherical harmonic coefficients 440. Consistent with the above disclosure, the position-specific spherical harmonic coefficients 440 indicate the lighting conditions of a specified location within the digital scene 428, such as the specified location identified by the local location indicator 430.
[0094] After generating such lighting parameters, the lighting estimation system 110 renders the modified digital scene 442 including the virtual object 432, which is illuminated at the specified location according to the location-specific spherical harmonic coefficients 440. For example, in some embodiments, the lighting estimation system 110 overlays or otherwise integrates a computer-generated image of the virtual object 432 within the digital scene 428. As part of the rendering, the lighting estimation system 110 selects and renders pixels of the virtual object 432 that reflect the lighting, shading, or appropriate hue indicated by the location-specific spherical harmonic coefficients 440.
[0095] As mentioned above, Figure 5A and Figure 5BAnother embodiment of an illumination estimation system 110 is depicted that trains and applies a local illumination estimation neural network to generate location-specific illumination parameters. Figure 5A As shown, for example, the illumination estimation system 110 iteratively trains the local illumination estimation neural network 504. As an overview of the training iterations, the illumination estimation system 110 extracts a global feature training map from the digital training scene using a first set of network layers 506 of the local illumination estimation neural network 504. The illumination estimation system 110 also generates a plurality of local location training indicators for specified locations within the digital training scene based on the feature maps output by each layer of the first set of network layers. The illumination estimation system 110 then modifies the global feature training map based on the local location training indicators.
[0096] Based on the combined feature training map obtained by modifying the global feature training map, the illumination estimation system 110 generates position-specific spherical harmonic coefficients for the specified location using the second set of network layers 516 of the local illumination estimation neural network 504. The illumination estimation system 110 then modifies the network parameters of the local illumination estimation neural network 504 based on a comparison of the position-specific spherical harmonic coefficients for the specified location within the digital training scene with the underlying true spherical harmonic coefficients.
[0097] For example, Figure 5A As shown, the illumination estimation system 110 feeds a digital training scene 502 to a local illumination estimation neural network 504. After receiving the digital training scene 504 as training input, a first set of network layers 506 extracts a global feature training map 508 from the digital training scene 502. In some embodiments, the first set of network layers 506 includes layers from a DenseNet and outputs the global feature training map 508 in the form of a dense feature map, such as the output of the various dense blocks described by Huang. In some such implementations, the first set of network layers 506 also includes several convolutional layers and max pooling layers, which reduce the resolution of the global feature training map 508. In an alternative to the DenseNet layers, the first set of network layers 506 can include an encoder and output the global feature training map 508 in the form of an encoded feature map, such as the output of the encoder described by Gardner. Regardless of its form, the global feature training map 508 can represent visual features of the digital training scene 502, such as the location of the light sources and the global geometry of the digital training scene 502.
[0098] In addition to generating the global feature training map 508, the lighting estimation system 110 also identifies feature training maps from each of the individual layers of the first set of network layers 506. Figure 5AAs shown, for example, the layers from the first set of network layers 506 collectively output a feature training map 510 that includes each such feature training map. The lighting estimation system 110 selects a training pixel from each feature training map of the feature training maps 510 that corresponds to a specified location within the digital training scene 502. For example, the lighting estimation system 110 optionally selects the coordinates of each training pixel corresponding to the specified location from each feature training map. Each training pixel (or corresponding coordinates) represents a local position training indicator for the specified location.
[0099] like Figure 5A As further shown, illumination estimation system 110 generates a hypercolumn training map 512 based on the selected training pixels from feature training map 510. For example, in some embodiments, illumination estimation system 110 combines or concatenates each selected training pixel, such as by concatenating the features of each selected pixel, to form hypercolumn training map 512. Thus, hypercolumn training map 512 represents the local features of a specified location within digital training scene 502. In some such embodiments, the lighting estimation system 110 generates a hypercolumn map, as described in Aayush Bansal et al., “Marr Revisited: 2D-3D Model Alignment via Surface Normal Prediction,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2016), or in Bharath Hariharan et al., “Hypercolumns for Object Segmentation and Fine-Grained Localization,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2015), both of which are incorporated herein by reference in their entireties.
[0100] In some embodiments, when generating the hypercolumn training map 512, the lighting estimation system 110 concatenates the global feature training map 508 and the hypercolumn training map 512 to form a combined feature training map 514. To form the combined feature training map 514, the lighting estimation system 110 optionally (i) couples the global feature training map 508 and the hypercolumn training map 512 together to form a dual or stacked feature map, or (ii) combines rows of values from the global feature training map 508 with rows of values from the hypercolumn training map 512. However, any suitable concatenation method may be used.
[0101] like Figure 5AAs further shown, the lighting estimation system 110 feeds the combined feature training map 514 to a second set of network layers 516 of the local lighting estimation neural network 504. In some implementations, for example, the second set of network layers 516 includes fully connected layers. As part of a training iteration with such layers, the lighting estimation system 110 optionally passes the combined feature training map 514 through several fully connected layers with batch normalization and ELU learning functions.
[0102] After passing the combined feature training map 514 through the second set of network layers 516, the local illumination estimation neural network 504 outputs position-specific spherical harmonic training coefficients 518. Consistent with the disclosure above, the position-specific spherical harmonic training coefficients 518 indicate the lighting conditions for a specified location within the digital training scene 502. For example, the position-specific spherical harmonic training coefficients 518 indicate the lighting conditions for a specified location within the digital training scene 502 identified by the local location training indicator from the feature training map 510.
[0103] After generating the position-specific spherical harmonic training coefficients 518, the illumination estimation system 110 compares the position-specific spherical harmonic training coefficients 518 with the base real spherical harmonic coefficients 522. The base real spherical harmonic coefficients 522 represent the spherical harmonic coefficients projected from the cube map corresponding to the sample location in the digital training scene 502, i.e., the same specified location indicated by the local location training indicator from the feature training map 510.
[0104] like Figure 5A As further shown, the illumination estimation system 110 uses a loss function 520 to compare the position-specific spherical harmonic training coefficients 518 with the basis true spherical harmonic coefficients 522. In some embodiments, the illumination estimation system 110 uses an MSE function as the loss function 520. Alternatively, in some implementations, the illumination estimation system 110 uses an L2 loss function, a mean absolute error function, a mean absolute percentage error function, a rendering loss function, a root mean square error function, or another suitable loss function as the loss function 520.
[0105] After determining the loss from the loss function 520, the lighting estimation system 110 modifies network parameters (e.g., weights or values) of the local lighting estimation neural network 504 to reduce the loss of the loss function 520 in subsequent training iterations. For example, the lighting estimation system 110 may increase or decrease weights or values from some (or all) of the first set of network layers 506 or the second set of network layers 516 within the local lighting estimation neural network 504 to reduce or minimize the loss in subsequent training iterations.
[0106] After modifying the network parameters of the local illumination estimation neural network 504 for the initial training iteration, the illumination estimation system 110 can perform additional training iterations. In subsequent training iterations, for example, the illumination estimation system 110 extracts additional global feature training maps for additional digital training scenes, generates additional local position training indicators for specified locations within the additional digital training scenes, and modifies the additional global feature training maps based on the additional local position training indicators. Based on the additional combined feature training maps, the illumination estimation system 110 generates additional position-specific spherical harmonic training coefficients for the specified locations.
[0107] Illumination estimation system 110 then modifies network parameters of local illumination estimation neural network 504 based on a loss from loss function 520 that compares the additional location-specific spherical harmonic training coefficients with the additional basis true spherical harmonic coefficients for the specified location within the additional digital training scene. In some cases, illumination estimation system 110 performs training iterations until the values or weights of local illumination estimation neural network 504 do not change significantly during the training iterations or a convergence criterion is met.
[0108] Figure 5B An example of generating location-specific lighting parameters using a local lighting estimation neural network 504 is depicted. Figure 5B As shown, the lighting estimation system 110 identifies a request to render a virtual object 526 at a specified location within the digital scene 524. The lighting estimation system 110 uses the local lighting estimation neural network 504 to analyze both the digital scene 524 and the plurality of local position indicators to generate position-specific spherical harmonic coefficients 536 for the specified location. Based on the request, the lighting estimation system 110 renders a modified digital scene 538 that includes the virtual object 526 illuminated at the specified location according to the position-specific spherical harmonic coefficients 536.
[0109] As just noted, the lighting estimation system 110 identifies a request to draw a virtual object 526 at a specified location within the digital scene 524. For example, the lighting estimation system 110 may identify a digital request from a computing device to draw a virtual animal or character (or other virtual item) at a specific location on a landscape (or another real item) depicted in the digital scene 524. Figure 5B Not shown, but the request to draw virtual object 526 includes an indication of a specified location within digital scene 524 at which to draw virtual object 526.
[0110] Based on the received Figure 5B, the lighting estimation system 110 inputs the digital scene 524 into the local lighting estimation neural network 504. In particular, the lighting estimation system 110 feeds the digital scene 524 to the first set of network layers 506 to extract a global feature map 528 from the digital scene 524. As described above, the first set of network layers 506 optionally includes layers from DenseNet and outputs a dense feature map as the global feature map 528. Alternatively, the first set of network layers 506 optionally includes an encoder and outputs an encoded feature map as the global feature map 528.
[0111] like Figure 5B As further shown, the lighting estimation system 110 identifies a feature map from each of the layers of the first set of network layers 506. Figure 5B As shown, for example, the various layers from the first set of network layers 506 collectively output feature maps 530, including each such feature map. The lighting estimation system 110 selects a pixel corresponding to a specified location within the digital scene 524 from each feature map of the feature maps 510. For example, the lighting estimation system 110 selects coordinates for each pixel corresponding to the specified location from each feature map.
[0112] like Figure 5B As further shown, the lighting estimation system 110 generates a hypermap 532 based on the pixels selected from the feature map 530. For example, in some embodiments, the lighting estimation system 110 combines or concatenates each selected pixel, such as by concatenating the features from each selected pixel, to form the hypermap 532. The hypermap 532 accordingly represents the local features of the specified location within the digital scene 524.
[0113] In some embodiments, when generating the hypercolumn graph 532, the lighting estimation system 110 concatenates the global feature map 528 and the hypercolumn graph 532 to form a combined feature map 534. To form the combined feature map 534, the lighting estimation system 110 can use any of the concatenation methods described above. The lighting estimation system 110 then feeds the combined feature map 534 to the second set of network layers 516.
[0114] By passing the combined feature map 534 through the second set of network layers 516, the local illumination estimation neural network 504 outputs position-specific spherical harmonic coefficients 536. Consistent with the above disclosure, the position-specific spherical harmonic coefficients 536 indicate the lighting conditions at a specified location within the digital scene 524, such as the specified location identified by the local location indicator from the feature map 530.
[0115] After generating such lighting parameters, the lighting estimation system 110 renders the modified digital scene 538 including the virtual object 526, which is illuminated at the specified location according to the location-specific spherical harmonic coefficients 536. For example, in some embodiments, the lighting estimation system 110 overlays or otherwise integrates a computer-generated image of the virtual object 526 within the digital scene 524. As part of the rendering, the lighting estimation system 110 selects and renders pixels of the virtual object 526 that reflect the lighting, shading, or appropriate hue indicated by the location-specific spherical harmonic coefficients 536.
[0116] In addition to accurately depicting the lighting conditions at a given location, Figure 4B and Figure 5B The position-specific spherical harmonic coefficients generated in can dynamically capture the lighting from different perspectives of a specified location within a digital scene. In some embodiments, when the perspective of the digital scene changes in camera viewpoint, model orientation, or other perspective adjustments, the position-specific spherical harmonic coefficients can accurately indicate the lighting conditions at the specified location despite such perspective changes. For example, in some embodiments, the lighting estimation system 110 identifies a perspective adjustment request to draw the digital scene from a different viewpoint, such as by detecting movement of a mobile device that redirects the digital scene or identifying user input that modifies the perspective of the digital scene (e.g., camera movement that adjusts the perspective). Based on the perspective adjustment request, the lighting estimation system 110 can draw a modified digital scene from a different viewpoint, the scene including the scene according to Figure 4B or Figure 5B The position-specific spherical harmonic coefficients generated in the virtual object are illuminated at the specified position.
[0117] Alternatively, in some implementations, the lighting estimation system 110 adjusts or generates new location-specific lighting parameters in response to a perspective adjustment of the digital scene and a corresponding viewpoint change (e.g., a camera movement to adjust the perspective). For example, in some embodiments, the lighting estimation system 110 identifies a perspective adjustment request to draw a virtual object at a specified location within the digital scene from a new or different viewpoint. Based on such a perspective adjustment request, the lighting estimation system 110 may generate the same lighting parameters as above. Figure 4B or Figure 5B New consistent location-specific lighting parameters.
[0118] In some cases, for example, the lighting estimation system 110 generates new local position indicators for a specified position within the digital scene from a different viewpoint (e.g., as Figure 4B The new coordinates for the newly specified location are as shown, or as Figure 5BThe illumination estimation system 110 then modifies the global feature map of the digital scene based on the new local position indicators for the designated positions from the different viewpoints to form a new modified global feature map (e.g., concatenating the new global feature map and the new masked dense feature map to form a new combined feature map, as shown in FIG. 1 ). Figure 4B As shown, or concatenate the new global feature map and the new hypercolumn map to form a new combined feature map, such as Figure 5B Based on the new modified global feature map, the lighting estimation system 110 uses a second set of network layers (e.g., Figure 4B The second set of network layers 416 or Figure 5B The second set of network layers 516 in the
[0065] network layer generates new location-specific lighting parameters for the specified location from different viewpoints. In response to the viewpoint adjustment request, the lighting estimation system 110 accordingly renders the adjusted digital scene, which includes virtual objects illuminated at the specified location according to the new location-specific lighting parameters.
[0119] Apart from Figure 4B and Figure 5B In addition to the position-specific spherical harmonic coefficients and modified digital scene shown in , the lighting estimation system 110 can generate new position-specific spherical harmonic coefficients and adjusted digital scene in response to a position adjustment request to render the virtual object at the new specified position. Figure 4B The local illumination estimation neural network 406 of FIG, the illumination estimation system 110 can identify a new local position indicator from the position adjustment request and modify the global feature map based on the new local position indicator to generate a new masked dense feature map. Figure 5B With the local illumination estimation neural network 504, the illumination estimation system 110 may identify a plurality of new local position indicators from the feature maps output by the first set of network layers 506 and generate a hypercolumn graph based on pixels selected from the feature maps.
[0120] Use from Figure 4B or Figure 5B Using a neural network architecture, the lighting estimation system 110 can generate a new combined feature map and use a second set of network layers to generate new position-specific spherical harmonic coefficients. Based on the position adjustment request, the lighting estimation system 110 can then render an adjusted digital scene including a virtual object, illuminated at the newly specified position according to the new position-specific spherical harmonic coefficients. In response to the position adjustment request, the lighting estimation system 110 can adapt the lighting conditions to the different position to which the computing device has moved the virtual object.
[0121] In addition to updating the position-specific spherical harmonic coefficients and the digital scene in response to position adjustments, the lighting estimation system 110 can generate new position-specific spherical harmonic coefficients and an adjusted scene in response to changes or adjustments in lighting conditions, movement of objects in the scene, or other changes to the scene. Figure 4B The lighting estimation system 110 can extract a new global feature map from the scene and modify the new global feature map based on the local position indicators to generate a new masked dense feature map. Figure 5B By using the first set of network layers 506 , the lighting estimation system 110 may generate a new global feature map and a new hypercolumn map from pixels selected from the new global feature map.
[0122] Use from Figure 4B or Figure 5B Using a neural network architecture, the lighting estimation system 110 can generate a new combined feature map and use a second set of network layers to generate new location-specific spherical harmonic coefficients that reflect the adjusted lighting conditions. Thus, based on the adjustment of the lighting conditions for a specified location within the digital scene, the lighting estimation system 110 can render the adjusted digital scene including a virtual object, which is illuminated at the specified location according to the new location-specific spherical harmonic coefficients. In response to the adjustment of the lighting conditions, the lighting estimation system 110 can adapt the location-specific lighting parameters accordingly to reflect the adjustment of the lighting conditions.
[0123] Figures 6A-6C An embodiment of a lighting estimation system 110 is depicted in which a computing device, in response to different rendering requests, renders a digital scene including a virtual object at a specified location according to location-specific lighting parameters and new location-specific lighting parameters, respectively, and then renders a digital scene including a virtual object at the new specified location. As an overview, Figures 6A-6C Each depicts a computing device 600 including an augmented reality application for the augmented reality system 108 and the lighting estimation system 110. The augmented reality application includes causing the computing device 600 to execute Figures 6A-6C Computer-executable instructions for certain actions depicted in .
[0124] Rather than repeatedly describing computer-executable instructions in an augmented reality application as causing the computing device 600 to perform such actions, the present disclosure primarily describes the computing device 600 or the lighting estimation system 110 as shorthand for performing an action. Figures 6A-6C Various user interactions indicated, such as when computing device 600 detects user selection of a virtual object. Figures 6A-6C, but computing device 600 may alternatively be any type of computing device, such as a desktop, laptop, or tablet, and may also detect any appropriate user interaction, including but not limited to audio input from a microphone, gaming device button input, keyboard input, mouse clicks, stylus interaction with a touch screen, or touch gestures on a touch screen.
[0125] Now back Figure 6A , which depicts a computing device 600 presenting a graphical user interface 606a that includes a digital scene 610 within a screen 602. As shown, the digital scene 610 includes real objects 612a and 612b. The graphical user interface 606a also includes a selectable options bar 604 for virtual objects 608a-608c. By presenting the graphical user interface 606a, the computing device 600 provides a user with the option to request the lighting estimation system 110 to draw one or more virtual objects 608a-608c at a specified location within the digital scene 610.
[0126] For example, Figure 6A As shown, the computing device 600 detects a user interaction requesting the lighting estimation system 110 to draw a virtual object 608a at a specified location 614 within the digital scene 610. In particular, Figure 6A The computing device 600 is depicted detecting a drag-and-drop gesture to move the virtual object 608a to a designated location 614. Figure 6A A user using a drag-and-drop gesture is illustrated, but computing device 600 may detect any suitable user interaction requesting lighting estimation system 110 to draw a virtual object within a digital scene.
[0127] Based on receiving the request for the lighting estimation system 110 to draw the virtual object 608a within the digital scene 610 , the augmented reality system 108 , in conjunction with the lighting estimation system 110 , draws the virtual object 608a at the designated location 614 . Figure 6B An example of such a rendering is depicted. As shown, computing device 600 presents a graphical user interface 606b including a modified digital scene 616 within screen 602. Consistent with the above disclosure, computing device 600 renders modified digital scene 616 including virtual object 608a, which is illuminated at designated location 614 according to location-specific lighting parameters generated by lighting estimation system 110.
[0128] To generate such location-specific lighting parameters, the lighting estimation system 110 optionally performs Figure 4B or Figure 5B The action shown. Figure 6B, the location-specific lighting parameters indicate realistic lighting conditions for virtual object 608a, where lighting and shadows are consistent with real objects 612a and 612b. Shadows of virtual object 608a and real objects 612a and 612b always reflect light from light sources outside the perspective shown in modified digital scene 616.
[0129] As described above, the lighting estimation system 110 may generate new location-specific lighting parameters and an adjusted digital scene in response to a location adjustment request to render the virtual object at the new specified location. Figure 6C An example of such an adjusted digital scene reflecting the new location-specific lighting parameters is depicted. Figure 6C As shown, computing device 600 presents graphical user interface 606c including adjusted digital scene 618 within screen 602. Computing device 600 detects a user interaction including a position adjustment request to move virtual object 608a from designated position 614 to a new designated position 620.
[0130] Upon receiving the request from the lighting estimation system 110 to move the virtual object 608a, the augmented reality system 108, in conjunction with the lighting estimation system 110, draws the virtual object 608a at the new designated location 620. Figure 6C The computing device 600 is depicted rendering an adjusted digital scene 618 including a virtual object 608 a illuminated at a new designated location 620 according to new location-specific lighting parameters generated by the lighting estimation system 110 .
[0131] To generate such new location-specific lighting parameters, the lighting estimation system 110 optionally modifies the global feature map and uses a lighting estimation neural network to generate the following: Figure 4B or Figure 5B The position-specific spherical harmonic coefficients are shown in . Figure 6C As shown, the location-specific lighting parameters indicate realistic lighting conditions for the virtual object 608a, where the virtual object 608a has adjusted lighting and shadows consistent with the real objects 612a and 612b. Figures 6B to 6C As shown in the transition of , the lighting estimation system 110 can adapt the lighting conditions to different positions in real time (or near real time) in response to a position adjustment request to move a virtual object in the virtual scene.
[0132] As described above, the lighting estimation system 110 can generate location-specific lighting parameters that indicate accurate and realistic lighting conditions for locations within a digital scene. To test the accuracy and realism of the lighting estimation system 110, the researchers modified digital scenes from the SUNCG dataset (described above) and applied a local lighting estimation neural network to generate location-specific lighting parameters for various locations in such digital scenes. Figure 7A and Figure 7B Examples of such accuracy and realism in various renderings of digital scenes with locations illuminated according to base real lighting parameters from the lighting estimation system 110 and locations illuminated according to location-specific lighting parameters from the lighting estimation system 110 are illustrated.
[0133] for Figure 7A and Figure 7B In both cases, the researchers trained Figure 4A The local illumination estimation neural network depicted, where the local illumination estimation neural network includes DenseNet blocks from the first set of network layers of DenseNet120 and initializes the network parameters using weights trained on ImageNet. Figure 7A and Figure 7B In both cases, researchers based Figure 4B The shown actions further generate position-specific spherical harmonic coefficients by applying a trained local illumination estimation neural network.
[0134] For example, Figure 7A As shown, the researchers modified a digital scene 702 from the SUNCG dataset. The digital scene 702 includes a designated location 704 indicating a designated location that serves as a target for estimating lighting conditions. For comparison purposes, the researchers plotted a red, green, and blue ("RBG") representation 706 and a light intensity representation 710 of the designated location based on the underlying real spherical harmonic coefficients. Consistent with the disclosure above, the researchers projected the underlying real spherical harmonic coefficients of the designated location 704 from a cubemap. The researchers further used the lighting estimation system 110 to generate location-specific spherical harmonic coefficients for the designated location 704. After generating such lighting parameters, the augmented reality system 108, in conjunction with the lighting estimation system 110, plotted the RBG representation 708 and the light intensity representation 712 of the designated location 704 based on the location-specific spherical harmonic coefficients.
[0135] As a pair Figure 7A As shown in FIG1 , the comparison of the RGB representation and the light intensity representation indicates that the lighting estimation system 110 generates location-specific spherical harmonic coefficients that accurately and realistically estimate the lighting conditions of an object located at a specified location 704. Unlike some conventional augmented reality systems, the lighting estimation system 110 accurately estimates the lighting emitted by light sources outside the viewpoint of the digital scene 702. Even though the strongest light source is from behind the camera capturing the digital scene 702, the lighting estimation system 110 estimates the lighting conditions shown in the RBG representation 708 and the light intensity representation 712 with similar accuracy and realism to those shown in the RBG representation 706 and light intensity representation 710, which reflect the ground truth.
[0136] Figure 7B A modified digital scene 714 is illustrated, which includes virtual objects 718a-718d at specified locations that are illuminated according to the underlying real spherical harmonic coefficients. Consistent with the above disclosure, the researchers projected the underlying real spherical harmonic coefficients from the cube map for the specified locations in the modified digital scene 714. Figure 7B Further illustrated is a modified digital scene 716 that includes virtual objects 720a-720d at specified locations illuminated according to location-specific spherical harmonic coefficients generated by the lighting estimation system 110. To determine estimated lighting conditions for the virtual objects 720a-720d at the specified locations, the lighting estimation system 110 generates location-specific spherical harmonic coefficients for each specified location in the modified digital scene 716.
[0137] For comparison purposes, the researchers used the lighting estimation system 110 to render the metal spheres for the virtual objects 718a-718d in the modified digital scene 714 and the virtual objects 720a-720d in the modified digital scene 716. Figure 7B As shown, both modified digital scene 714 and modified digital scene 716 include virtual objects 718a-718d and virtual objects 720a-720d, respectively, at the same designated locations.
[0138] As indicated by a comparison of the lighting of the virtual objects in the modified digital scenes 714 and 716, the lighting estimation system 110 generates position-specific spherical harmonic coefficients that accurately and realistically estimate the lighting conditions of the virtual objects 720a-720d at each object's corresponding designated location in the modified digital scene 716. Although the light intensities of the virtual objects 720a-720d differ slightly from the light intensities of the virtual objects 718a-720d, the trained local lighting estimation neural network detects sufficient geometric context from the underlying scene of the modified digital scene 716 to generate coefficients that both (i) darken the occluded metal sphere and (ii) reflect strong directional light on the metal sphere when exposed to light from sources outside the viewing angle of the modified digital scene 716.
[0139] Now go to Figure 8 and Figure 9 ,These figures provide an overview of the environments in which lighting estimation systems can operate and architectural examples of lighting estimation systems. In particular, Figure 8 A block diagram illustrating an exemplary system environment ("environment") 800 in which a lighting estimation system 806 may operate is depicted, according to one or more embodiments. Specifically, Figure 8An environment 800 is illustrated, which includes server(s) 802, third-party server(s) 810, a network 812, a client device 814, and a user 818 associated with the client device 814. Figure 8 One client device and one user are illustrated, but in alternative embodiments, environment 800 may include any number of computing devices and associated users. Figure 8 A particular arrangement of server(s) 802, third-party server(s) 810, network 812, client devices 814, and users 818 is illustrated, but various additional arrangements are possible.
[0140] like Figure 8 As shown, server(s) 802, third party server(s) 810, network 812, and client device 814 may be directly or indirectly communicatively coupled to one another, such as via network 812, as will be discussed below with respect to Figure 12 Further described. The server(s) 802 and client device 814 may include any type of computing device, including one or more computing devices, as described below with respect to Figure 12 Further discussion.
[0141] like Figure 8 As shown, the server(s) 802 can generate, store, receive, and / or transmit any type of data, including user input for inputting a digital scene into a neural network or requesting the drawing of a virtual object to create an augmented reality scene. For example, the server(s) 802 can receive user input from a client device 814 requesting the drawing of a virtual object at a specified location within a digital scene, and then utilize a local illumination estimation neural network to generate location-specific lighting parameters for the specified location. After generating such parameters, the server(s) 802 can further draw a modified digital scene including a virtual object, where the virtual object is illuminated at the specified location according to the location-specific lighting parameters. In some embodiments, the server(s) 802 include a data server, a communication server, or a web hosting server.
[0142] like Figure 8As further shown, the server(s) 802 may include an augmented reality system 804. Generally, the augmented reality system 804 facilitates the generation, modification, sharing, access, storage, and / or deletion of digital content (e.g., a two-dimensional digital image of a scene or a three-dimensional digital model of a scene) in an augmented reality-based image. For example, the augmented reality system 804 may use the server(s) 802 to generate a modified digital image or model including a virtual object or to modify an existing digital scene. In some implementations, the augmented reality system 804 uses the server(s) 802 to receive user input identifying a digital scene, a virtual object, or a specified location within the digital scene from a client device 814, or to transmit data representing the digital scene, the virtual object, or the specified location to the client device 814.
[0143] In addition to augmented reality system 804, server(s) 802 also include an illumination estimation system 806. Illumination estimation system 806 is an embodiment of illumination estimation system 110 described above (and can perform functions, methods, and processes). For example, in some embodiments, illumination estimation system 806 uses server(s) 802 to identify a request to draw a virtual object at a specified location within a digital scene. Illumination estimation system 806 further uses server(s) 802 to extract a global feature map from the digital scene using a first set of network layers of a local illumination estimation neural network. In some implementations, illumination estimation system 806 also uses server(s) 802 to generate a local location indicator for the specified location and modify the global feature map of the digital scene based on the local location indicator. Based on the modified global feature map, illumination estimation system 806 further uses server(s) 802 to (i) generate location-specific lighting parameters for the specified location using a second set of layers of the local illumination estimation neural network, and (ii) draw the modified digital scene, which includes the virtual object at the specified location illuminated according to the location-specific lighting parameters.
[0144] As suggested by the previous embodiments, the lighting estimation system 806 may be implemented in whole or in part by various elements of the environment 800. Figure 8 The lighting estimation system 806 is illustrated as being implemented within the server(s) 802, but components of the lighting estimation system 806 may be implemented in other components of the environment 800. For example, in some embodiments, the client device 814 includes the lighting estimation system 806 and performs all of the functions, methods, and processes of the lighting estimation system 806 described above and below. Figure 9 Components of the lighting estimation system 806 are further described.
[0145] like Figure 8As further shown, in some embodiments, client device 814 comprises a computing device that allows user 818 to send and receive digital communications. For example, client device 814 may comprise a desktop computer, laptop computer, smartphone, tablet computer, or other electronic device. In some embodiments, client device 814 also includes one or more software applications (e.g., augmented reality application 816) that allow user 818 to send and receive digital communications. For example, augmented reality application 816 may be a software application installed on client device 814 or hosted on server(s) 802. When hosted on server(s) 802, augmented reality application 816 may be accessed by client device 814 through another application, such as a web browser. In some implementations, augmented reality application 816 comprises instructions that, when executed by a processor, cause client device 814 to present one or more graphical user interfaces, such as a user interface that includes digital scenes and / or virtual objects, for user 818 to select as input when generating location-specific lighting parameters or a modified digital scene, or for lighting selection system 806 to include as input.
[0146] Also like Figure 8 As shown, the augmented reality system 804 is communicatively coupled to the augmented reality database 808. In one or more embodiments, the augmented reality system 804 accesses and queries the augmented reality database 808 for data associated with the request from the lighting estimation system 806. For example, the augmented reality system 804 can access the digital scene, the virtual object, the specified location within the digital scene, or the location-specific lighting parameters of the lighting estimation system 806. Figure 8 As shown, augmented reality database 808 is maintained separately from server(s) 802. Alternatively, in one or more embodiments, augmented reality system 804 and augmented reality database 808 comprise a single combined system or subsystem within server(s) 802.
[0147] Now go to Figure 9 , which provides additional details regarding the components and features of the lighting estimation system 806. In particular, Figure 9 A computing device 900 is shown that implements the augmented reality system 804 and the lighting estimation system 806. In some embodiments, the computing device 900 includes one or more servers (e.g., server(s) 802). In other embodiments, the computing device 900 includes one or more client devices (e.g., client device 814).
[0148] like Figure 9As shown, computing device 900 includes augmented reality system 804. In some embodiments, augmented reality system 804 uses its components to provide tools for generating digital scenes or other augmented reality-based images or modifying existing digital scenes or other augmented reality-based images in a user interface of augmented reality application 816. Additionally, in some cases, augmented reality system 804 facilitates the generation, modification, sharing, access, storage, and / or deletion of digital content in augmented reality-based images.
[0149] like Figure 9 As further shown, computing device 900 includes lighting estimation system 806. Lighting estimation system 806 includes, but is not limited to, digital scene manager 902, virtual object manager 904, neural network trainer 906, neural network operator 908, augmented reality renderer 910, and / or storage manager 912. The following paragraphs describe each of these components in turn.
[0150] As just mentioned, the lighting estimation system 806 includes a digital scene manager 902. The digital scene manager 902 receives input regarding digital scenes, identifies and analyzes the digital scenes. For example, in some embodiments, the digital scene manager 902 receives user input identifying a digital scene and presents the digital scene from an augmented reality application. Additionally, in some embodiments, the digital scene manager 902 identifies multiple digital scenes for presentation as part of an image sequence (e.g., an augmented reality sequence).
[0151] like Figure 9 As further shown, the virtual object manager 904 receives input regarding virtual objects, identifies, and analyzes the virtual objects. For example, in some embodiments, the virtual object manager 904 receives user input identifying a virtual object and requesting the lighting estimation system 110 to draw the virtual object at a specified location within the digital scene. Additionally, in some embodiments, the virtual object manager 904 provides selectable options for the virtual object, such as those displayed in a user interface of an augmented reality application.
[0152] like Figure 9As further shown, the neural network trainer 906 trains the local illumination estimation neural network 918. For example, in some embodiments, the neural network trainer 906 extracts a global feature training map from the digital training scene using a first set of network layers of the local illumination estimation neural network 918. Additionally, in some embodiments, the neural network trainer 906 generates a local location training indicator for a specified location within the digital training scene and modifies the global feature training map based on the local location training indicator for the specified location. Based on the modified global feature training map, the neural network trainer 906 (i) generates location-specific illumination training parameters for the specified location using a second set of network layers of the local illumination estimation neural network 918, and (ii) modifies network parameters of the local illumination estimation neural network based on a comparison of the location-specific illumination training parameters for the specified location within the digital training scene with the ground truth illumination parameters.
[0153] In some such embodiments, the neural network trainer 906 trains Figure 4A and Figure 5A Local lighting estimation neural network 918 is shown. In some embodiments, neural network trainer 906 also communicates with storage manager 912 to apply and / or access digital training scenes from digital scenes 914, ground truth lighting parameters from location-specific lighting parameters 920, and / or local lighting estimation neural network 918.
[0154] like Figure 9 As further shown, the neural network operator 908 applies a trained version of the local illumination estimation neural network 918. For example, in some embodiments, the neural network operator 908 extracts a global feature map from the digital scene using a first set of network layers of the local illumination estimation neural network 918. The neural network operator 908 also generates a local position indicator for the specified location and modifies the global feature map of the digital scene based on the local position indicator. Based on the modified global feature map, the neural network operator 908 also generates location-specific illumination parameters for the specified location using a second set of layers of the local illumination estimation neural network. In some such embodiments, the neural network operator 908 applies the following methods, respectively: Figure 4B and Figure 5B 918. In some embodiments, neural network operator 908 also communicates with storage manager 912 to apply and / or access digital scenes from digital scenes 914, virtual objects from virtual objects 916, location-specific lighting parameters from location-specific lighting parameters 920, and / or local lighting estimation neural network 918.
[0155] In addition to the neural network operator 908, in some embodiments, the lighting estimation system 806 also includes an augmented reality renderer 910. The augmented reality renderer 910 renders a modified digital scene that includes a virtual object. For example, in some embodiments, based on a request to render a virtual object at a specified location within the digital scene, the augmented reality renderer 910 renders a modified digital scene that includes the virtual object at the specified location illuminated according to the location-specific lighting parameters from the neural network operator 908.
[0156] In one or more embodiments, each component of the lighting estimation system 806 communicates with each other using any suitable communication technology. Additionally, the components of the lighting estimation system 806 can communicate with one or more other devices including one or more client devices described above. Figure 9 Components of the lighting estimation system 806 are shown as separate, but any subcomponents may be combined into fewer components, such as into a single component, or may be divided into more components for a particular implementation. Figure 9 , but at least some of the components used to perform operations in conjunction with the lighting estimation system 806 described herein can be implemented on other devices within the environment 800.
[0157] Each component 902-920 of the lighting estimation system 806 may include software, hardware, or both. For example, the components 902-920 may include one or more instructions stored on a computer-readable storage medium, and the one or more instructions may be executable by a processor of one or more computing devices (such as a client device or a server device). When executed by one or more processors, the computer-executable instructions of the lighting estimation system 806 may cause the computing device(s) to perform the methods described herein. Alternatively, the components 902-920 may include hardware, such as a dedicated processing device that performs a specific function or group of functions. Alternatively, the components 902-920 of the lighting estimation system 806 may include a combination of computer-executable instructions and hardware.
[0158] Furthermore, the components 902-920 of the lighting estimation system 806 can be implemented, for example, as one or more operating systems, as one or more standalone applications, as one or more generators of applications, as one or more plug-ins, as one or more library functions or functions callable by other applications, and / or as a cloud computing model. Thus, the components 902-920 can be implemented as standalone applications, such as desktop or mobile applications. Furthermore, the components 902-920 can be implemented as one or more web-based applications hosted on a remote server. The components 902-920 can also be implemented in a set of mobile device applications or "apps." For illustration, the components 902-920 can be implemented in software applications, including, but not limited to, Adobe ILLUSTRATOR, Adobe EXPERIENCE DESIGN, Adobe CREATIVE CLOUD, Adobe PHOTOSHOP, Project AERO, or Adobe LIGHTROOM. “ADOBE”, “ILLUSTRATOR”, “EXPERIENCE DESIGN”, “CREATIVE CLOUD”, “PHOTOSHOP”, “PROJECT AERO” and “LIGHTROOM” are registered trademarks or trademarks of Adobe Corporation in the United States and / or other countries.
[0159] Now go to Figure 10 , which illustrates a flow chart of a series of actions 1000 for training a local lighting estimation neural network to generate location-specific lighting parameters in accordance with one or more embodiments. Figure 10 The actions according to one embodiment are illustrated, but alternative embodiments may omit, add, reorder, and / or modify Figure 10 Any action shown. Figure 10 Alternatively, a non-transitory computer-readable storage medium may include instructions that, when executed by one or more processors, cause a computing device to perform Figure 10 In other embodiments, the system may perform the actions depicted in Figure 10 action.
[0160] like Figure 10 As shown, act 1000 includes act 1010, which extracts a global feature training map from a digital training scene using a local illumination estimation neural network. In particular, in some embodiments, act 1010 includes extracting a global feature training map from one of the digital training scenes using a first set of network layers of the local illumination estimation neural network. In some embodiments, the digital training scene includes a three-dimensional digital model or a digital viewpoint image of a real-world scene.
[0161] like Figure 10 As further shown, act 1000 includes act 1020 of generating a local position training indicator for a specified location within the digital training scene. For example, in some embodiments, generating the local position training indicator includes identifying local position training coordinates representing the specified location within the digital training scene. In contrast, in some implementations, generating the local position training indicator includes: selecting a first training pixel corresponding to the specified location from a first feature training map corresponding to a first layer of the local illumination estimation neural network; and selecting a second training pixel corresponding to the specified location from a second feature training map corresponding to a second layer of the local illumination estimation neural network.
[0162] like Figure 10 As further shown, action 1000 includes action 1030 of modifying a global feature training map of the digital training scene based on the local position training indicator. For example, in some implementations, action 1030 includes modifying the global feature training map to generate a modified global feature training map by: generating a masked feature training map from the local position training coordinates; multiplying the global feature training map and the masked feature training map for the local position training coordinates to generate a masked dense feature training map; and concatenating the global feature training map and the masked dense feature training map to form a combined feature training map.
[0163] like Figure 10 As further shown in FIG, act 1000 includes act 1040 of generating location-specific lighting training parameters for a specified location using a local lighting estimation neural network based on the modified global feature training map. In particular, in certain implementations, act 1040 includes generating location-specific lighting training parameters for the specified location using a second set of network layers of the local lighting estimation neural network based on the modified global feature training map. In some embodiments, the first set of network layers of the local lighting estimation neural network includes lower layers of a densely connected convolutional network, and the second set of network layers of the local lighting estimation neural network includes convolutional layers and fully connected layers.
[0164] As described above, in some implementations, generating location-specific lighting training parameters for a specified location includes generating location-specific spherical harmonic training coefficients that indicate lighting conditions at the specific location. In some such embodiments, generating the location-specific spherical harmonic training coefficients includes generating the location-specific spherical harmonic training coefficients five times for each color channel.
[0165] like Figure 10As further shown, act 1000 includes act 1050 of modifying network parameters of the local lighting estimation neural network based on the comparison of the location-specific training parameters with the ground truth lighting parameters. In particular, in some embodiments, act 1050 includes modifying network parameters of the local lighting estimation neural network based on the comparison of the location-specific lighting training parameters with a set of ground truth lighting parameters for a specified location within the digital training scene.
[0166] In addition to acts 1010-1050, in some cases, act 1000 further includes determining a set of ground-truth lighting parameters for the specified location by determining a set of ground-truth location-specific spherical harmonic coefficients indicative of lighting conditions at the specified location. Additionally, in one or more embodiments, act 1000 further includes generating location-specific lighting training parameters by providing the combined feature training map to a second set of network layers.
[0167] As described above, in some embodiments, act 1000 further includes determining a set of ground-truth lighting parameters for the specified location by determining a set of ground-truth position-specific spherical harmonic coefficients that indicate lighting conditions at the specified location. In some such implementations, determining the set of ground-truth position-specific spherical harmonic coefficients includes: identifying locations within the digital training scene; generating a cubemap for each location in the digital training scene; and projecting the cubemap for each location in the digital training scene onto the set of ground-truth position-specific spherical harmonic coefficients.
[0168] Now go to Figure 11 , which illustrates a flow diagram of a series of actions 1100 for applying a trained local illumination estimation neural network to generate location-specific illumination parameters in accordance with one or more embodiments. Figure 11 The actions according to one embodiment are illustrated, but alternative embodiments may omit, add, reorder, and / or modify Figure 11 Any action shown. Figure 11 Alternatively, a non-transitory computer-readable storage medium may include instructions that, when executed by one or more processors, cause a computing device to perform Figure 11 In other embodiments, the system may perform the actions depicted in Figure 11 action.
[0169] like Figure 11 As shown, act 1100 includes identifying a request to draw a virtual object at a specified location within a digital scene 1110. For example, in some embodiments, identifying the request includes receiving a request to draw a virtual object at a specified location from a mobile device.
[0170] like Figure 11As further shown, act 1100 includes extracting a global feature map from the digital scene using a local illumination estimation neural network 1120. In particular, in some embodiments, act 1120 includes extracting a global feature map from the digital scene using a first set of network layers of the local illumination estimation neural network.
[0171] like Figure 11 As further shown in , act 1100 includes act 1130 of generating a local position indicator for the specified position within the digital scene. In particular, in some embodiments, act 1130 includes generating a local position indicator for the specified position by identifying local position coordinates representing the specified position within the digital scene.
[0172] In contrast, in some implementations, action 1130 includes generating a local position indicator for the specified position by selecting a first pixel corresponding to the specified position from a first feature map corresponding to a first layer of the first set of network layers; and selecting a second pixel corresponding to the specified position from a second feature map corresponding to a second layer of the first set of network layers.
[0173] like Figure 11 As further shown, action 1100 includes action 1140 of modifying a global feature map of the digital scene based on the local position indicator for the specified position. Specifically, in some embodiments, action 1140 includes modifying the global feature map to generate a modified global feature map by: combining features of a first pixel and a second pixel corresponding to the specified position to generate a hypercolumn map; and concatenating the global feature map and the hypercolumn map to form a combined feature map.
[0174] like Figure 11 As further shown in FIG, act 1100 includes act 1150 of generating location-specific lighting parameters for the specified location using a local illumination estimation neural network based on the modified global feature map. In particular, in some embodiments, act 1140 includes generating location-specific lighting parameters for the specified location using a second set of network layers of the local illumination estimation neural network based on the modified global feature map. In some embodiments, the first set of network layers of the local illumination estimation neural network includes lower layers of a densely connected convolutional network, and the second set of network layers of the local illumination estimation neural network includes convolutional layers and fully connected layers.
[0175] As an example of action 1150, in some embodiments, generating location-specific lighting parameters for the specified location includes generating location-specific spherical harmonic coefficients that indicate lighting conditions of an object at the specific location. As another example, in some implementations, generating the location-specific lighting parameters includes providing the combined feature map to the second set of network layers.
[0176] like Figure 11 As further shown, act 1100 includes act 1160 of drawing a modified digital scene including a virtual object, the virtual object being illuminated at the specified location according to the location-specific lighting parameters. In particular, in some embodiments, act 1160 includes drawing, based on the request, a modified digital scene including the virtual object, the virtual object being illuminated at the specified location according to the location-specific lighting parameters. For example, in some cases, drawing the modified digital scene includes drawing, based on receiving the request from the mobile device, within a graphical user interface of the mobile device, the modified digital scene including the virtual object, the virtual object being illuminated at the specified location according to the location-specific lighting parameters.
[0177] In addition to actions 1110-1160, in some implementations, action 1100 also includes identifying a position adjustment request to move the virtual object from a specified position within the digital scene to a new specified position within the digital scene; generating a new local position indicator for the new specified position within the digital scene; modifying a global feature map of the digital scene based on the new local position indicator for the new specified position to form a new modified global feature map; generating new position-specific lighting parameters for the new specified position using a second set of network layers based on the new modified global feature map; and based on the position adjustment request, drawing the adjusted digital scene including the virtual object, wherein the virtual object is illuminated according to the new position-specific lighting parameters at the new specified position.
[0178] As suggested above, in one or more embodiments, action 1100 further includes: identifying a perspective adjustment request to render the digital scene from a different viewpoint; and based on the perspective adjustment request, rendering a modified digital scene from the different viewpoint, the modified digital scene including a virtual object illuminated at a specified location according to location-specific lighting parameters.
[0179] In addition, in some cases, action 1100 also includes identifying a perspective adjustment request from a different viewpoint to draw the virtual object at a specified location within the digital scene; generating a new local position indicator for the specified location within the digital scene from the different viewpoint; modifying a global feature map of the digital scene based on the new local position indicator for the specified location from the different viewpoint to form a new modified global feature map; generating new location-specific lighting parameters for the specified location from the different viewpoint using a second set of network layers based on the new modified global feature map; and drawing the adjusted digital scene based on the adjustment of the lighting conditions, the scene including the virtual object illuminated at the specified location according to the new location-specific lighting parameters.
[0180] Additionally, in some implementations, action 1100 further includes: identifying an adjustment to a lighting condition for a specified location within the digital scene; extracting a new global feature map from the digital scene using a first set of network layers of a local illumination estimation neural network; modifying the new global feature map of the digital scene based on a local position indicator of the specified location; generating new location-specific lighting parameters for the specified location using a second set of network layers based on the new modified global feature map; and rendering the adjusted digital scene based on the adjustment to the lighting condition, the scene including a virtual object illuminated at the specified location according to the new location-specific lighting parameters.
[0181] In addition to (or as an alternative to) the above actions, in some embodiments, action 1000 (or action 1100) includes a step for training a local illumination estimation neural network using a global feature training map for a digital training scene and a local position training indicator for a specified position within the digital training scene. Figure 4A or Figure 5A The described algorithms and actions may include corresponding actions for performing steps for training a local illumination estimation neural network using a global feature training map for a digital training scene and local location training indicators for specified locations within the digital training scene.
[0182] Additionally or alternatively, in some embodiments, action 1000 (or action 1100) includes a step for generating location-specific lighting parameters for a specified location by utilizing a trained local lighting estimation neural network. Figure 4B or Figure 5B The described algorithms and acts may include corresponding acts for performing the following steps: generating location-specific lighting parameters for a specified location by utilizing a trained local lighting estimation neural network.
[0183] Embodiments of the present disclosure may include or utilize a special-purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in more detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium, and the instructions are executable by one or more computing devices (e.g., any media content access device described herein). Typically, a processor (e.g., a microprocessor) receives instructions from a non-transitory computer-readable medium (e.g., a memory, etc.) and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.
[0184] Computer-readable media can be any available medium that can be accessed by a general-purpose or special-purpose computer system. A computer-readable medium that stores computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium that carries computer-executable instructions is a transmission medium. Thus, by way of example and not limitation, embodiments of the present disclosure may include at least two distinct types of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0185] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drives (“SSD”) (e.g., RAM-based), flash memory, phase-change memory (“PCM”), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code means in the form of computer-executable instructions or data structures that can be accessed by a general-purpose or special-purpose computer.
[0186] A "network" is defined as one or more data links that enable the transport of electronic data between computer systems and / or generators and / or other electronic devices. When information is transmitted or provided to a computer via a network or other communication connection (hardwired, wireless, or a combination of hardwired or wireless), the computer properly views the connection as a transmission medium. Transmission media may include networks and / or data links that can be used to carry desired program code components in the form of computer-executable instructions or data structures, and the desired program code components can be accessed by general-purpose or special-purpose computers. Combinations of the foregoing should also be included within the scope of computer-readable media.
[0187] Furthermore, upon reaching various computer system components, program code components in the form of computer-executable instructions or data structures can be automatically transferred from a transmission medium to a non-transitory computer-readable storage medium (device) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface generator (e.g., a "NIC") and then ultimately transferred to computer system RAM and / or a less volatile computer storage medium (device) within the computer system. Thus, it should be understood that a non-transitory computer-readable storage medium (device) can be included in computer system components that also (or even primarily) utilize a transmission medium.
[0188] Computer-executable instructions include, for example, instructions and data that, when executed on a processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a specific function or set of functions. In one or more embodiments, computer-executable instructions are executed on a general-purpose computer to transform the general-purpose computer into a special-purpose computer that implements the elements of the present disclosure. Computer-executable instructions can be, for example, binary files, intermediate format instructions (such as assembly language), or even source code. Although the subject matter has been described in language specific to structural marketing features and / or methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the marketing features or actions described above. Rather, the described marketing features and actions are disclosed as example forms of implementing the claims.
[0189] Those skilled in the art will appreciate that the present disclosure can be practiced in a network computing environment with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, etc. The present disclosure can also be practiced in a distributed system environment in which local and remote computer systems linked by a network (by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) all perform tasks. In a distributed system environment, the program generator can be located in local and remote memory devices.
[0190] Embodiments of the present disclosure may also be implemented in a cloud computing environment. In this specification, "cloud computing" is defined as a subscription model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be adopted in the market to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. This shared pool of configurable computing resources can be rapidly provisioned via virtualization, released with minimal management effort or service provider interaction, and then scaled accordingly.
[0191] A cloud computing subscription model can be comprised of various features, such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, metered services, and the like. A cloud computing subscription model can also expose various service subscription models, such as, for example, software as a service ("SaaS"), web services, platform as a service ("PaaS"), and infrastructure as a service ("IaaS"). A cloud computing subscription model can also be deployed using different deployment subscription models, such as private cloud, community cloud, public cloud, hybrid cloud, and the like. In this specification and claims, a "cloud computing environment" is an environment in which cloud computing is employed.
[0192] Figure 12 A block diagram of an exemplary computing device 1200 that may be configured to perform one or more of the above-described processes is shown. Figure 12 As shown, computing device 1200 may include a processor 1202, a memory 1204, a storage device 1206, an I / O interface 1208, and a communication interface 1210, which may be communicatively coupled via a communication infrastructure 1212. In some embodiments, Figure 12 The computing device 1200 may include fewer or more components than those shown. Figure 12 Components of computing device 1200 are shown.
[0193] In one or more embodiments, the processor 1202 includes hardware for executing instructions, such as those comprising a computer program. By way of example and not limitation, to execute instructions for digitizing a real-world object, the processor 1202 may retrieve (or fetch) instructions from internal registers, internal cache memory, memory 1204, or storage device 1206, and decode and execute them. The memory 1204 may be volatile or non-volatile memory for storing data, metadata, and programs for execution by the processor(s). The storage device 1206 includes a storage device such as a hard disk, a flash disk drive, or other digital storage device for storing data or instructions related to the object digitization process (e.g., a digital scan, a digital model).
[0194] The I / O interface 1208 allows a user to provide input to the computing device 1200, receive output from the computing device 1200, and otherwise transfer data to and receive data from the computing device 1200. The I / O interface 1208 may include a mouse, a keypad or keyboard, a touch screen, a camera, an optical scanner, a network interface, a modem, other known I / O devices, or a combination of such I / O interfaces. The I / O interface 1208 may include one or more devices for presenting output to the user, including but not limited to a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In some embodiments, the I / O interface 1208 is configured to provide graphical data to the display for presentation to the user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may serve a particular implementation.
[0195] The communication interface 1210 may include hardware, software, or both. In any case, the communication interface 1210 may provide one or more interfaces for communication between the computing device 1200 and one or more other computing devices or networks (such as, for example, packet-based communication). By way of example and not limitation, the communication interface 1210 may include a network interface controller ("NIC") or network adapter for communicating with an Ethernet or other wired-based network, or a wireless NIC ("WNIC") or wireless adapter for communicating with a wireless network (such as WI-FI).
[0196] In addition, the communication interface 1210 can facilitate communication with various types of wired or wireless networks. The communication interface 1210 can also facilitate communication using various communication protocols. The communication infrastructure 1212 can also include hardware, software, or both that couple the components of the computing device 1200 to each other. For example, the communication interface 1210 can use one or more networks and / or protocols to enable multiple computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the digitization process described herein. To illustrate, the image compression process can allow multiple devices (e.g., a server device for performing image processing tasks for a large number of images) to exchange information using various communication networks and protocols for exchanging information about selected workflows and image data for multiple images.
[0197] In the foregoing description, the present disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the present disclosure have been described with reference to the details discussed herein, and the accompanying drawings illustrate various embodiments. The above description and drawings are illustrative of the present disclosure and should not be construed as limiting the present disclosure. Numerous specific details have been described to provide a thorough understanding of the various embodiments of the present disclosure.
[0198] Without departing from the spirit or essential characteristics of the present invention, the present invention may be embodied in other specific forms. The described embodiments should be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with fewer or more steps / actions, or the steps / actions may be performed in a different order. In addition, the steps / actions described herein may be repeated or performed in parallel with each other or with different instances of the same or similar steps / actions. Therefore, the scope of this application is indicated by the appended claims rather than the foregoing description. All changes that fall within the equivalent meaning and scope of the claims should be included within their scope.
Claims
1. A non-transitory computer-readable medium storing thereon instructions that, when executed by at least one processor, cause a computing device to: Identifying a request to draw an object at a specified location within a digital scene; generating global features for the digital scene using a neural network; generating a local feature for the designated location in the digital scene; generating location-specific lighting parameters for the specified location from the global features and the local features using the neural network; and A modified digital scene is rendered, the modified digital scene including the object illuminated at the specified location according to the location-specific lighting parameters.
2. The non-transitory computer-readable medium of claim 1 , wherein the instructions, when executed by the at least one processor, cause the computing device to generate the global features by extracting a global feature map from the digital scene using a first set of layers of the neural network.
3. The non-transitory computer-readable medium of claim 2 , wherein the instructions, when executed by the at least one processor, cause the computing device to generate the location-specific lighting parameters by concatenating the global feature map and the local features and processing the concatenation of the global feature map and the local features using a second set of layers of the neural network. 4 . The non-transitory computer-readable medium of claim 1 , wherein the instructions, when executed by the at least one processor, cause the computing device to generate the local feature by masking the global feature based on the specified location.
5. The non-transitory computer-readable medium of claim 4 , wherein the instructions, when executed by the at least one processor, cause the computing device to generate the location-specific lighting parameters for the specified location by concatenating the global features and the masked global features and regressing the location-specific lighting parameters from the concatenation of the global features and the masked global features using the neural network. 6 . The non-transitory computer-readable medium of claim 4 , wherein masking the global feature based on the specified location comprises multiplying the global feature by a one-hot vector having a value in a group at the specified location.
7. The non-transitory computer-readable medium of claim 1 , wherein the instructions, when executed by the at least one processor, cause the computing device to generate the local feature by generating a hypercolumn of features by extracting values corresponding to the specified position from a feature map from a layer of the neural network used to generate the global feature.
8. The non-transitory computer-readable medium of claim 7 , wherein the instructions, when executed by the at least one processor, cause the computing device to concatenate the global features and the hypercolumn of features and to generate the location-specific lighting parameters for the specified location by regressing the location-specific lighting parameters from the concatenation of the global features and the hypercolumn of features using the neural network.
9. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by the at least one processor, cause the computing device to dynamically update the location-specific lighting parameters as the perspective of the digital scene moves due to movement of a camera of the computing device capturing the digital scene, and to dynamically render a modified digital scene based on the updated location-specific lighting parameters.
10. A system comprising: one or more memory devices storing a local illumination estimation neural network; as well as At least one server device is configured to enable the system to: receiving an indication of a specified location within the digital scene; generating a global feature for the digital scene using the local illumination estimation neural network; generating a masked global feature by masking the global feature based on the specified location within the digital scene; as well as Location-specific lighting parameters are generated for the specified location by regressing the masked global features using the local lighting estimation neural network.
11. The system of claim 10, wherein the at least one server device is configured to cause the system to generate a stitched feature by stitching the global feature and the masked global feature, and wherein regressing the masked global feature comprises regressing the stitched feature.
12. The system of claim 10, wherein the at least one server device is configured to cause the system to generate the location-specific lighting parameters for the specified location by generating location-specific spherical harmonic coefficients indicative of lighting conditions at the specified location.
13. The system of claim 10, wherein the at least one server device is configured to cause the system to generate a modified digital scene by drawing an object within the digital scene at the specified location, the object being illuminated according to the location-specific lighting parameters for the specified location.
14. The system of claim 13, wherein the at least one server device is configured to cause the system to render the object as an augmented reality object within the digital scene captured by a camera of the system.
15. A computer-implemented method for estimating lighting conditions for a virtual object, comprising: receiving an indication of a specified location within the digital scene; generating a series of feature maps from the digital scene using a neural network; generating a hypercolumn of features by extracting a value corresponding to the specified position from the feature map; as well as Location-specific lighting parameters are generated for the specified location by regressing the hypercolumn of features using the neural network.
16. The computer-implemented method of claim 15, further comprising drawing a modified digital scene in a graphical user interface of the mobile device, the modified digital scene including a virtual object illuminated at the specified location according to the location-specific lighting parameters. 17 . The computer-implemented method of claim 16 , wherein drawing the modified digital scene comprises drawing an augmented reality scene in real time on a mobile device as a camera of the mobile device captures the digital scene.
18. The computer-implemented method of claim 15, further comprising generating global features from the feature map for the digital scene.
19. The computer-implemented method of claim 18, further comprising generating a concatenated feature by concatenating the global feature and the supercolumn of features, and wherein regressing the supercolumn of features comprises regressing the concatenated feature.
20. The computer-implemented method of claim 15, wherein generating the location-specific lighting parameters for the specified location comprises generating location-specific spherical harmonic coefficients indicative of lighting conditions at the specified location.