Neural optical transmission

A neural network trained on UV texture maps from controlled lighting conditions addresses the limitations of image capture devices in handling varying viewpoints and lighting, enabling advanced image rendering and manipulation.

CN114514561BActive Publication Date: 2025-07-15GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080069485.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-03
Filing Date
2020-05-04
Publication Date
2025-07-15
Estimated Expiration
2040-05-04

AI Technical Summary

Technical Problem

When existing image capturing devices process images captured under low light conditions, it is difficult to effectively correct the problem of insufficient lighting on objects, which limits the development of full-angle re-illumination systems.

Method used

By training neural networks to learn light transmission functions, using data sets under multi-view and multi-light lighting conditions, a texture map that can synthesize objects from novel perspectives and lighting conditions is generated, realizing digital re-illumination and free viewpoint rendering.

Benefits of technology

It realizes efficient digital re-illumination and free viewpoint rendering of three-dimensional objects in the image, and can generate realistic object images under different lighting conditions, improving image quality and visual effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114514561B_ABST
    Figure CN114514561B_ABST
Patent Text Reader

Abstract

The example relates to an implementation of neural light transport. A computing system can obtain data indicating multiple UV texture maps and geometries of an object. Each UV texture map depicts the object from one of multiple viewpoints. The computing system can use the data to train a neural network to learn a light transport function. The light transport function can be a continuous function specifying how light interacts with the object when the object is viewed from the multiple viewpoints. The computing system can generate an output UV texture map depicting the object from a synthetic viewpoint based on the application of the light transport function by the trained neural network.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 910,265, filed on October 3, 2019, the entire content of which is incorporated herein by reference. Technical Field

[0003] This application relates to neural light transport. Background Art

[0004] Many modern computing devices, including mobile phones, personal computers, and tablet computers, include image - capture devices, such as still cameras and / or video cameras. The image - capture devices can capture images, such as images including people, animals, landscapes, and / or objects.

[0005] Some image - capture devices and / or computing devices can correct or otherwise modify the captured images. For example, some image - capture devices can provide "red - eye" correction, which removes artifacts in the form of red eyes in images of people and animals that may be present in images captured using bright lights (such as flash illumination). After the captured image has been corrected, the corrected image can be saved, displayed, sent, printed on paper, and / or otherwise utilized. In some cases, the image of an object may be affected by insufficient illumination during image capture. Summary of the Invention

[0006] Disclosed herein are embodiments related to the development of neural light transport that enables digital relighting and free - viewpoint rendering of three - dimensional (3D) objects captured in images. In particular, to train a neural network to learn a light - transport function, a computing system can use a data set associated with a collection of UV texture maps that depict an object captured using an illumination stage. The data set can specify, for each UV texture map within the collection of UV texture maps, the perspective of the camera and the position of the light that illuminates the object. By using this data set, one or more neural networks can develop a neural light transport that can subsequently be used to synthesize the texture of an object from novel viewpoints and / or novel illuminations. The synthesized texture map can then be applied to a 3D model of the object for relighting to produce an output texture map of the object from the synthetic viewpoints (e.g., novel viewpoints and illuminations).

[0007] In one aspect, the present application describes a method. The method involves obtaining, at a computing system, data indicative of a plurality of UV texture maps and geometry of an object. Each UV texture map depicts the object from one of a plurality of viewpoints. The method may also involve the computing system using the data to train a neural network to learn a light transport function. The light transport function specifies how light interacts with the object when the object is viewed from the plurality of viewpoints. The method may also involve the computing system generating an output UV texture map depicting the object from a synthetic viewpoint based on the application of the light transport function by the trained neural network.

[0008] In another aspect, the present application describes a system. The system includes a sensor and a computing system. The computing system is configured to obtain data indicative of a plurality of UV texture maps and geometry of an object. Each UV texture map depicts the object from one of a plurality of viewpoints, and the sensor captures data indicative of the geometry of the object. The computing system is further configured to use the data to train a neural network to learn a light transport function. The light transport function specifies how light interacts with the object when the object is viewed from the plurality of viewpoints. The computing system is further configured to generate an output UV texture map depicting the object from a synthetic viewpoint based on the application of the light transport function by the trained neural network.

[0009] In yet another example, the present application describes a non-transitory computer-readable medium configured to store instructions that, when executed by a computing system including one or more processors, cause the computing system to perform operations. The operations involve obtaining data indicative of a plurality of UV texture maps and geometry of an object. Each UV texture map depicts the object from one of a plurality of viewpoints. The operations also involve using the data to train a neural network to learn a light transport function. The light transport function specifies how light interacts with the object when the object is viewed from the plurality of viewpoints. The operations also involve generating an output UV texture map depicting the object from a synthetic viewpoint based on the application of the light transport function by the trained neural network.

[0010] In another aspect, the present application describes a system that includes means for implementing neural light transport. The system includes means for obtaining data indicative of a plurality of UV texture maps and geometry of an object. Each UV texture map depicts the object from one of a plurality of viewpoints. The system also includes means for using the data to train a neural network to learn a light transport function. The light transport function specifies how light interacts with the object when the object is viewed from the plurality of viewpoints. The system also includes means for generating an output UV texture map depicting the object from a synthetic viewpoint based on the application of the light transport function by the trained neural network.

[0011] The foregoing Summary is illustrative only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, other aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 FIG. illustrates a schematic diagram of a computing device according to an example embodiment.

[0013] Figure 2 FIG. illustrates a schematic diagram of a cluster of server devices according to an example embodiment.

[0014] Figure 3A FIG. depicts an ANN architecture according to an example embodiment.

[0015] Figure 3B FIG. depicts training an ANN according to an example embodiment.

[0016] Figure 4A FIG. depicts a convolutional neural network (CNN) architecture according to an example embodiment.

[0017] Figure 4B FIG. depicts a convolution according to an example embodiment

[0018] Figure 5 FIG. depicts a system involving an ANN and a mobile device according to an example embodiment.

[0019] Figure 6 FIG. illustrates a system for implementing neural optical transmission according to an example embodiment.

[0020] Figure 7A FIG. shows a set of operations for implementing neural optical transmission according to an example embodiment.

[0021] Figure 7B FIG. illustrates ray casting for determining 3D pixel projections according to an example embodiment.

[0022] Figure 7C FIG. illustrates the connection of an object rendered into its counterpart in the UV space according to an example embodiment.

[0023] Figure 7D FIG. illustrates a rendering process according to an example embodiment.

[0024] Figure 8 FIG. is a flowchart of a method for implementing a neural optical transmission function according to an example embodiment.

[0025] Figure 9 FIG. is a schematic diagram of a conceptual partial view of a computer program for executing a computer process on a computing system arranged according to at least some embodiments presented herein. Detailed Implementation Manner

[0026] The present disclosure describes example methods, devices, and systems. It should be understood that the terms "example" and "exemplary" as used herein mean "serving as an example, instance, or illustration". Any embodiment or feature described as "example" or "exemplary" herein is not necessarily to be construed as more preferred or advantageous than other embodiments or features. Other embodiments may be utilized, and other changes may be made without departing from the scope of the subject matter presented herein.

[0027] A light stage and / or other hardware may be used to capture a fully relightable object. However, during the capture time, the object remains static. In addition, the light stage may include a limited number of mounted lights. Thus, the light stage may capture the object only from predefined views, thus limiting the development of a full-angle relighting system.

[0028] The example presented herein describes methods and systems for implementing neural light transport. A computing system may train a neural network to learn a light transport function, also referred to herein as neural light transport. The light transport function may be a function (e.g., a continuous function) that specifies how light interacts with an object when the object is viewed from various perspectives. For example, the light transport function may enable the computing system to describe how light interacts with the material of the object at the perspective of an observer. Thus, the light transport function may be used to generate a representation (e.g., an image or a UV texture map) of the object depicting the object from a synthetic perspective. For example, the synthetic perspective may show the object using novel illumination (i.e., illumination from a light source at a new position), where one or more materials of the object are modified or changed, and / or using a novel perspective (e.g., a viewpoint of the object that has not been previously captured and recorded via a camera).

[0029] An example method may involve obtaining data indicative of an image of an object. Each image may depict the object from a different perspective. For example, a light stage may be used to collect the data. In particular, the light stage may enable the development of a One-Light-at-a-Time dataset that represents a UV texture map generated based on images captured from various fixed and known perspectives (e.g., dozens of perspectives) while illuminating the object with lights positioned near the light stage in a known order (e.g., one light at a time). As a result, the data may specify, for each image, information about the perspective of the camera (e.g., the camera pose) and the pose of one or more specific lights that illuminate the object. In addition, one or more sensors may provide data representing the geometry of the object. The UV texture map and the geometry information may together form a dataset that the computing system may use to train one or more neural networks to learn the light transport function.

[0030] The trained neural network can then generate an output UV texture map that depicts the object from a synthetic perspective. For example, the light transport function can enable a computing system to synthesize textures of novel views and novel illumination of an object, which can then be applied to a 3D model of the object for relighting or novel view synthesis. Additionally, in some examples, the synthetic perspective can be used to show an object with one or more different materials.

[0031] I. Example Computing Devices and Cloud-Based Computing Environments

[0032] The following embodiments describe the architectural and operational aspects of example computing devices and systems that can employ the disclosed ANN implementations, as well as their features and advantages.

[0033] Figure 1 is a simplified block diagram illustrating a computing system 100, showing some of the components that can be included in a computing device arranged to operate in accordance with embodiments herein. The computing system 100 can be a client device (e.g., a device actively operated by a user), a server device (e.g., a device that provides computing services to client devices), or some other type of computing platform. Some server devices can operate as client devices from time to time to perform specific operations, and some client devices can incorporate server characteristics.

[0034] In this example, the computing system 100 includes a processor 102, a memory 104, a network interface 106, and an input / output unit 108, all of which can be coupled by a system bus 110 or a similar mechanism. In some embodiments, the computing system 100 can include other components and / or peripherals (e.g., removable storage devices, printers, etc.).

[0035] The processor 102 can be one or more of any type of computer processing element, such as a central processing unit (CPU), a coprocessor (e.g., a math, graphics, or cryptographic coprocessor), a digital signal processor (DSP), a network processor, and / or an integrated circuit or controller that performs processor operations in one form. In some cases, the processor 102 can be one or more single-core processors. In other cases, the processor 102 can be one or more multi-core processors having multiple independent processing units. The processor 102 can also include register memory for temporarily storing instructions and associated data being executed, as well as cache memory for temporarily storing recently used instructions and data.

[0036] The memory 104 can be any form of computer-usable memory, including but not limited to random access memory (RAM), read-only memory (ROM), and non-volatile memory. This can include flash memory, hard disk drives, solid state drives, rewritable optical discs (CDs), rewritable digital video discs (DVDs), and / or tape storage devices, just as a few examples.

[0037] The computing system 100 can include fixed memory as well as one or more removable memory units, the latter including but not limited to various types of Secure Digital (SD) cards. Thus, the memory 104 represents both the main memory unit and the long-term storage device. Other types of memory can include biological memory.

[0038] The memory 104 can store program instructions and / or the data on which the program instructions can operate. By way of example, the memory 104 can store these program instructions on a non-transitory computer-readable medium such that the instructions can be executed by the processor 102 to perform any method, process, or operation disclosed in this specification or the accompanying drawings.

[0039] As Figure 1 shown, the memory 104 can include firmware 104A, a kernel 104B, and / or an application 104C. The firmware 104A can be program code for booting or otherwise starting some or all of the computing system 100. The kernel 104B can be an operating system, including modules for memory management, scheduling and management of processes, input / output, and communication. The kernel 104B can also include device drivers that allow the operating system to communicate with the hardware modules of the computing system 100 (such as memory units, networking interfaces, ports, and buses). The application 104C can be one or more user-space software programs, such as a web browser or an email client, and any software libraries used by these programs. In some examples, the application 104C can include one or more neural network applications. The memory 104 can also store the data used by these and other programs and applications.

[0040] The network interface 106 can take the form of one or more wired interfaces, such as Ethernet (e.g., Fast Ethernet, Gigabit Ethernet, etc.). The network interface 106 can also support communication over one or more non-Ethernet media (such as coaxial cable or power line) or over wide area media (such as Synchronous Optical Network (SONET) or Digital Subscriber Line (DSL) technology). The network interface 106 can also take the form of one or more wireless interfaces, such as IEEE 802.11 (Wifi), , a Global Positioning System (GPS) or a wide area wireless interface. However, other forms of physical layer interfaces and other types of standard or proprietary communication protocols may be used on the network interface 106. Additionally, the network interface 106 may include multiple physical interfaces. For example, some embodiments of the computing system 100 may include Ethernet, and a Wifi interface.

[0041] The input / output unit 108 may facilitate the interaction between the user and peripheral devices with the computing system 100 and / or other computing systems. The input / output unit 108 may include one or more types of input devices, such as a keyboard, a mouse, one or more touchscreens, sensors, biosensors, etc. Similarly, the input / output unit 108 may include one or more types of output devices, such as a screen, a monitor, a printer, and / or one or more light-emitting diodes (LEDs). Additionally or alternatively, the computing system 100 may communicate with other devices using, for example, a Universal Serial Bus (USB) or a High-Definition Multimedia Interface (HDMI) port interface.

[0042] In some embodiments, one or more instances of the computing system 100 may be deployed to support a cluster architecture. The exact physical location, connectivity, and configuration of these computing devices may be unknown and / or unimportant to the client device. Thus, the computing devices may be referred to as "cloud-based" devices, which may be located in various remote data center locations. Additionally, the computing system 100 may enable the execution of the embodiments described herein, including the use of neural networks and the implementation of neuro-optical transmission.

[0043] Figure 2 depicts a cloud-based server cluster 200 according to an example embodiment. In Figure 2 , one or more operations of the computing devices (e.g., the computing system 100) may be distributed among the server devices 202, the data storage 204, and the router 206, all of which may be connected by a local cluster network 208. The number of server devices 202, data storage 204, and routers 206 in the server cluster 200 may depend on the (one or more) computing tasks and / or applications assigned to the server cluster 200. In some examples, the server cluster 200 may perform one or more operations described herein, including the use of neural networks and the implementation of neuro-optical transmission functions.

[0044] The server device 202 can be configured to perform various computing tasks of the computing system 100. For example, one or more computing tasks can be distributed among one or more server devices 202. To the extent that these computing tasks can be executed in parallel, such distribution of tasks can reduce the total time to complete these tasks and return the results. For simplicity, both the server cluster 200 and the stand-alone server device 202 can be referred to as "server devices". This nomenclature should be understood to imply that one or more different server devices, data storage devices, and cluster routers may be involved in the operation of the server device.

[0045] The data storage device 204 can be a data storage array that includes a drive array controller configured to manage read and write access to a group of hard disk drives and / or solid state drives. The drive array controller, alone or in combination with the server device 202, can also be configured to manage backup or redundant copies of the data stored in the data storage device 204 to protect against drive failures or other types of failures of the units that prevent one or more server devices 202 from accessing the cluster data storage device 204. Other types of memory other than drives can be used.

[0046] The router 206 can include networking equipment configured to provide internal and external communication for the server cluster 200. For example, the router 206 can include one or more packet switching and / or routing devices (including switches and / or gateways) configured to provide (i) network communication between the server device 202 and the data storage device 204 via the cluster network 208, and / or (ii) network communication between the server cluster 200 and other devices via the communication link 210 to the network 212.

[0047] In addition, the configuration of the cluster router 206 can be at least partially based on the data communication requirements of the server device 202 and the data storage device 204, the latency and throughput of the local cluster network 208, the latency, throughput, and cost of the communication link 210, and / or other factors that may contribute to the cost, speed, fault tolerance, resilience, efficiency, and / or other design goals of the system architecture.

[0048] As a possible example, the data storage device 204 can include any form of database, such as a Structured Query Language (SQL) database. Various types of data structures can store information in such a database, including but not limited to tables, arrays, lists, trees, and tuples. In addition, any database in the data storage device 204 can be monolithic or distributed across multiple physical devices.

[0049] The server device 202 can be configured to send data to and receive data from the cluster data storage device 204. Such transmission and retrieval can respectively take the form of SQL queries or other types of database queries and the outputs of these queries. Additional text, images, video, and / or audio may also be included. Additionally, the server device 202 can organize the received data into a web page representation. Such a representation can take the form of a markup language (such as Hypertext Markup Language (HTML), Extensible Markup Language (XML)) or some other standardized or proprietary format. Further, the server device 202 can have the ability to execute various types of computerized scripting languages, such as but not limited to Perl, Python, PHP Hypertext Preprocessor (PHP), Active Server Pages (ASP), JavaScript, etc. Computer program code written in these languages can facilitate providing web pages to client devices and interacting with the client devices of the web pages.

[0050] II. Artificial Neural Networks

[0051] A. Example ANN

[0052] An artificial neural network (ANN) is a computational model in which multiple simple units that work independently in parallel and without central control can be combined to solve complex problems. The ANN is represented as multiple nodes arranged into multiple layers, with connections between the nodes of adjacent layers.

[0053] Figure 3A An example ANN 300 is shown. In particular, the ANN 300 represents a feedforward multi-layer neural network, but similar structures and principles are also used in, for example, convolutional neural networks (CNNs), recurrent neural networks, and recursive neural networks. The ANN 300 can represent an ANN trained to perform a specific task, such as image processing techniques (e.g., segmentation, semantic segmentation, image enhancement) or learning the neural optical transfer function described herein. In further examples, the ANN 300 can learn to perform other tasks, such as computer vision, risk assessment, etc.

[0054] As Figure 3A shown, the ANN 300 consists of four layers: an input layer 304, a hidden layer 306, a hidden layer 308, and an output layer 310. Three nodes of the input layer 304 respectively receive X1, X2, and X3 as initial input values 302. Two nodes of the output layer 310 respectively produce Y1 and Y2 as final output values 312. Thus, the ANN 300 is a fully connected network because the nodes of each layer except the input layer 304 receive inputs from all the nodes in the previous layer.

[0055] Solid arrows between nodes represent connections through which intermediate values flow and are each associated with a respective weight applied to the corresponding intermediate value. Each node performs an operation on its input values and their associated weights (e.g., values between 0 and 1, including 0 and 1) to produce an output value. In some cases, this operation can involve a dot product sum of the product of each input value and its associated weight. An activation function can be applied to the result of the dot product sum to produce the output value. Other operations are possible.

[0056] For example, if a node receives input values {x1, x2,..., x n} on n connections, with corresponding weights {w1, w2,..., w n}, then the dot product sum d can be determined as:

[0057]

[0058] where b is a bias specific to the node or specific to the layer.

[0059] Notably, by giving a value of 0 to one or more weights, the fully - connected nature of ANN 300 can be used to effectively represent a partially - connected ANN. Similarly, the bias can also be set to 0 to eliminate the b term.

[0060] Activation functions, such as the logistic function, can be used to map d to an output value y between 0 and 1, including 0 and 1:

[0061]

[0062] Functions other than the logistic function, such as the sigmoid or tanh functions, can be used instead.

[0063] Then, y can be used on the output connections of each node and will be modified by its respective weight. In particular, in ANN 300, the input values and weights are applied to the nodes of each layer, from left to right until the final output value 312 is produced. If ANN 300 has been fully trained, the final output value 312 is a proposed solution to the problem that ANN 300 was trained to solve. To obtain a meaningful, useful, and reasonably accurate solution, ANN 300 requires at least a certain degree of training.

[0064] B. Training

[0065] Training an ANN can involve providing the ANN with some form of supervised training data, i.e., a collection of input values and desired or ground truth output values. For example, supervised training that enables an ANN to perform an image processing task can involve providing pairs of images, where the pairs include a training image and a corresponding ground truth mask representing the desired output of the training image (e.g., a desired segmentation). For ANN 300, this training data can include m sets of input values paired with output values. More formally, the training data can be represented as:

[0066]

[0067] where i = 1...m, and and are the input values X 1,i 、X 2,i and X 3,i 's desired output values.

[0068] The training process involves applying the input values from such a set to ANN 300 and generating an associated output value. A loss function can be used to evaluate the error between the generated output value and the ground truth output value. In some cases, this loss function can be the sum of differences, the mean squared error, or some other metric. In some cases, error values are determined for all m sets, and the error function involves computing an aggregation (e.g., an average) of these values.

[0069] Once the error is determined, the weights on the connections are updated to try to reduce the error. Put simply, this update process should reward "good" weights and punish "bad" weights. Thus, the update should distribute the "responsibility" for the error through ANN 300 in such a way that future iterations of the training data result in lower error. For example, the update process can involve modifying at least one weight of ANN 300 such that subsequent applications of ANN 300 to the training image generate a new output that more closely matches the ground truth mask corresponding to the training image.

[0070] The training process continues to apply the training data to ANN 300 until the weights converge. Convergence occurs when the error is less than a threshold or when the change in error between consecutive training iterations is small enough. At this point, ANN 300 is said to be "trained" and can be applied to new sets of input values in order to predict unknown output values. When trained to perform an image processing technique, ANN 300 can produce an output for an input image that is very similar to the ground truth (i.e., the desired result) created for the input image.

[0071] Many training techniques for ANNs utilize some form of backpropagation. During backpropagation, the input signal propagates forward through the network to the output, and then the network error is calculated with respect to the target variable and the network error propagates backward toward the input. In particular, backpropagation distributes the error layer by layer from right to left through ANN 300. Thus, first, the weights of the connections between the hidden layer 308 and the output layer 310 are updated, second, the weights of the connections between the hidden layer 306 and the hidden layer 308 are updated, and so on. This update is based on the derivative of the activation function.

[0072] To further explain error determination and backpropagation, it is helpful to look at an example of the process in action. However, representing backpropagation can become very complex except on the simplest ANNs. Thus, Figure 3B a very simple ANN 330 is introduced to provide an illustrative example of backpropagation.

[0073] Weight Node Weight Node <![CDATA[w1]]> I1, H1 <![CDATA[w5]]> H1, O1 <![CDATA[w2]]> I2, H1 <![CDATA[w6]]> H2, O1 <![CDATA[w3]]> I1, H2 <![CDATA[w7]]> H1, O2 <![CDATA[w4]]> I2, H2 <![CDATA[w8]]> H2, O2

[0074] Table 1

[0075] ANN 330 consists of three layers, an input layer 334, a hidden layer 336, and an output layer 338, with two nodes in each layer. Initial input values 332 are provided to the input layer 334, and the output layer 338 produces a final output value 340. Weights have been assigned to each connection, and in some examples biases (e.g., Figure 3B b1, b2 as shown) may also be applied to the net input of each node in the hidden layer 336. For clarity, Table 1 maps the weights to the pairs of nodes with the connections to which these weights are applied. As an example, w2 is applied to the connection between nodes I2 and H1, w7 is applied to the connection between nodes H1 and O2, and so on.

[0076] The goal of training ANN 330 is to update the weights through a number of forward and backpropagation iterations until the final output value 340 is close enough to the specified desired output. Note that training ANN 330 effectively with a single set of training data is only for that set. If multiple sets of training data are used, ANN 330 will also be trained according to these sets.

[0077] 1. Example forward pass

[0078] To initiate the forward pass, the net input to each node in the hidden layer 336 is calculated. From the net input, the output of these nodes can be found by applying the activation function. For node H1, the net input net H1 is:

[0079] net H1 = w1X1 + w2X2 + b1 (4)

[0080] Applying an activation function (here, the logistic function) to this input determines the output out of node H1 H1 which is

[0081]

[0082] For node H2, following the same process, the output out can also be determined H2 . The next step in the feed - forward iteration is to perform the same calculation for the nodes in output layer 338. For example, the net input net to node O1 O1 is

[0083] net O1 = w5out H1 + w6out H2 + b2 (6)

[0084] Therefore, the output out of node O1 O1 is

[0085]

[0086] For node O2, following the same process, the output out can be determined O2 . At this point, the total error Δ can be determined based on the loss function. For example, the loss function can be the sum of the squared errors of the nodes in output layer 508. In other words

[0087]

[0088] The multiplicative constant in each term is used to simplify the differentiation during backpropagation. Since the overall result will be scaled by the learning rate anyway, this constant does not negatively affect the training. Anyway, at this point, the feed - forward iteration is complete and backpropagation begins

[0089] 2. Backpropagation

[0090] As mentioned above, the goal of backpropagation is to use Δ (i.e., the total error determined based on the loss function) to update the weights such that they contribute less error in future feed - forward iterations. As an example, consider the weight w5. The goal involves determining how much a change in w5 affects Δ. This can be expressed as the partial derivative Using the chain rule, this term can be expanded as

[0091]

[0092] Therefore, the effect of a change in w5 on Δ is equivalent to (i) the effect of a change in out O1 on Δ, (ii) the effect of a change in net O1The change of to out O1 The impact of, and (iii) the change of w5 on net O1 The impact of. Each of these multiplicative terms can be determined independently. Intuitively, this process can be considered as isolating the impact of w5 on net o1 The impact of net O1 on out O1 The impact of out O1 on Δ.

[0093] For other weights fed to the output layer 338, this process can be repeated. Note that the weights are not updated until the updates for all weights have been determined at the end of backpropagation. Then, all weights are updated before the next feedforward iteration.

[0094] After calculating the updates for the remaining weights w1, w2, w3, and w4, continue the backpropagation pass towards the hidden layer 336. This process can be repeated for other weights fed to the output layer 338. At this point, the backpropagation iteration ends and all weights have been updated. The ANN 330 can be continuously trained through subsequent feedforward and backpropagation iterations. In some cases, after several feedforward and backpropagation iterations (e.g., thousands of iterations), the error can be reduced to produce a result close to the original desired result. At this point, the values of Y1 and Y2 will be close to the target values. As shown, by using a differentiable loss function, the total error of the prediction output by the ANN 330 compared to the desired result can be determined and used to modify the weights of the ANN 330 accordingly.

[0095] In some cases, if the hyperparameters of the system (e.g., biases b1 and b2 and learning rate α) are adjusted, an equivalent amount of training can be completed with fewer iterations. For example, setting the learning rate closer to a specific value may result in a faster reduction of the error rate. Additionally, the biases can be updated in a manner similar to updating the weights as part of the learning process.

[0096] In any case, the ANN 330 is just a simplified example. Arbitrarily complex ANNs can be developed by adjusting the number of nodes in each input and output layer to solve specific problems or goals. Additionally, more than one hidden layer can be used, and there can be any number of nodes in each hidden layer.

[0097] III. Convolutional Neural Network

[0098] A convolutional neural network (CNN) is similar to an ANN in that a CNN can consist of a certain number of layers of nodes with weighted connections between them and possibly biases for each layer. The weights and biases can be updated through the feedforward and backpropagation processes discussed above. A loss function can be used to compare the output values of the feedforward process with the desired output values.

[0099] On the other hand, CNNs are typically designed under the explicit assumption that the initial input values are derived from one or more images. In some embodiments, each color channel of each pixel in a patch of images is a separate initial input value. Assuming each pixel has three color channels (e.g., red, green, and blue), even a small 32x32 pixel patch will result in 3072 incoming weights for each node in the first hidden layer. Clearly, using a naive ANN for image processing can lead to a very large and complex model that will take a long time to train.

[0100] In contrast, CNNs are designed to take advantage of the inherent structure that can be found in almost all images. In particular, the nodes in a CNN are only connected to a few nodes in the previous layer. This CNN architecture can be considered three-dimensional, where the nodes are arranged in blocks with width, height, and depth. For example, the previously mentioned 32x32 pixel patch with 3 color channels can be arranged as an input layer with 32 nodes in width, 32 nodes in height, and 3 nodes in depth.

[0101] Figure 4A An example CNN 400 is shown. The initial input values 402 represented as pixels X1...X m are provided to the input layer 404. As discussed above, the input layer 404 can have three dimensions based on the width, height, and number of color channels of the pixels X1...X m The input layer 404 provides the values into one or more sets of feature extraction layers, each set containing instances of a convolutional layer 406, a RELU layer 408, and a pooling layer 410. The output of the pooling layer 410 is provided to one or more classification layers 412. The final output values 414 can be arranged in a feature vector representing a concise characterization of the initial input values 402.

[0102] The convolutional layer 406 can transform these input values by sliding one or more filters around the three-dimensional space of its input values. The filters are represented by biases applied to the nodes and weights for the connections between them, and typically have a width and height that are smaller than the width and height of the input values. The result of each filter can be a two-dimensional patch of output values (referred to as a feature map), where the width and height can have the same size as the width and height of the input values, or one or more of these dimensions can have different sizes. The combination of the outputs of each filter results in a layer of feature maps in the depth dimension, where each layer represents the output of one of the filters.

[0103] Applying a filter can involve computing the dot product sum between the entries in the filter and a two-dimensional depth slice of the input values. An example of this is shown in Figure 4Bis shown. Matrix 420 represents the input to the convolutional layer and can thus be, for example, image data. The convolution operation overlays filter 422 on matrix 420 to determine output 424. For example, when filter 422 is positioned at the upper left corner of matrix 420 and the dot product sum of each entry is calculated, the result is 4. This is placed at the upper left corner of output 424.

[0104] Returning to Figure 4A , the CNN learns the filters during training so that these filters can ultimately identify certain types of features at specific locations in the input values. As an example, convolutional layer 406 can include filters that are ultimately able to detect edges and / or colors in the image patches from which the initial input values 402 are derived. A hyperparameter called the receptive field determines the number of connections between each node in convolutional layer 406 and input layer 404. This allows each node to focus on a subset of the input values.

[0105] The RELU layer 408 applies an activation function to the output provided by the convolutional layer 406. In fact, it has been determined that the rectified linear unit (RELU) function or its variants exhibit strong results in CNNs. The RELU function is a simple threshold function defined as f(x) = max(0, x). Thus, the output is 0 when x is negative and x when x is non - negative. A smooth, differentiable approximation of the RELU function is the softplus function. It is defined as f(x) = log(1 + e x ). Nevertheless, other functions can also be used in this layer.

[0106] The pooling layer 410 reduces the spatial size of the data by downsampling each two - dimensional depth slice of the output from the RELU layer 408. One possible method is to apply a 2x2 filter to each 2x2 block of the depth slice with a stride of 2. This reduces the width and height of each depth slice by a factor of 2, thereby reducing the overall size of the data by 75%.

[0107] The classification layer 412 calculates the final output value 414 in the form of a feature vector. As an example, in a CNN trained as an image classifier, each entry in the feature vector can encode the probability that an image patch contains an item of a specific class (e.g., face, cat, beach, tree, etc.).

[0108] In some embodiments, there are multiple sets of feature extraction layers. Thus, an instance of the pooling layer 410 can provide an output to an instance of the convolutional layer 406. Additionally, for each instance of the pooling layer 410, there can be multiple instances of the convolutional layer 406 and the RELU layer 408.

[0109] The CNN 400 represents a general structure that can be used in image processing. Similar to the layers in the ANN 300, the convolutional layer 406 and the classification layer 412 apply weights and biases, and these weights and biases can be updated during backpropagation so that the CNN 400 can learn. On the other hand, the RELU layer 408 and the pooling layer 410 typically apply fixed operations and thus may not be learnable.

[0110] Similar to the ANN, the CNN can include a different number of layers than those shown in the examples herein, and each of these layers can include a different number of nodes. Thus, the CNN 400 is for illustrative purposes only and should not be considered to limit the structure of the CNN.

[0111] Figure 5 Depicted is a system 500 involving an ANN operating on a computing system 502 and a mobile device 510 according to an example embodiment.

[0112] The ANN operating on the computing system 502 can correspond to the above-mentioned ANN 300 or ANN 330. For example, the ANN can be configured to execute instructions to perform the described operations, including learning one or more neural light transports. In some examples, the ANN can represent a CNN (e.g., CNN 400), a feedforward ANN, a gradient descent-based activation function ANN, or a recurrent feedback ANN, among other types.

[0113] As an example, the ANN can determine multiple processing parameters or techniques based on data derived from a UV texture map and geometry obtained from an object using an illumination stage. For example, the ANN 502 can undergo a machine learning process to "learn" how to manipulate the texture, perspective, and illumination of one or more objects like a human professional. The size of the dataset used may vary depending on the example.

[0114] In some examples, the dataset can depend on the arrangement of the illumination stage. For example, the number of lights and the number of captured perspectives can vary depending on the illumination stage used to develop the dataset.

[0115] Figure 6 Illustrated is a system that implements neural light transport according to an example embodiment. The system 600 can be implemented by one or more computing systems (e.g., Figure 1 the computing system 100 shown) and can involve one or more features such as an illumination stage 602, neural light transport 604, material modeling 606, relighting 608, and synthetic perspective 610. In other examples, the system 600 can include other features or aspects in different arrangements.

[0116] System 600 may represent an example system that uses one or more computing systems to one or more neural networks to model how light propagates in a 3D scene. In particular, system 600 may implement the execution of material editing, relighting, and novel view synthesis of one or more objects using a trained neural network. In some examples, the trained neural network may be executed on various computing devices, such as wearable computing devices, smartphones, laptops, and servers. For example, a first computing system may train a neural network and provide the trained neural network to a second computing system.

[0117] The lighting stage 602 may be involved in the development of data (also referred to herein as a data set) that can be used to train one or more neural networks. Data can be developed using a physical lighting stage environment that includes lights positioned at various locations relative to the stage and one or more cameras positioned at multiple perspectives to capture the UV texture map of an object. Thus, during the lighting stage 602, a physical object may be placed in the structured lighting stage environment while one or more cameras capture the UV texture map of the physical object from different perspectives. When each UV texture map is captured, one or more lights positioned relative to the physical object may illuminate the physical object. Thus, the data captured during the lighting stage 602 may indicate the perspective of the camera during each image represented in the data and the pose of one or more lights used to illuminate the object. In some examples, capturing the UV texture map of an object may involve one or more cameras capturing images of the object that can be used to develop the UV texture map.

[0118] To further illustrate, an example embodiment may involve using a lighting stage equipped with a certain number of lights (e.g., 330 lights) arranged in different poses relative to the area where the object to be analyzed is placed. Measurements (e.g., images, sensor readings) of the object may be captured from various camera perspectives (e.g., 55 different perspectives) while the lights illuminate the object in a known configuration (e.g., one light at a time). In this way, data generated from sensors or camera measurements, as well as the known poses and perspectives of the (one or more) lights and cameras for each image, can be collected to develop a data set to train one or more neural networks to learn neural light transport as shown in neural light transport 604.

[0119] Once the data is obtained, system 600 may fit one or more neural networks to the observations within the data to train the (one or more) neural networks. Training the (one or more) neural networks may enable the (one or more) networks to learn functions, such as the following function:

[0120] f(x, ω i , ω o ) (10)

[0121] The function, which can be a continuous function, is referred to herein as the neural light transport or light transport function. The function can be determined and implemented by one or more neural networks. As shown above, the light transport function is a six-dimensional function and is arranged as follows: (i) two degrees of freedom represented by x that describe the position on the object surface; (ii) two degrees of freedom represented by ω i that define the incident light direction, and (iii) two remaining degrees of freedom represented by ω o that describe the viewing direction.

[0122] After training the neural network to learn the neural light transport function, the system 600 can query the function to perform different operations. For example, querying the function with x can cause the neural network to perform material modeling 606. Material modeling 606 can involve using spatially-varying material modeling to model an object, where each pixel in the image can change according to the material, the camera viewpoint, and the illumination direction. For most real-world objects, if we traverse on the object surface (e.g., the surface of a kitchen knife), we will observe multiple materials, such as a metal blade and a wooden handle, (thus, "is spatially-varying").

[0123] Querying the function with ω i can cause the neural network to render a scene of a physical object with novel illumination during relighting 608. Querying the function with ω o , the neural network can use the UV texture map that defines the camera view in the query to generate a perspective of the scene, as shown in the synthesis operation 610. The final image rendering is obtained by applying the inferred UV texture map to the 3D object.

[0124] Figure 7A 、 7B Figures 7C and 7D illustrate implementations using neural light transport according to an example embodiment. In particular, the implementation can be performed by the system 600 shown in Figure 6 and / or one or more computing systems (e.g., the computing system 100 shown in Figure 1 ). Thus, the example implementations illustrate the development and use of neural light transport with respect to a rabbit and a dragon that serve as objects. In other examples, different objects can be used to develop neural light transport.

[0125] Figure 7A shows a set of operations for implementing neural light transport according to an example embodiment. As shown, a system (e.g., the system 600 shown in Figure 6 ) can use a neural network to model how light transports in a 3D scene involving a rabbit and a dragon. This enables the system or another device to use the trained network discussed above with respect to the system 600 shown in Figure 6 to perform material editing, relighting, and novel view synthesis.

[0126] As shown, the lighting stage 702 can involve using lights and cameras positioned in different poses to capture images of a rabbit and a dragon to develop data for training a neural network. For example, the lighting stage 702 can involve positioning the rabbit and the dragon in a lighting stage setup that enables the lights and cameras to illuminate and capture images of the rabbit and the dragon from different perspectives while using various lighting techniques (e.g., one light at a time).

[0127] Neural light transport 704 can be developed by one or more neural networks executed on one or more computing systems. In particular, the data generated during the lighting stage 702 can enable the neural network to develop the neural light transport f(x, ω Figure 6 described above with respect to i , ω o ).

[0128] A light transport function can be defined on the surface of an object. Thus, light transport can be expressed as a high-dimensional UV map. UV mapping corresponds to a 3D modeling process of projecting a two-dimensional (2D) image onto the surface of a 3D model for texture mapping. Thus, the letters "U" and "V" are used to represent the axes of the 2D texture because "X", "Y", and "Z" are typically used to represent the axes of a 3D object in model space. UV textures can allow the use of colors (and other surface properties) from ordinary images to paint (or re-design) the polygons that make up a 3D object. This image is typically referred to as a UV texture map.

[0129] The UV mapping process can involve assigning pixels in an image to surface mappings on a polygon, typically done by "programmatically" copying triangular patches of the image map and pasting them onto the triangles of the object. UV textures represent an alternative to projection mapping, which involves using any pair of X, Y, Z coordinates or any transformation of the positions of the model. UV textures involve mapping into texture space rather than mapping into the geometric space of the object. Thus, rendering calculations use UV texture coordinates to determine how to draw a 3D surface. For each UV position, there is a four-dimensional function that takes the lighting direction (ω i ) and the viewing direction (ω o ) as inputs and outputs red, green, and blue (RGB) colors.

[0130] As shown, the variables of the light transport function can be queried to manipulate the output of the neural network. For example, querying x can enable the neural network to model spatially varying materials 706. This enables application material editing: changing the material of the dragon to the material of the rabbit. This can enable the neural network to determine how light can affect the appearance of different materials from different perspectives.

[0131] Querying ω iIt enables the neural network to adjust the illumination applied to the dragon and the rabbit, as shown for relighting 708. The relighting 708 can enable the neural network to show how the rabbit and the dragon might appear under different lighting conditions. Query ω o It enables the neural network to provide the dragon and the rabbit from a synthetic perspective, as shown in novel view synthesis 710. For novel view synthesis 710, the neural network can show the dragon and / or the rabbit from different perspectives (e.g., rotated 180 degrees), with or without the application of novel illumination.

[0132] Figure 7B Illustrated is ray casting for determining 3D pixel projections according to an example embodiment. The system can perform ray casting 722 using the knowledge of the geometry of the objects in the image 720 and the position of the camera relative to the light. In particular, the ray casting can determine where each pixel projects in the 3D space 724.

[0133] Using ray casting 722, each pixel can be traced to a 3D point on the object surface. It is also predefined which UV position each 3D point maps to. Linking these two together gives a mapping from each pixel to a UV position. This correspondence is used to generate the UV counterpart of the object rendering.

[0134] Figure 7C Illustrated is the connection of the object rendering to the counterpart of the object in the UV space according to an example embodiment. As shown, the neural network can use the information within the data (e.g., light and camera pose information) and the image to cause the light transport function to generate a UV map of red, green, and blue (RGB) values 726. In particular, each texel can encode a 4D function as follows:

[0135]

[0136] Accordingly, the neural network can use the image 720 to perform the light transport function to generate the UV map 726. This UV texture provides multi-view correspondence across different views without the need for explicit search among the views.

[0137] For each 3D point on the shown surface, the system can map the point to the UV map based on a predefined UV unwrapping process. Thus, the system can estimate where each pixel on the original RGB rendering should go on the UV map. By rearranging the pixel values, the system can determine the RGB map 726 in the UV space as shown.

[0138] Figure 7DIllustrates the rendering process according to an example embodiment. The system can repeat the above process to generate a UV buffer for the scene, which is an intermediate buffer used by the graphics engine to generate the final rendering. That is, the system can project the viewing direction, lighting direction, normal, and cosine terms 730 represented by action 732 into the UV space. These mappings can be described as a "UV buffer", similar to the "Z buffer" in traditional graphics. Thus, the network receives these UV buffers and aims to generate UVRGB 734 (as shown in the right figure).

[0139] Figure 8 Is a flowchart of a method 800 for implementing a neural light transport function according to an example embodiment. Method 800 may include one or more operations, functions, or actions as shown in one or more of blocks 802, 804, and 806. Although these blocks are shown in sequential order, in some cases these blocks may be executed in parallel and / or in a different order than described herein. Additionally, various blocks may be combined into fewer blocks, divided into additional blocks, and / or removed based on the desired implementation.

[0140] Furthermore, for method 800 and other processes and methods disclosed herein, the flowchart illustrates the functions and operations of one possible implementation of this embodiment. In this regard, each block may represent a module, segment, or portion of program code that includes one or more instructions executable by a processor to implement a specific logical function or step in the process. The program code may be stored on any type of computer-readable medium or memory, for example, a storage device such as a disk or hard drive.

[0141] The computer-readable medium may include a non-transitory computer-readable medium, for example, a computer-readable medium that stores data for a short period of time, such as register memory, processor cache, and random access memory (RAM). The computer-readable medium may also include a non-transitory medium or memory, for example, a secondary or persistent long-term memory, such as read-only memory (ROM), optical disk, or magnetic disk, compact disc read-only memory (CD-ROM).

[0142] The computer-readable medium may also be any other volatile or non-volatile storage system. For example, the computer-readable medium may be considered a computer-readable storage medium, a tangible storage device, or other article. Furthermore, for method 800 and other processes and methods disclosed herein, Figure 8 each block in may represent a circuit wired to perform a specific logical function in the process.

[0143] At block 802, method 800 involves obtaining data indicative of multiple UV texture maps and geometry of an object. Each UV texture map can depict the object from one of a variety of perspectives. A computing system such as a smartphone, camera, or server can obtain data representing an image of the object.

[0144] In some examples, a lighting stage can be used to obtain the multiple UV texture maps. In particular, the lighting stage can include a number of lights (e.g., dozens, hundreds) to illuminate the object in a sequential order (e.g., one at a time) when one or more cameras capture images of the object. These images can be used for subsequent generation of the UV texture maps. Thus, the lighting stage can enable data to specify information to be associated with each image, such as which light is illuminating the object and from which perspective the image was taken. By using various lights to illuminate the object one light at a time, in a known sequential order, and capturing images from various perspectives, data can be accumulated for subsequent use.

[0145] Additionally, the computing system can obtain data indicative of the geometry of the object from one or more sensors. For example, the computing system can obtain data from photometric stereo and depth sensors. The one or more sensors can include various types of sensors configured to measure physical aspects of the object. In some examples, these sensors can be part of the lighting stage.

[0146] At block 804, method 800 involves using the data to train a neural network to learn a light transport function. For example, the computing system can train the neural network to learn the light transport function based on information specifying the light positions and perspectives associated with each UV texture map. Additionally, the neural network can be trained such that the output of the light transport function depends on one or more materials of the object.

[0147] The light transport function can be a continuous function that specifies how light interacts with the object when viewed from multiple perspectives. As indicated above, the data can associate specific lighting and perspectives with each image to generate the UV texture maps. By collecting and analyzing data from multiple images (e.g., dozens, hundreds, thousands) and geometry information, one or more neural networks can learn how to express the information in the form of a light transport function. Thus, the light transport function can enable estimation of novel lighting and perspectives of the object.

[0148] At block 806, method 800 involves generating an output UV texture map depicting an object from a synthetic perspective based on the application of a light transport function by a trained neural network. In some cases, the synthetic perspective can include novel illumination applied to the object and / or a novel view of the object. For example, the synthetic perspective can include applications for spatially-varying material modeling of the object. Additionally, the output image can involve relighting applications that illuminate the object within the output image. In some examples, the synthetic perspective can represent an object with one or more modifications to the material of the object.

[0149] In some examples, generating the output UV texture map can involve determining the synthesis of the object texture from a particular perspective with a particular illumination. For example, the particular perspective can be different from the plurality of perspectives. Additionally, the computing system can also relight the 3D model of the object based on the determined synthesis of the object texture and generate an output image depicting the object based on the relighting of the 3D model such that the object includes a new material.

[0150] In some examples, method 800 also involves determining an output image depicting the synthetic perspective of the object based on the output UV texture map and displaying the output image on a display interface. For example, the computing system (or another computing system) can include a display interface to display the output image. Additionally, method 800 can also involve providing the trained neural network to a second computing system. For example, a server can train the neural network and send the trained neural network to a smartphone for local execution.

[0151] Figure 9 is a schematic diagram showing a conceptual partial view of a computer program for executing a computer process on a computing system arranged according to at least some embodiments presented herein. In some embodiments, the disclosed method can be implemented as computer program instructions encoded in a machine-readable format on a non-transitory computer-readable storage medium or encoded on other non-transitory media or articles.

[0152] In one embodiment, an example computer program product 900 is provided using a signal-bearing medium 902, which can include one or more programming instructions 904 that, when executed by one or more processors, can provide the above regarding Figures 1 - 8The described functionality or portions thereof. In some examples, the signal-bearing medium 902 can encompass a non-transitory computer-readable medium 906, such as, but not limited to, a hard disk drive, a compact disc (CD), a digital video disc (DVD), a digital tape, a memory, etc. In some embodiments, the signal-bearing medium 902 can encompass a computer-recordable medium 908, such as, but not limited to, a memory, a read / write (R / W) CD, an R / W DVD, etc. In some embodiments, the signal-bearing medium 902 can encompass a communication medium 910, such as, but not limited to, digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, etc.). Thus, for example, the signal-bearing medium 902 can be transmitted via a wireless form of the communication medium 910.

[0153] One or more programming instructions 904 can be, for example, computer-executable and / or logic-implemented instructions. In some examples, a computing device of the computer system 100, such as Figure 1 can be configured to provide various operations, functions, or actions in response to the programming instructions 904 transmitted to the computer system 100 by one or more of the computer-readable medium 906, the computer-recordable medium 908, and / or the communication medium 910.

[0154] The non-transitory computer-readable medium can also be distributed among multiple data storage elements that can be remotely located from one another. Alternatively, the computing device that executes some or all of the stored instructions can be another computing device, such as a server.

[0155] The detailed description above, with reference to the accompanying drawings, describes various features and functions of the disclosed systems, devices, and methods. Although various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent. The various aspects and embodiments disclosed herein are for illustrative purposes and are not intended to be limiting, and the true scope is indicated by the appended claims.

[0156] It should be understood that the arrangements described herein are for illustrative purposes only. Thus, those skilled in the art will understand that other arrangements and other elements (e.g., machines, devices, interfaces, functions, groupings of commands and functions, etc.) can be used in place of, and some elements can be omitted altogether depending on the desired results. Additionally, many of the elements described are functional entities that can be implemented as discrete or distributed components or in combination with other components in any suitable combination and location.

Claims

1. A method for implementing neural light transport, comprising: obtaining, at a computing system, data indicative of a plurality of UV texture maps and geometry of an object, wherein each UV texture map depicts the object from one of a plurality of viewpoints; training, by the computing system, a neural network using training data, wherein the training data includes data indicative of a plurality of UV texture maps and geometry of the object, wherein the training causes the neural network to learn a light transport function, wherein the light transport function specifies how light interacts with the object when the object is viewed from the plurality of viewpoints; and generating, by the computing system, an output UV texture map depicting the object from a synthetic viewpoint based on an application of the light transport function by the trained neural network.

2. The method according to claim 1, wherein Obtaining data indicative of a plurality of UV texture maps and geometry of the object includes: obtaining the plurality of UV texture maps using an illumination stage, wherein the illumination stage includes a plurality of lights arranged to illuminate the object in sequential order when one or more cameras capture an image of the object, the image of the object being used to subsequently generate the plurality of UV texture maps.

3. The method according to claim 1, wherein Obtaining data indicative of a plurality of UV texture maps and geometry of the object includes: obtaining data indicative of the geometry of the object from a sensor.

4. The method according to claim 3, wherein Obtaining data indicative of the geometry of the object from a sensor includes: obtaining data indicative of the geometry of the object from a photometric stereo and depth sensor.

5. The method according to claim 1, wherein, The training data further includes: information specifying the positions of the lights and the viewpoints of the cameras capturing images of the object.

6. The method according to claim 1, wherein, Training the neural network includes: training the neural network such that the output of the light transport function depends on one or more materials of the object.

7. The method according to claim 1, wherein Generating an output UV texture map depicting the object from a synthetic viewpoint includes: generating an output UV texture map depicting the object such that the synthetic viewpoint includes modeling a spatially varying material applied to the object.

8. The method according to claim 7, wherein, Generating an output UV texture map depicting the object such that the synthetic viewpoint includes modeling a spatially varying material applied to the object includes: modifying the material of the object for the synthetic viewpoint.

9. The method according to claim 1, wherein Generating an output UV texture map depicting the object from a synthetic viewpoint includes: generating an output UV texture map depicting the object from a synthetic viewpoint such that the synthetic viewpoint includes a novel view of the object.

10. The method according to claim 1, wherein, Generating an output UV texture map depicting the object from a synthetic viewpoint includes: generating an output UV texture map depicting the object from a synthetic viewpoint such that the synthetic viewpoint includes a relighting application of illuminating the object from a specific viewpoint among the plurality of viewpoints.

11. The method according to claim 1, further comprising: determining an output image depicting the synthetic viewpoint of the object based on the output UV texture map; and displaying the output image on a display interface.

12. The method according to claim 1, further comprising: providing the trained neural network to a second computing system.

13. The method according to claim 1, wherein Generating an output UV texture map depicting the object from a synthetic viewpoint includes: determining a synthesis of the object texture from a specific viewpoint with a specific illumination, wherein the specific viewpoint is different from the plurality of viewpoints; relighting a three-dimensional (3D) model of the object based on the determined synthesis of the object texture; and generating an output image depicting the object such that the object includes a new material.

14. A system for implementing neural light transport, comprising: a sensor; a computing system configured to: Obtain data indicating multiple UV texture maps and the geometry of an object, where each UV texture map depicts the object from one of multiple perspectives, and where a sensor captures data indicating the geometry of the object; Train a neural network using training data, where the training data includes data indicating multiple UV texture maps and the geometry of the object, where the training causes the neural network to learn a light transport function, where the light transport function specifies how light interacts with the object when the object is viewed from the multiple perspectives; and Generate an output UV texture map depicting the object from a synthetic perspective based on the application of the light transport function by the trained neural network.

15. The system according to claim 14, further comprising: A lighting stage, where the computing system is configured to use the lighting stage to obtain data representing multiple UV texture maps, and where the lighting stage includes multiple lights arranged to illuminate the object in a sequential order when one or more cameras capture an image of the object, the image of the object being used to subsequently generate the multiple UV texture maps.

16. The system according to claim 14, wherein, The computing system is further configured to: Generate an output UV texture map depicting the object from a synthetic perspective such that the synthetic perspective includes modeling a spatially varying material applied to the object.

17. The system according to claim 14, wherein The computing system is further configured to: Generate an output UV texture map depicting the object from a synthetic perspective such that the synthetic perspective includes a novel view of the object.

18. The system according to claim 14, wherein The computing system is further configured to: Provide the trained neural network to a second computing system.

19. The system according to claim 14, wherein, The computing system is further configured to: Determine an output image depicting the synthetic perspective of the object based on the output UV texture map; and Display the output image on a display interface.

20. A non-transitory computer-readable medium configured to store instructions that, when executed by a computing system including one or more processors, cause the computing system to perform operations, the operations including: Obtain data indicating multiple UV texture maps and the geometry of an object, where each UV texture map depicts the object from one of multiple perspectives; Train a neural network using training data, where the training data includes data indicating multiple UV texture maps and the geometry of the object, where the training causes the neural network to learn a light transport function, where the light transport function specifies how light interacts with the object when the object is viewed from the multiple perspectives; and Generate an output UV texture map depicting the object from a synthetic perspective based on the application of the light transport function by the trained neural network.

Citation Information

Patent Citations

  • Systems and methods for rendering avatars with deep appearance models

    US20190213772A1