Deformable Neural Radiance Field
D-NeRF addresses the challenge of rendering non-rigidly deforming objects by incorporating a deformation model that maps between observation and canonical frames, resulting in accurate and artifact-free synthetic views.
Patent Information
- Application Number
- JP2023528508
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-16
- Filing Date
- 2021-01-14
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2041-01-14
AI Technical Summary
Conventional Neural Radiance Fields (NeRF) struggle to accurately synthesize new views of scenes containing non-rigidly deforming objects, such as people, due to their inability to handle movement and materials like hair, glasses, and earrings.
The development of a deformable neural radiance field (D-NeRF) that includes a deformation model to map positions between an observation frame and a canonical frame, allowing for the accurate representation of non-rigid movements and transformations.
D-NeRF effectively predicts synthetic views of non-rigidly deforming scenes without artifacts, enabling realistic rendering of dynamic objects and environments.
Smart Images

Figure 0007695357000020 
Figure 0007695357000021 
Figure 0007695357000022
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application is a non - provisional application claiming priority to U.S. Provisional Patent Application No. 63 / 198,841, filed on November 16, 2020, entitled "DEFORMABLE NEURAL RADIANCE FIELDS", the entire content of which is incorporated by reference herein.
[0002] Technical Field This specification relates to image synthesis using neural radiance fields (NeRF).
Background Art
[0003] Background Some computers configured to render computer - graphic objects can render an object in a specified view assuming a plurality of existing views. For example, in the case of several depth images and color images captured from a camera regarding a scene including such computer - graphic objects, it may be a goal to synthesize new views of the scene seen from different viewpoints. The scene may be a real - world scene where views are captured using physical color sensors and depth sensors, or a synthetic scene where views are captured using rendering algorithms such as rasterization. In the case of a real - world scene, there are many depth - sensing technologies such as time - of - flight sensors, structured - light - based sensors, and stereo or multi - view stereo algorithms. Such technologies may include visible or infrared sensors that use passive or active illumination patterns whose patterns can change over time.
Summary of the Invention
[0004] Summary In a general aspect, the method may include obtaining image data representing a plurality of images. Each of the plurality of images includes an image of a scene within an observation frame, and the scene includes a non-rigidly deforming object viewed from respective viewpoints. The method may also include generating a deformation model based on the image data. The deformation model describes the movement performed by the non-rigidly deforming object while the image data was being generated, and the deformation model is represented by a mapping between positions within the observation frame and positions within a canonical frame. The method may further include generating a deformable neural radiance field (D-NeRF) based on the position and viewing direction of a projection ray passing through a position within the canonical frame. The D-NeRF provides a mapping between the position and viewing direction to colors and optical densities at each position within the observation frame. The colors and optical densities at each position within the observation frame enable viewing the non-rigidly deforming object from a new viewpoint.
[0005] In another general aspect, a computer program product comprises a non-transitory storage medium, the computer program product including instructions that, when executed by a processing circuit of a computing device, cause the processing circuit to perform a method. The method may include obtaining image data representing a plurality of images. Each of the plurality of images includes an image of a scene within an observation frame, the scene including a non-rigidly deforming object viewed from respective viewpoints. The method may also include generating a deformation model based on the image data. The deformation model describes the motion performed by the non-rigidly deforming object while the image data was being generated, and the deformation model is represented by a mapping between positions within the observation frame and positions within a canonical frame. The method may further include generating a deformable neural radiance field (D-NeRF) based on the position and viewing direction of a projection ray passing through the position within the canonical frame. The D-NeRF provides a mapping between the position and viewing direction to colors and optical densities at each position within the observation frame. The colors and optical densities at each position within the observation frame enable viewing the non-rigidly deforming object from a new viewpoint.
[0006] In another general aspect, an electronic device includes a memory and a control circuit coupled to the memory. The control circuit may be configured to obtain image data representing a plurality of images. Each of the plurality of images includes an image of a scene within an observation frame, and the scene includes a non-rigidly deforming object viewed from respective viewpoints. The control circuit may also be configured to generate a deformation model based on the image data. The deformation model describes the motion performed by the non-rigidly deforming object while the image data was being generated, and the deformation model is represented by a mapping between positions within the observation frame and positions within a canonical frame. The control circuit may further be configured to generate a deformable neural radiance field (D-NeRF) based on the position and viewing direction of a projection ray passing through the position within the canonical frame. The D-NeRF provides a mapping between the position and the viewing direction to a color and an optical density at each position within the observation frame. The color and the optical density at each position within the observation frame enable viewing the non-rigidly deforming object from a new viewpoint.
[0007] Details of one or more implementations are set forth in the accompanying drawings and the description below. Other features will be apparent from the following description, the accompanying drawings, and the appended claims.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 5
Mode for Carrying Out the Invention
[0009] Detailed Description Conventional techniques for synthesizing a new view of a scene employ Neural Radiance Fields (NeRF). In this technique, a static scene is represented as a continuous 5D function that outputs the radiance emitted in each direction (θ, ψ) at each point (x, y, z) in space, and a per-point density that acts like a differential opacity controlling how much radiance is accumulated by the ray passing through (x, y, z). This technique often optimizes a fully-connected deep neural network without convolutional layers, often referred to as a multilayer perceptron (MLP), to represent the 5D function by regressing from a single 5D coordinate (x, y, z, θ, ψ) to a single volumetric density and a view-dependent RGB color.
[0010] To render this 5D function or NeRF, 1) advance camera rays into the scene to generate a sampling set of 3D points, 2) use these points and their corresponding 2D viewing directions as inputs to a neural network to generate an output set of colors and densities, and 3) using conventional volume rendering techniques, these colors and densities can be accumulated into a 2D image. Since this process is naturally distinguishable, gradient descent can be used to optimize NeRF by minimizing the error between each observed image and the corresponding view rendered from our representation. By minimizing this error across multiple views, the network is encouraged to predict a coherent model of the scene by assigning high volume densities and accurate colors to locations containing the true underlying scene content.
[0011] NeRF is generally excellent at synthesizing views in scenes containing rigid inanimate objects. In contrast, NeRF is not as good at synthesizing views in scenes containing people (more generally, non-rigid objects that tend to move). Technical problems arise when using a handheld camera to model people. This technical problem is difficult due to reasons such as 1) being non-rigid and unable to remain completely stationary, and 2) troublesome materials such as hair, glasses, and earrings that violate the assumptions used in most reconstruction methods. For example, if NeRF is used to synthesize the view of a person taking a photo of themselves (i.e., a selfie) with a mobile phone camera, the synthesized view may contain artifacts resulting from the inaccuracies introduced by NeRF, which cannot handle non-rigid and troublesome materials.
[0012] Unlike conventional approaches to solving the above technical problems, the technical solution to the above technical problems involves generating a motion deformation model that a target experiences in a non-rigid deformation scene that defines how the target is distorted by movement. For example, when an image synthesis system uses NeRF, the system takes in multiple poses of the target as input for training data. Unlike conventional NeRF, the technical solution first represents the position of the target from various viewpoints in the observation frame. The technical solution then involves deriving a deformation model, i.e., a mapping between the observation frame and the canonical frame in which the movement of the target is considered. This mapping is achieved using pose-specific latent deformation codes determined using a multi-layer perceptron (MLP). NeRF is then derived from positions within the canonical frame and projection ray directions using another MLP. A new pose for the target can then be derived using NeRF.
[0013] The technical advantage of the above technical solution is that the technical solution accurately predicts the synthetic view of a new scene without artifacts caused by not considering the movement of the target.
[0014] In some implementations, the deformation model is conditioned on frame-specific latent codes, and the latent codes encode the state of the scene within the frame.
[0015] In some implementations, the deformation model includes a rotation, a center point corresponding to the rotation, and a translation. In some implementations, the rotation is encoded as a pure logarithmic quaternion. In some implementations, the deformation model includes (i) a center transformation for the difference between the position and the similar point, (ii) the center point, and (iii) a sum translation.
[0016] In some embodiments, the deformation model includes a multi-layer perceptron (MLP) within a neural network. In some embodiments, the elastic loss function component for the MLP is based on the norm of the matrix representing the deformation model. In some embodiments, the matrix is the Jacobian of the deformation model with respect to positions in the observation frame. In some embodiments, the elastic loss function component is based on the singular value decomposition of the matrix representing the deformation model. In some embodiments, the elastic loss function component is based on the logarithm of the singular value matrix resulting from the singular value decomposition. In some embodiments, the elastic loss function component is composed of rational functions to generate a robust elastic loss function.
[0017] In some embodiments, the background loss function component requires specifying points in the scene as static points with a penalty on motion. In some embodiments, the background loss function component is based on the difference between the static points and the mapping of the static points in the observation frame to the canonical frame according to the deformation model. In some embodiments, the step of generating the deformation model includes applying position encoding to the position coordinates in the scene to generate a periodic function of position, and the periodic function has a frequency that increases with the training iterations for the MLP. In some embodiments, the periodic function of the position encoding is multiplied by a weight indicating whether the training iteration includes a specific frequency.
[0018] NeRF is a continuous volume representation. This is
[0019] [Number]
[0020] the case. By combining with volume rendering techniques, NeRF can represent scenes with a realistic quality like photos. For this reason, NeRF is constructed to address the problems in realistically capturing people like in photos.
[0021] The NeRF training procedure relies on the fact that, given a 3D scene, two intersecting rays from two different cameras should result in the same color. Ignoring specular reflection and transmission, this assumption holds for all scenes with static structures. Unfortunately, people have been found to lack the ability to remain completely still. This can be verified as follows. If one tries to take a self - portrait video while remaining completely still, it will be found that one's line of sight naturally follows the camera, and even parts that seem to be stationary are moving relative to the background.
[0022] Understanding that there are such limitations, NeRF is extended to enable the reconstruction of non - rigidly deforming scenes. Instead of directly projecting rays through NeRF, a canonical template of the scene is used. This template contains the relative structure and appearance of the scene, and rendering will use a non - rigid transformation version of the said template. Additionally, there are those that can model the deformation for each template and frame, where this deformation is defined on mesh points and voxel grids respectively, while being modeled as a continuous function using an MLP.
[0023] The deformation from observation to canonical is adopted for every frame i ∈ {1,…,n}, where n is the number of observed frames. This defines a mapping T i : x → x′ that maps all observed space coordinates x to canonical space coordinates x′. In practice, the deformation field is modeled for all time steps using a single MLP T: (x, ω i ) → x′ conditioned on the latent code ω i learned for each frame. The per - frame latent code models the state of the scene at that frame. Given the canonical space radiance field F and the observation - canonical mapping T, the observation space radiance field is
[0024]
Equation
[0025] It can be evaluated as. During rendering, the ray and the sample point are simply projected into the viewing frame, and then the deformation field is used to map the sampled point to the point on the template.
[0026] FIG. 1 is a diagram showing an exemplary electronic environment 100 in which the improved technology described above can be realized. As shown in FIG. 1, the exemplary electronic environment 100 includes a computer 120.
[0027] The computer 120 includes a network interface 122, one or more processing units 124, and a memory 126. The network interface 122 includes, for example, an Ethernet (registered trademark) adapter or the like for converting an electronic signal and / or an optical signal received from a network into an electronic format that can be used by the computer 120. The set of processing units 124 includes one or more processing chips and / or assemblies. The memory 126 includes both volatile memory (e.g., RAM) and non-volatile memory such as one or more ROMs, disk drives, solid-state drives, etc. The set of processing units 124 and the memory 126 together form a control circuit configured and arranged to execute various methods and functions as described herein.
[0028] In some embodiments, one or more of the components of the computer 120 may include a processor (e.g., the processing unit 124) configured to process instructions stored in the memory 126. Examples of instructions as shown in FIG. 1 include an image acquisition manager 130, a deformation model manager 140, and a template NeRF manager 150. Further, the memory 126 is configured to store various data described for each manager that uses such data, as shown in FIG. 1.
[0029] The image acquisition manager 130 is configured to acquire image data 132 to be input to the deformation model manager 140. In some implementation examples, the image acquisition manager 130 receives the image data 132 via the network interface 122, that is, via the network. In some implementation examples, the image acquisition manager 130 receives the image data 132 from a local storage (e.g., a disk drive, a flash drive, an SSD, etc.).
[0030] The image data 132 represents a plurality of images of scenes 134(1), 134(2), …, 134(N). For example, the user may generate the images 134(1), 134(2), …, 134(N) by using a mobile phone camera to record their own images, i.e., “selfie photos,” from various viewpoints 136(1), 136(2), …, 136(N).
[0031] Modeling people with a handheld camera is particularly difficult due to both 1) the non-rigidity of not being able to remain completely still and 2) troublesome materials such as hair, glasses, and earrings that violate the assumptions used in most reconstruction methods. To model non-rigid deformation scenes, NeRF can be generalized by introducing additional components. The additional components are the canonical NeRF model that serves as a template for all observations supplemented by a per-observation deformation field that distorts the 3D points in the reference system of the observation to the reference system of the canonical model.
[0032] The deformation model manager 140 is configured to generate a deformation model that provides a mapping between the coordinates in the observation space of the scene and the coordinates in the canonical space to which the NeRF model is applied. For this purpose, the deformation model manager 140 is configured to generate observation frame position data 141 and latent deformation code data 142.
[0033] The observation frame data 141 represents the coordinates of points in the observation frame, that is, the coordinate frames of images 134(1), 134(2), …, 134(N). For example, the observation frame data 141 may represent a point x within a volume representing the extent of the space in which the image data 132 is generated. The observation frame can be visualized in FIG. 2.
[0034] FIG. 2 is a diagram showing an exemplary system architecture 200 for generating a deformable neural radiance field. FIG. 2 shows the observation frame 220 as a set of points within a three-dimensional volume. FIG. 2 also shows the camera viewpoint 210 towards which the ray 212 is directed.
[0035] The latent deformation code data 142 represents the latent deformation code represented by the symbol ω in FIG. 2. Each image 134(1), …, 134(N) is associated with its own latent deformation code. The per-image latent deformation code associated with an image of the scene models the state of the scene in that image. In some implementations, the per-image latent deformation codes are learned. The number of dimensions of each per-image latent deformation code is small, and in some implementations, each latent deformation code has 8 dimensions.
[0036] The deformation model manager 140 is also configured to generate a deformation model based on the observation frame data 141 and the latent deformation code data 142. In some implementations, the deformation model is derived using a neural network. In some implementations, the neural network does not include convolutional layers. As shown in FIG. 2, the deformation model is derived using a multi-layer perceptron (MLP) 230. As shown in FIG. 1, the deformation model is derived using the deformation field MLP data 143.
[0037] The deformation field MLP data 143 represents the values that define the deformation field MLP. An exemplary deformation field MLP in the context of the technical solutions described herein has six layers (one input, one output, and four hidden layers). In this example, the size of the hidden layers (i.e., the number of nodes) is 128, there is a skip connection to the fourth layer, and the Softplus activation function log(1 + e x ) is present. The deformation field MLP data 143 further includes loss function data 144 and dense data 145.
[0038] The deformation model adds ambiguities that can make optimization more difficult. For example, an object moving backward is visually equivalent to an object that is shrinking in size, with a great deal of resolution in between. These ambiguities lead to the problem of insufficient constraints on optimization, resulting in unrealistic deformations and artifacts. Therefore, prior art that provides more reasonable solutions is introduced.
[0039] The loss function data 144 represents the loss function components used to determine the values of the nodes of the deformation field MLP (i.e., MLP230 in FIG. 2) for each training iteration. As shown in FIG. 2, the loss function components for MLP230 include elastic loss data 144(1) and background loss data 144(2).
[0040] The elastic loss data 144(1) represents the value of the elastic loss function used to determine the deformation model. Modeling non-rigid deformations using elastic energy that measures the deviation of local deformations from rigid body motion is common in geometric processing and physical simulations. Such energy has been widely used for the reconstruction and tracking of non-rigid scenes and objects. Therefore, elastic energy is a suitable candidate for such approaches. Elastic energy has been most commonly used for discretized surfaces, such as meshes, but similar concepts can be applied in the context of the continuous deformation fields included in the deformation model.
[0041] a certain latent signature ωi In this case, the continuous deformation field T is
[0042]
Number
[0043] a non - linear mapping to. Nevertheless, such a non - linear mapping can be approximated in matrix form.
[0044] In some embodiments, the non - linear mapping is also differentiable. In this case,
[0045]
Number
[0046] represents a suitable linear approximation of the transformation at that point. Therefore, the local behavior of the deformation model is controllable by the Jacobian of T. Note that, unlike other methods using discretized surfaces, this continuous / differentiable formulation enables the direct calculation of the Jacobian of this mapping via automatic differentiation of the deformation field MLP.
[0047] Jacobian J from the rigid - body transformation T There are several ways to impose a penalty on the deviation of. The Jacobian J T =UΣV T Considering the singular - value decomposition of, several methods penalize the deviation from the closest rotation as
[0048]
Number
[0049] where R = VU T and · F is the Frobenius norm. In some embodiments, the elastic loss component is J TIt is based on the singular values, and the elastic loss component includes a measure of the deviation of the singular value matrix Σ from the identity I. The logarithm of the singular values was found to give equal weight to the stretching and shrinking of the same factors and to function more appropriately. Thus, the point x derived from the deviation of the logarithmic singular values from zero i The following penalty is imposed on the elastic loss component at
[0050]
Equation
[0051] where log represents the matrix logarithm. In some embodiments, the elastic loss component is remapped to a more robust loss function. For example, most of the human body is rigid, but there are some movements that can break our assumptions about local rigid bodies, such as facial expressions that locally stretch or contract the skin. The elastic energy defined above can then be remapped using the robust loss component.
[0052]
Equation
[0053] where ρ(·) is the Geman-McClure robust error function implemented with the hyperparameter c = 0.03, and w i is the weight. Over multiple points, the net robust loss component L elastic-r is the weighted average of the robust loss components at each of those multiple points. The robust loss component reduces the influence of outliers during training by reducing the loss gradient to zero when the value of the argument is large.
[0054] The background loss data 144(2) represents the value of the background loss function used to determine the deformation model. The deformation field T is not constrained, so everything can move freely. In some embodiments, a regularization term is added to prevent the background from moving. Assuming a set of 3D points within a scene known to be static, a penalty can be imposed on any deformation at these points. For example, by aligning cameras using 3D shape reconstruction (structure from motion) from multi-view images, a set of 3D feature points that behave rigidly over at least some sets of observations is generated. Given these static 3D points {x1…,x K , the movement is penalized as
[0055]
Number
[0056] In addition to keeping the background points from moving, this regularization also has the advantage of aligning the observation coordinate frame with the canonical coordinate frame.
[0057] The coarse-to-fine data 145 represents coarse-to-fine deformation regularization. The core component of the NeRF architecture is positional encoding. A similar concept is adopted for the deformation field MLP,
[0058]
Number
[0059] The hyperparameter m controls the number of frequency bands (and thus the highest frequency) used in the mapping. It has been found that this controls the smoothness of the network. The higher the value of m, the more high-frequency details can be modeled, but as a result, there is a possibility of NeRF overfitting and modeled image noise as 3D structures.
[0060] It has been observed that jointly optimizing NeRF with the deformation field gives rise to an optimization problem that tends to result in minima. At the beginning of training, neither NeRF nor the deformation field contains meaningful information. When using a large value for m, this means that the deformation field may overfit to the incomplete NeRF template. For example, when the subject rotates their head sideways, a network using a large m will often choose to keep the head in the forward position and encode the appearance changes using the view direction component of NeRF. On the other hand, when using a small value for m, the network may not be able to model deformations that require high-frequency details such as facial expressions or moving hair strands.
[0061] It has been found that the positional encoding used in NeRF has a simple interpretation regarding the neural tangent kernel (NTK) of NeRF's MLP. As a result, a stationary interpolation kernel is obtained where m controls the adjustable "bandwidth" of its interpolation kernel. When the number of frequencies is small, it results in a wide kernel that underfits the data, while when the number of frequencies is large, it results in a narrow kernel that overfits the data. Considering this, a method has been proposed to smoothly anneal the bandwidth of the NTK by introducing a parameter α that windows the frequency band of the positional encoding. The weights are defined as
[0062]
Number
[0063] and defined as such, where linearly annealing the parameter α ∈ [0, m] can be interpreted as sliding a partially truncated incomplete Hann window across the frequency band (where the left side is fixed at 1 and the right side is fixed at 0). Then, the positional encoding becomes
[0064]
Number
[0065] where t is the current training iteration and N is a hyperparameter as to when α should reach the maximum number of frequency m.
[0066]
Number
[0067] Along these lines, the deformation field MLP data 143 also includes the SE(3) transformation data 146. The SE(3) transformation data 146 is encoded in the MLP and represents the transformation field as described above. The SE(3) transformation data 146 includes rotation data 147 representing the rotation q, and center point s representing center point data 148, and translation data representing the translation t. In some embodiments, the rotation data 147, center the point data 148 and the translation data 149 are represented in quaternion form. In some embodiments, the rotation data 147, center the point data 148, and the translation data 149 are represented in another form, such as matrix form.
[0068] The template NeRF manager 150 is configured to
[0069]
Number
[0070] generate a 5D representation of. In some embodiments, for each image, the appearance code Ψ iis provided to modulate color output and handle appearance changes between input frames, such as exposure and white balance. As shown in FIG. 2, a canonical frame 240 is visualized. Points along a ray 212 within an observation frame 220 are mapped to points along a curve 242 within the canonical frame 240 using a deformation field MLP 230. Each position and ray / camera viewpoint is mapped to color and density along with an appearance code by a NeRF MLP 250.
[0071] As shown in FIG. 1, a template NeRF manager 150 is configured to generate canonical frame position data 151, direction data 152, latent appearance code data 153, and template NeRF MLP data 154 and output output data 160 including color data 162 and density data 163. The canonical frame position data 152 represents positions within the canonical frame mapped from an observation frame using a deformation field MLP (e.g., MLP 230 in FIG. 2). The direction data 152 represents rays passing through each point within the canonical frame or camera angles or direction cosines. The latent appearance code data 153 represents per-image latent appearance codes.
[0072] The template NeRF MLP data 154 represents values defining a NeRF MLP. An exemplary NeRF MLP in the context of the technical solutions described herein has six layers (one input, one output, and four hidden layers). In this example, the size of the hidden layers (i.e., the number of nodes) is 128, there is a skip connection in the fourth layer, and there is a ReLU activation function. The deformation field MLP data 154 further includes color loss function data 155 and coarse density data 156.
[0073] The color loss function data 155 represents values of a color loss function defined as follows. For each optimization iteration, a batch of camera rays is randomly sampled from the set of all pixels within the dataset and then N c samples from the coarse network and N c +N fHierarchical sampling continues to be performed to query samples. The volume rendering procedure is used to render the color of each ray from both sets of samples. Color loss is the sum of squared errors between the rendered pixel color and the true pixel color for both the coarse and dense renderings.
[0074]
Number
[0075] The coarse-to-dense data 156 represents the same coarse-to-dense deformation regularization as the coarse-to-dense data 145. However, in the case of the NeRF MLP, the sine and cosine are not weighted.
[0076] Returning to the elastic loss function, the deformation field T can behave freely within the empty space. This is because an object moving relative to the background may require non-rigid deformation somewhere within the space. Thus, in some embodiments, the elastic loss function is weighted only by the contribution to the rendered view at each point, as follows.
[0077] The 5D neural radiance field represents the scene as the volume density and directional radiance at any point in space. The color of any ray passing through the scene is rendered using the principles from conventional volume rendering.
[0078]
Number
[0079] The function T(t) represents the cumulative transmittance along the ray from t n to t, i.e., the probability that the ray travels from t n to t without colliding with other particles. By rendering views from our continuous neural radiance field, it is necessary to estimate this integral C(r) for the camera rays traced through each pixel of the desired virtual camera.
[0080] This continuous integral is numerically estimated using quadrature methods. Typically, the deterministic quadrature methods used to render a discretized voxel grid would substantially limit the resolution of our representation, as the MLP can only be queried at a fixed set of discrete positions. Instead, we use a hierarchical sampling approach. In this case, the interval [t n , t f is partitioned into M bins placed at equal intervals, and then one sample is randomly drawn uniformly from each bin.
[0081]
Number
[0082] A discrete set of samples is used to estimate the integral, but hierarchical sampling enables the representation of continuous scenes. This is because hierarchical sampling results in the MLP being evaluated at continuous positions over the optimization process. We estimate C(r) as follows using these samples.
[0083]
Number
[0084] Rendering strategies that densely evaluate the neural radiance field network at M query points along each camera ray are inefficient, as free space and occluded regions that do not contribute to the rendered image are still repeatedly sampled. Therefore, instead of representing the scene using a single network, two networks, one "coarse" and the other "fine", are optimized simultaneously. First, N cA set of positions is sampled using hierarchical sampling, and the "coarse" network at these positions is evaluated as described in equations (13), (14), and (15). Based on the output of this "coarse" network, more informative sampling for points is provided along each ray. In this case, the samples are biased towards relevant parts of the volume. For this purpose, first, the alpha compositing color is in equation (14),
[0085] [Number]
[0086] The output data 160 represents the output of the NeRF MLP (i.e., the MLP 250 in Figure 2). The output data 160 includes color data 162 and density data 164. As described above (e.g., in equation (11)), the color data 162 depends on the position and the ray angle, while the density data 164 depends only on the ray angle. It is also possible to obtain a view of the scene from an arbitrary viewpoint from the color data 162 and the density data 164.
[0087] The system 200 shown in Figure 2 may be optimized as a unit, i.e., the MLP 230 and the MLP 250 may be identified by a single loss function that is the sum of the loss components described above.
[0088] [Number]
[0089] where λ and μ are weights. In some embodiments, λ = μ = 10 -3 is true. FIG. 3 is a flowchart illustrating an exemplary method 300 for generating a deformable NeRF. Method 300 may be executed by a software configuration described in connection with FIG. 1 that resides in memory 126 of computer 120 and is executed by a set of processing units 124, or may be executed by a software configuration that resides in memory of a computing device different from computer 120 (e.g., remote from computer 120).
[0090] At 310, image acquisition manager 130 acquires image data (e.g., image data 132) representing a plurality of images (e.g., images 134(1),..., 134(N)), each of the plurality of images including an image of a scene in an observation frame (e.g., observation frame position data 141), the scene including a non-rigidly deformable object viewed from respective viewpoints (e.g., viewpoints 136(1),..., 136(N)).
[0091] At 320, deformation model manager 140 generates a deformation model (e.g., deformation field MLP data 143) based on the image data, the deformation model describing the movement performed by the non-rigidly deformable object while the image data was being generated, the deformation model represented by a mapping between a position in the observation frame and a position in the canonical frame (e.g., canonical frame position data 151).
[0092] At 330, template NeRF manager 150 generates a deformable neural radiance field (D-NeRF) based on the position and viewing direction of projection rays passing through positions in the canonical frame, the D-NeRF providing a mapping between the position and the viewing direction to colors (e.g., color data 162) and optical densities (e.g., density data 164) at each position in the observation frame, the colors and optical densities at each position in the observation frame enabling viewing of the non-rigidly deformable object from a new viewpoint.
[0093] The influence of the deformation field on the object is shown in FIGS. 4A, 4B, and 4C. FIG. 4A is a diagram showing an exemplary posture 400 of a person who is the object in the observation frame. FIG. 4B is a diagram showing an exemplary posture 450 of a person who is the object in the canonical frame. In both FIGS. 4A and 4B, the object is shown with an inset showing the orthographic projection views in the forward and leftward directions. In FIG. 4A (observation frame), note that there are displacements from right to left and from front to back between the observation model and the canonical model, and these displacements are modeled by the deformation field for this observation.
[0094] FIG. 4C is a diagram showing an exemplary interpolation 470 between a keyframe (in white) without elastic regularization (upper) and a keyframe (in white) with elastic regularization (lower). FIG. 4C shows a new figure synthesized without elastic regularization by linearly interpolating the observation deformation codes and a new figure synthesized with elastic regularization. In the absence of elastic regularization, the intermediate state shows distortion, for example, the spacing between facial features has changed from the original image.
[0095] FIG. 5 shows examples of a general-purpose computer device 500 and a general-purpose mobile computer device 550 that can be used with the techniques described herein.
[0096] As shown in FIG. 5, the computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers, etc. The computing device 550 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices, etc. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not meant to limit the embodiments of the invention described and / or claimed herein.
[0097] The computing device 500 includes a processor 502, a memory 504, a storage device 506, a high-speed interface 508 that connects to the memory 504 and a high-speed expansion port 510, and a low-speed interface 512 that connects to a low-speed bus 514 and the storage device 506. Each of the components 502, 504, 506, 508, 510, and 512 can be interconnected using various buses and can be implemented on a common motherboard or in other suitable manners. The processor 502 can process instructions to be executed within the computing device 500, including instructions stored in the memory 504 or the storage device 506, to display graphical information regarding a GUI on an external input / output device such as a display 516 coupled to the high-speed interface 508. In other embodiments, multiple processors and / or multiple buses can be used as appropriate, along with multiple memories and multiple types of memories. Also, multiple computing devices 500 can be connected, and each device can provide a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).
[0098] The memory 504 stores information within the computing device 500. In one embodiment, the memory 504 is one or more volatile memory units. In another embodiment, the memory 504 is one or more non-volatile memory units. The memory 504 can also be another form of computer-readable medium, such as a magnetic disk or an optical disk.
[0099] Storage device 506 can provide large-capacity storage for computing device 500. In one implementation example, storage device 506 can be an array of devices including a computer-readable medium, such as a floppy (registered trademark) disk device, a hard disk device, an optical disk device, or a tape device, etc., or a flash memory or other similar solid-state memory device, or a storage area network or a device within other configurations, or can include them. A computer program product can be tangibly embodied in an information carrier. The computer program product can also include instructions that, when executed, perform one or more methods such as those described above. The information carrier is a computer-readable medium or a machine-readable medium such as memory 504, storage device 506, or memory on processor 502.
[0100] High-speed controller 508 manages bandwidth-intensive operations for computing device 500, and low-speed controller 512 manages lower bandwidth-intensive operations. Such an assignment of functions is merely an example. In one implementation example, high-speed controller 508 is coupled to memory 504, display 516 (e.g., via a graphics processor or an accelerator), and is also coupled to a high-speed expansion port 510 that can receive various expansion cards (not shown). In this implementation example, low-speed controller 512 is coupled to storage device 506 and a low-speed expansion port 514. The low-speed expansion port, which can include various communication ports (e.g., USB, Bluetooth (registered trademark), Ethernet (registered trademark), wireless Ethernet (registered trademark)), can be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or a router, etc., via, for example, a network adapter.
[0101] As shown in the figure, the computing device 500 can be implemented in several different forms. The computing device 500 can be implemented, for example, as a standard server 520 or multiple times within a group of such servers. Also, the computing device 500 may be implemented as part of a rack server system 524. In addition, the computing device 500 can be implemented in a personal computer such as a laptop computer 522. Alternatively, components from the computing device 500 may be combined with other components within a mobile device (not shown) such as device 550. Each of such devices may include one or more of the computing devices 500, 550, and the entire system may be composed of a plurality of computing devices 500, 550 that communicate with each other.
[0102] The computing device 550 includes, among other components, a processor 552, a memory 564, input / output devices such as a display 554, a communication interface 566, and a transceiver 568. The device 550 may also include a storage device such as a microdrive or other device to provide additional storage. Each of the components 550, 552, 564, 554, 566, and 568 are interconnected using various buses, and some of these components may be implemented on a common motherboard or otherwise as appropriate.
[0103] The processor 552 can execute instructions, including instructions stored in the memory 564, within the computing device 450. The processor may be implemented as a chipset of chips including separate multiple analog and digital processors. The processor can provide coordination of other components of the device 550, such as, for example, control of the user interface, applications executed by the device 550, and wireless communication by the device 550.
[0104] Processor 552 may communicate with a user via a control interface 558 and a display interface 556 coupled to a display 554. The display 554 may be, for example, a Thin-Film-Transistor Liquid Crystal Display (TFT LCD) or an Organic Light Emitting Diode (OLED) display, or other suitable display technology. The display interface 556 may comprise appropriate circuitry for driving the display 554 to present graphical and other information to the user. The control interface 558 may receive commands from the user and convert the commands for presentation to the processor 552. Additionally, an external interface 562 may be provided to communicate with the processor 552 to enable short-range communication between the device 550 and other devices. The external interface 562 may provide, for example, wired communication in some implementations, or wireless communication in other implementations, and multiple interfaces may be used.
[0105] Memory 564 stores information within computing device 550. Memory 564 can be implemented as one or more of one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. Extended memory 574 may be provided and can be connected to device 550 via an expansion interface 572 that can include, for example, a Single In Line Memory Module (SIMM) card interface. Such extended memory 574 can provide additional storage space for device 550 or store applications or other information for device 550. Specifically, extended memory 574 can include instructions for executing or supplementing the processes described above and can also include security-protected information. Thus, for example, extended memory 574 may be provided as a security module for device 550 and may be programmed with instructions that enable secure use of device 550. In addition, security-protected applications can be provided via the SIMM card along with additional information, such as by placing identification information on the SIMM card in a non-hackable manner.
[0106] The memory can include, for example, flash memory and / or NVRAM memory as described below. In one implementation example, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, execute one or more methods such as the methods described above. The information carrier is a computer-readable medium or machine-readable medium such as memory 564, extended memory 574, or memory on processor 552 and can be received, for example, via transceiver 568 or external interface 562.
[0107] Device 550 may communicate wirelessly via a communication interface 566 that may include a digital signal processing circuit as needed. The communication interface 566 can provide communication under various modes or protocols such as, among others, GSM (registered trademark) voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA (registered trademark), CDMA2000, or GPRS. Such communication may be performed, for example, through a radio frequency transceiver 568. In addition, short-range communication may be performed using Bluetooth (registered trademark), WiFi, or other such transceivers (not shown). In addition, a Global Positioning System (GPS) receiver module 570 can provide the device 550 with additional navigation-related and location-related wireless data that can be appropriately used by applications executed on the device 550.
[0108] Device 550 may also communicate in a voice-recognizable manner using a voice codec 560, which can receive speech information from the user and convert it into usable digital information. The voice codec 560 can similarly generate audible sounds for the user, for example, through a speaker within the handset of the device 550. Such sounds may include sounds from a voice telephone call, recorded sounds (e.g., voice messages, music files, etc.), or sounds generated by an application operating on the device 550.
[0109] As shown in the figure, the computing device 550 can be implemented in several different forms. For example, the computing device 550 may be implemented as a mobile phone 580. The computing device 550 may also be implemented as part of a smartphone 582, a personal digital assistant, or other similar mobile device.
[0110] Various implementations of the systems and techniques described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed application specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs that are executable and / or interpretable on a programmable system that includes at least one programmable processor coupled to receive and transmit data and instructions from and to a storage system, at least one input device, and at least one output device, which programmable processor may be either special purpose or general purpose.
[0111] (Also known as a program, software, software application, or code) These computer programs include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine language. The terms "machine-readable medium" and "computer-readable medium," as used herein, refer to any computer program product, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0112] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to the user and a keyboard and pointing device (e.g., a mouse or trackball) that the user can use to provide input to the computer. Interaction with the user can be performed using other types of devices. For example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and the input from the user can be received in any form, including acoustic input, voice input, or tactile input.
[0113] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or includes middleware components (e.g., an application server), or includes front-end components (e.g., a client computer having a graphical user interface or a web browser that enables a user to interact with an implementation of the systems and techniques described herein), or includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0114] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact with each other through a communication network. The relationship between the client and the server is created by computer programs that are executed on respective computers and have a client-server relationship with each other.
[0115] Returning to FIG. 1, in some embodiments, the memory 126 can be any type of memory, such as random access memory, disk drive memory, flash memory, etc. In some embodiments, the memory 126 can be implemented as a plurality of memory components (e.g., a plurality of RAM components or disk drive memory) associated with the components of the compression computer 120. In some embodiments, the memory 126 can be a database memory. In some embodiments, the memory 126 can be or can include non-local memory. For example, the memory 126 can be or can include memory shared by a plurality of devices (not shown). In some embodiments, the memory 126 can be associated with a server device (not shown) within a network and can be configured to function for the components of the compression computer 120.
[0116] The components of the compression computer 120 (e.g., modules, processing unit 124) can be configured to operate based on one or more platforms (e.g., one or more similar platforms or different platforms) that can include one or more types such as hardware, software, firmware, operating system, runtime library, etc. In some embodiments, the components of the compression computer 120 can be configured to operate within a cluster of devices (e.g., a server farm). In such embodiments, the functions and processing of the components of the compression computer 120 can be distributed among some of the devices in the cluster of devices.
[0117] The components of computer 120 can be, or can include, any type of hardware and / or software configured to process attributes. In some implementations, one or more portions of the components shown for computer 120 in FIG. 1 can be, or can include, a hardware-based module (e.g., a digital signal processor (DSP), a field programmable gate array (FPGA), memory), a firmware module, and / or a software-based module (e.g., a module of computer code, a set of computer-readable instructions executable on a computer). For example, in some implementations, one or more portions of the components of computer 120 can be, or can include, a software module configured for execution by at least one processor (not shown). In some implementations, the functionality of the components can be included in different modules and / or different components than those shown in FIG. 1.
[0118] Although not shown, in some implementations, the components (or portions thereof) of computer 120 may be configured to operate, for example, within a data center (such as a cloud computing environment), a computer system, one or more servers / host devices, etc. In some implementations, the components (or portions thereof) of computer 120 may be configured to operate within a network. Thus, the components (or portions thereof) of computer 120 can be configured to function within various types of network environments that may include one or more devices and / or one or more server devices. For example, the network can be, or can include, a local area network (LAN), a wide area network (WAN), etc. The network can be, or can include, a wireless network and / or can be a wireless network implemented using, for example, a gateway device, a bridge, a switch, etc. The network can include one or more segments and / or can have portions based on various protocols such as Internet Protocol (IP) and / or proprietary protocols. The network can include at least a portion of the Internet.
[0119] In some embodiments, one or more of the components of computer 120 can be, or can include, a processor configured to process instructions stored in a memory. For example, depth image manager 130 (and / or a portion thereof), viewpoint manager 140 (and / or a portion thereof), ray casting manager 150 (and / or a portion thereof), SDV manager 160 (and / or a portion thereof), aggregation manager 170 (and / or a portion thereof), root finder manager 180 (and / or a portion thereof), and depth image generation manager 190 (and / or a portion thereof) can be a combination of a processor and a memory configured to execute instructions related to a process for realizing one or more functions.
[0120] Although numerous embodiments have been described, it will be understood that various changes may be made without departing from the spirit and scope of this specification.
[0121] Also, when an element is referred to as being on, connected to, electrically connected to, coupled to, or electrically coupled to another element, it should be understood that the element may be directly on, connected to, or coupled to the other element, or there may be one or more intervening elements. In contrast, when an element is referred to as being directly on, directly connected to, or directly coupled to another element, there are no intervening elements. The expression "directly on," "directly connected," or "directly coupled" may not be used throughout the detailed description, but an element shown as being directly on, directly connected, or directly coupled may be referred to as such. The claims of this application may be amended to recite the exemplary relationships described herein or shown in the figures.
[0122] Some features of the described implementations are illustrated as described herein, but those skilled in the art will envision many modifications, alternatives, variations, and equivalents. Accordingly, it should be understood that the appended claims are intended to cover all such modifications and variations as fall within the scope of the implementations. These are presented by way of example only and not by way of limitation, and it should be understood that various changes in form and detail may be made. Any part of the apparatus and / or method described herein may be combined in any combination, except mutually exclusive combinations. The implementations described herein may include various combinations and / or partial combinations of the functions, components, and / or features of the various implementations described.
[0123] In addition, the logical flow shown in the figures does not require the specific order or sequential order shown to achieve the desired result. In addition, other steps may be provided, or steps may be eliminated from the described flow, other components may be added to or removed from the described system. Accordingly, other embodiments are within the scope of the appended claims.
Claims
A method of generating a system for synthesizing a new view of a scene, which is implemented by a computer, comprising: obtaining image data representing a plurality of images, each of the plurality of images including an image of a scene, the scene including objects viewed from each of a plurality of points, the method further comprising: generating a deformation model based on the image data, the deformation model describing the movement performed by the object while the image data was being generated, the deformation model mapping the position of each point in the scene to a distorted position by the movement performed by the object, the method further comprising: generating a neural radiance field (NeRF) model that outputs the new view from a certain viewpoint of the scene based on the position and viewing direction of each point of the plurality of points of the scene having positions distorted by the movement performed by the object described by the deformation model. **Claim 2** The method according to claim 1, wherein the deformation model is modeled for a plurality of time steps using a latent deformation code associated with a frame giving a view of the scene from the certain viewpoint. **Claim 3** The method according to claim 1, wherein the deformation model includes a rotation, a center point corresponding to the rotation, and a translation. **Claim 4** The method according to claim 3, wherein the rotation is encoded as the logarithm of a quaternion. **Claim 5** The method according to claim 3, wherein the deformation model includes (i) a similarity transformation for the difference between the position of each point of the plurality of points of the scene and the center point, (ii) the center point, and (iii) the sum of the translations. **Claim 6** The method according to claim 1, wherein the deformation model includes a multilayer perceptron (MLP) in a neural network.
7. The method according to claim 6, wherein the loss function for the MLP is based on the norm of the matrix representing the deformation model.
8. The method according to claim 7, wherein the matrix is a Jacobian.
9. The method according to claim 7, wherein the loss function is based on the norm of the second of three matrices obtained by singular value decomposition of the matrix representing the deformation model.
10. The method according to claim 9, wherein the loss function is based on the logarithm of the norm of the second of three matrices obtained by the singular value decomposition.
11. The method according to claim 7, wherein the loss function is composed of rational functions.
12. The method according to claim 6, wherein the background loss function component for the MLP includes designating points in the scene as static points with a penalty related to movement.
13. The method according to claim 12, wherein the background loss function component is based on the difference between a static point and the point to which the static point is mapped according to the deformation model.
14. The method according to claim 6, wherein the step of generating the deformation model includes applying positional encoding to the position coordinates in the scene to generate a periodic function of position, and the periodic function has a frequency that increases with the training iterations for the MLP.
15. The method according to claim 14, wherein the step of generating the deformation model includes multiplying the periodic function of the positional encoding by a weight indicating whether the training iteration includes a specific frequency.
16. A computer program which, when executed by a processing circuit of a computing device, causes the processing circuit to perform the method according to any one of claims 1 to 15.
17. An electronic device, a memory for storing the computer program according to claim 16, and a control circuit coupled to the memory, the control circuit being configured to execute the computer program stored in the memory. An electronic device.