Deformable neural radiation field
By establishing a deformation model between the observation frame and the canonical frame, and optimizing it using a multilayer perceptron and loss function, the artifact problem in the synthesized view of non-rigid deformable objects using NeRF is solved, and more accurate image synthesis is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2021-01-14
- Publication Date
- 2026-05-12
AI Technical Summary
Existing Neural Radiation Field (NeRF) technology struggles to accurately synthesize views when dealing with non-rigid deformable objects, especially the human body, often resulting in artifacts introduced due to the inability to account for subject movement.
A deformable model is generated, and the deformable neural radiation field (D-NeRF) is derived by observing the mapping between the frame and the canonical frame. A multilayer perceptron (MLP) is used to describe the movement of the subject, and the elastic loss function and background loss function are combined for optimization to ensure the accuracy of the synthesized view.
Accurate prediction of new composite views of non-rigid deformation objects avoids artifacts introduced by subject movement, thus improving the accuracy of image synthesis.
Smart Images

Figure CN116324895B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application is a non-provisional application filed on November 16, 2020, entitled “DEFORMABLE NEURAL RADIANCEFIELDS”, and claims priority thereto, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] This description relates to image synthesis using neural radiation fields (NeRF). Background Technology
[0004] Some computers configured to render computer graphics objects can render objects at a specified viewpoint given multiple existing views. For example, given several depth and color images captured from a camera about a scene including such a computer graphics object, the target could be to synthesize new views of the scene from different viewpoints. The scene could be realistic, in which case physical color and depth sensors are used to capture the views, or it could be synthetic, in which case rendering algorithms such as rasterization are used to capture the views. For realistic scenes, many depth sensing techniques exist, such as time-of-flight sensors, structured light-based sensors, and stereo or multi-view stereo algorithms. These techniques can involve visible or infrared sensors with passive or active illumination patterns, where these patterns can vary over time. Summary of the Invention
[0005] In one general aspect, a method may include acquiring image data representing a plurality of images, each of which includes an image of a scene within a viewing frame, the scene comprising non-rigid deformable objects viewed from a corresponding viewpoint. The method may further include generating a deformable model based on the image data, the deformable model describing the movement performed by the non-rigid deformable objects during the generation of the image data, the deformable model being represented by a mapping between positions within the viewing frame and positions within the canonical frame. The method may further include generating a deformable neural radiation field (D-NeRF) based on the position and viewing direction of a projection ray passing through a position within the canonical frame, the D-NeRF providing a mapping between the position and viewing direction to color and optical density at each position within the viewing frame, the color and optical density at each position within the viewing frame enabling viewing of the non-rigid deformable objects from new perspectives.
[0006] In another general aspect, a computer program product includes a non-transferable storage medium, the computer program product including code that, when executed by processing circuitry of a computing device, causes the processing circuitry to perform a method. The method may include acquiring image data representing a plurality of images, each of the plurality of images including an image of a scene within a viewing frame, the scene including a non-rigid deformable object observed from a corresponding viewpoint. The method may further include generating a deformable model based on the image data, the deformable model describing the movement performed by the non-rigid deformable object during the generation of the image data, the deformable model being represented by a mapping between positions in the viewing frame and positions in a canonical frame. The method may further include generating a deformable neural radiation field (D-NeRF) based on the position and viewing direction of a projected ray passing through a position in the canonical frame, the D-NeRF providing a mapping between the position and viewing direction to color and optical density at each position in the viewing frame, the color and optical density at each position in the viewing frame enabling observation of the non-rigid deformable object from a new viewpoint.
[0007] In another general aspect, an electronic device includes a memory and control circuitry coupled to the memory. The control circuitry can be configured to acquire image data representing a plurality of images, each image including an image of a scene within a viewing frame, the scene including a non-rigid deformable object observed from a corresponding viewpoint. The control circuitry can also be configured to generate a deformable model based on the image data, the deformable model describing the movement performed by the non-rigid deformable object during the generation of the image data, the deformable model being represented by a mapping between positions in the viewing frame and positions in a canonical frame. The control circuitry can also be configured to generate a deformable neural radiation field (D-NeRF) based on the position and viewing direction of a projection ray passing through a position in the canonical frame, the D-NeRF providing a mapping between the position and viewing direction to color and optical density at each position in the viewing frame, the color and optical density at each position in the viewing frame enabling observation of the non-rigid deformable object from a new viewpoint.
[0008] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features will become apparent from the specification, the drawings, and the claims. Attached Figure Description
[0009] Figure 1 This is a diagram illustrating an example electronic environment used to implement the technical solutions described herein.
[0010] Figure 2 The diagram is used in Figure 1 A diagram of an example system architecture for generating deformable neural radiation fields within an electronic environment is shown.
[0011] Figure 3 The diagram is in Figure 1The flowchart illustrates an example method for implementing a technical solution within an electronic environment.
[0012] Figure 4A This is an illustration of an example pose of a human subject within an observation frame.
[0013] Figure 4B It is a diagram illustrating an example pose of a human subject within a pictorial specification framework.
[0014] Figure 4C This is a diagram illustrating an example interpolation between the keyframes with and without elastic regularization (top) and with elastic regularization (bottom).
[0015] Figure 5 The illustrations show examples of computer devices and mobile computer devices that can be used with the circuits described herein. Detailed Implementation
[0016] A conventional approach to synthesizing new views of a scene employs Neural Radiation Fields (NeRF). In this approach, the static scene is represented as a continuous five-dimensional function whose output is along each direction at every point (x, y, z) in space. The emitted radiation and its density at each point act as differential opacity, controlling how much radiation accumulates through the (x, y, z) coordinates. This method utilizes a single 5D coordinate system... A deeply fully connected neural network that reverts to a single volume density and view-dependent RGB color to optimize the representation of a five-dimensional function without any convolutional layers is often called a multilayer perceptron (MLP).
[0017] To render this five-dimensional function, or NeRF, one can: 1) allow camera rays to travel through the scene to generate a sample set of 3D points; 2) use these points and their corresponding 2D view orientations as input to a neural network to produce a set of color and density outputs; and 3) accumulate these colors and densities into a 2D image using classic volumetric rendering techniques. Because this process is naturally differentiable, gradient descent can be used to optimize NeRF by minimizing the error between each observed image and the corresponding view rendered from our representation. Minimizing this error across multiple views encourages the network to predict a coherent model of the scene by assigning high volumetric density and accurate colors to locations containing realistic underlying scene content.
[0018] NeRF typically excels at synthesizing views in scenes involving rigid, inanimate objects. Conversely, it struggles with synthesizing views in scenes containing non-rigid objects, which are generally prone to movement. The technical challenge involves modeling people using a handheld camera. This is challenging due to 1) non-rigidity—they cannot remain perfectly still—and 2) challenging materials such as hair, glasses, and earrings, which violate assumptions used in most reconstruction methods. For example, when NeRF is used to synthesize views of people by taking a selfie using a mobile phone camera, the synthesized view may have artifacts resulting from inaccuracies introduced by NeRF's inability to handle non-rigid and challenging materials.
[0019] Compared to conventional methods for solving the aforementioned technical problems, this technical solution involves generating a deformation model of the movement experienced by a subject in a non-rigid deformation scene, defining how the movement distorts the subject. For example, when an image synthesis system uses NeRF, the system takes multiple poses of the subject as input for training data. Compared to conventional NeRF, this solution first represents the subject's position from different angles within the viewing frame. Then, this solution involves deriving a deformation model, i.e., a mapping between a canonical frame that considers the subject's movement and the viewing frame. This mapping is accomplished using latent deformation codes for each pose determined using a multilayer perceptron (MLP). The NeRF is then derived from the position and projection ray direction within the canonical frame using another MLP. The NeRF can then be used to derive new poses for the subject.
[0020] The technical advantage of the above-mentioned solution is that it accurately predicts the new composite view of the scene without introducing artifacts due to the inability to consider the movement of the subject.
[0021] In some implementations, the deformation model is tuned based on a latent code per frame that encodes the state of the scene within the frame.
[0022] In some implementations, the deformation model includes rotation, a pivot point corresponding to the rotation, and translation. In some implementations, the rotation is encoded as a pure logarithmic quaternion. In some implementations, the deformation model includes the sum of: (i) a similarity transformation with respect to the difference between the position and the pivot point, (ii) the pivot point, and (iii) the translation.
[0023] In some implementations, the deformable model includes a multilayer perceptron (MLP) within a neural network. In some implementations, the elasticity loss function component for the MLP is based on the norm of the matrix representing the deformable model. In some implementations, the matrix is a Jacobian representation of the deformable model relative to its position within the observation frame. In some implementations, the elasticity loss function component is based on the singular value decomposition of the matrix representing the deformable model. In some implementations, the elasticity loss function component is based on the logarithm of the singular value matrix produced by the singular value decomposition. In some implementations, the elasticity loss function component consists of rational functions to produce a robust elasticity loss function.
[0024] In some implementations, the background loss function component involves designating points in the scene as static points with a penalty for movement. In some implementations, the background loss function component is based on the difference between the static point and the mapping from the static point in the observation frame to the canonical frame according to the deformable model. In some implementations, generating the deformable model includes applying position encoding to position coordinates within the scene to produce a periodic function of position, which has a frequency that increases with training iterations of the MLP. In some implementations, the periodic function of position encoding is multiplied by weights indicating whether training iterations include a specific frequency.
[0025] NeRF is a continuous volumetric representation. It is a function F: (x, d) → (c, σ) that maps 3D position x = (x, y, z) and viewing direction d = (φ, θ) to RGB color c = (r, g, b) and density F. Combined with volumetric rendering techniques, NeRF can represent scenes with photorealistic quality. To this end, NeRF was developed to address the problem of photorealistic human capture.
[0026] The NeRF training process relies on the fact that, given a 3D scene, two intersecting rays from two different cameras should produce the same color. Ignoring specular reflection and transmission, this assumption holds true for all scenes with static structures. Unfortunately, it has been found that humans lack the ability to remain still. This can be verified as follows: when a person attempts to film themselves while remaining completely still, it will be observed that their gaze naturally follows the camera, and even parts that the person perceives as still are actually moving relative to the background.
[0027] Understanding these limitations, NeRF was extended to allow the reconstruction of non-rigidly deformable scenes. Instead of directly projecting rays through NeRF, it is used as a standard template for the scene. This template contains the relative structure and appearance of the scene, and the rendering will use a non-rigidly transformed version of the template. Alternatively, both the template and per-frame deformation can be modeled, but the deformation is defined on mesh points and voxel meshes respectively, and it is modeled as a continuous function using an MLP.
[0028] For each frame i ∈ {1, ..., n}, an observed canonical deformation is employed, where n is the number of observed frames. This defines a mapping T that maps all observed spatial coordinates x to canonical spatial coordinates x′. i :x→x′. In practice, a single MLPT is used: (x, ω i →x′ models the deformation field over all time steps, and this single MLP is based on the latent code ω learned per frame. i To regulate. The latent code for each frame models the state of the scene within that frame. Given the canonical space radiation field F and the observation-to-canonical mapping T, the observation-space radiation field can be evaluated as...
[0029] G(x, d, ω) i )=F(T(x,ω i ), d). (1)
[0030] During rendering, rays and sampling points are simply projected into the observation frame, and then the deformation field is used to map the sampling points onto points on the template.
[0031] Figure 1 This is a diagram illustrating an example electronic environment 100 in which the aforementioned improved techniques can be implemented. (See diagram for example.) Figure 1 As shown, the example electronic environment 100 includes a computer 120.
[0032] Computer 120 includes a network interface 122, one or more processing units 124, and memory 126. Network interface 122 includes, for example, an Ethernet adapter, etc., for converting electronic and / or optical signals received from network 150 into an electronic form for use by computer 120. The collection of processing units 124 includes one or more processing chips and / or components. Memory 126 includes volatile memory (e.g., RAM) and non-volatile memory, such as one or more ROMs, disk drives, solid-state drives, etc. The collection of processing units 124 and memory 126 together form control circuitry configured and arranged to perform the various methods and functions described herein.
[0033] In some embodiments, one or more components of the computer 120 may include a processor (e.g., processing unit 124) configured to process instructions stored in memory 126. Figure 1 Examples of such instructions described include an image acquisition manager 130, a variant model manager 140, and a template NeRF training manager 150. Additionally, as... Figure 1 As shown, memory 126 is configured to store various types of data, and the corresponding management of such data is described in relation to the use of such data.
[0034] Image acquisition manager 130 is configured to acquire image data 132 for input into deformable model manager 140. In some embodiments, image acquisition manager 130 receives image data 132 via network interface 122, i.e., via a network. In some embodiments, image acquisition manager 130 receives image data 132 from local storage (e.g., disk drive, flash drive, SSD, etc.).
[0035] Image data 132 represents multiple images 134(1), 134(2), ..., 134(N) of a scene. For example, a user can generate images 134(1), 134(2), ..., 134(N) by recording images from different perspectives 136(1), 136(2), ..., 136(N) using a mobile phone camera, i.e., "selfie".
[0036] Modeling people using handheld cameras is particularly challenging due to 1) their non-rigid nature—they cannot remain perfectly still—and 2) challenging materials such as hair, glasses, and earrings, which violate the assumptions used in most reconstruction methods. To model scenes with non-rigid deformation, NeRF can be generalized by introducing the following additional component: a canonical NeRF model that serves as a template for all observations, supplemented by a deformation field for each observation that warps 3D points in the observation's reference frame to the canonical model's reference frame.
[0037] The Deformation Model Manager 140 is configured to generate a deformation model that provides a mapping between the scene's view space and the canonical space where the NeRF model is applied. To this end, the Deformation Model Manager 140 is configured to generate view frame position data 141 and potential deformation code data 142.
[0038] The observation frame data 141 represents the coordinates of points within the observation frame, i.e., the coordinate frame of images 134(1), 134(2), ..., 134(N). For example, the observation frame data 141 can represent a point x within the volume representing the spatial extent of the generated image data 132. The observation frame can be... Figure 2 Visualization in Chinese.
[0039] Figure 2 This is a diagram illustrating an example system architecture 200 used to generate deformable neural radiation fields. Figure 2 The observation frame 220 is shown as a collection of points in a three-dimensional volume. Figure 2 The camera viewpoint 210 of the guide light 212 is also shown.
[0040] Potential transformation code data 142 indicates that it is composed of Figure 2The symbol ω represents the latent deformable code shown. Each image 134(1), 134(2), ..., 134(N) is associated with its own latent deformable code. The latent deformable code of each image associated with an image of the scene simulates the state of the scene in that image. In some implementations, the latent deformable code of each image is learned. Each image latent deformable code has a low number of dimensions; in some implementations, each latent deformable code has eight dimensions.
[0041] The deformable model manager 140 is also configured to generate deformable models based on observation frame data 141 and potential deformable code data 142. In some embodiments, a neural network is used to derive the deformable model. In some embodiments, the neural network does not include convolutional layers. Figure 2 As shown, a multilayer perceptron (MLP) 230 is used to derive the deformation model. Figure 1 As shown, the deformation model is derived using deformation field MLP data 143.
[0042] The deformable field MLP data 143 represents the value defining the deformable field MLP. The example deformable field MLP in the context of the technique described in this paper has six layers (one input, one output, and four hidden layers). In this example, the size of the hidden layers (i.e., the number of nodes) is 128, there is a skip connection in the fourth layer, and the Softplus activation function is log(1+e^(-1 / 2)). x The deformation field MLP data 143 also includes loss function data 144 and coarse-to-fine data 145.
[0043] Deformation models introduce ambiguity that can make optimization more challenging. For example, an object moving backward is visually equivalent to a scaled-down object, with an infinite number of solutions between them. This ambiguity leads to unconstrained optimization problems that produce unavoidable deformations and artifacts. Therefore, priors leading to more feasible solutions are introduced.
[0044] Loss function data 144 represents the loss function used to determine the deformation field MLP (i.e., ...) for each training iteration. Figure 2 The loss function components of the node values in MLP230. For example... Figure 2 As shown, the loss function components of MLP 230 include elastic loss data 144(1) and background loss data 144(2).
[0045] Elastic loss data 144(1) represents the value of the elastic loss function used to determine the deformation model. In geometric processing and physical simulations, elastic energy, which measures the deviation of local deformation from rigid motion, is commonly used to model non-rigid deformation. This energy has been widely used to reconstruct and track non-rigid scenes and objects; therefore, elastic energy is a good candidate for this approach. While elastic energy is most commonly used for discrete surfaces such as meshes, similar concepts can be applied to scenarios including continuous deformation fields in the deformation model.
[0046] For a fixed potential code ω i The continuous deformation field T is from observation coordinates to The nonlinear mapping of the normalized coordinates. However, this nonlinear mapping can be approximated using matrix representation.
[0047] In some implementations, the nonlinear mapping is also differentiable. In this case, at the point... The Jacobian expression J of the nonlinear mapping at that point T (x) describes a good linear approximation of the transformation at that point. Therefore, the local behavior of the deformation model can be controlled by the Jacobian of T. Note that, unlike other methods using discrete surfaces, this continuous / differentiable formulation allows the Jacobian of this mapping to be directly computed via automatic differentiation of the deformation field MLP.
[0048] There are several ways to punish Jacobi J. T The deviation from the rigid transformation. Consider the Jacobian J. T =UΣV T Singular value decomposition, multiple methods penalize the deviation from the nearest rotation as Where R = VU T and· F This is the Frobenius norm. In some implementations, the elastic loss component is based on J... T The singular values: The elastic loss component comprises a measure of the deviation of the singular value matrix Σ from the unit I. The logarithm of the singular values assigns equal weights to contractions and expansions of the same factor and is found to perform better. Therefore, at point x i The elastic loss component derived from the deviation of the logarithmic singularity from zero is penalized as follows:
[0049] Where log represents the matrix logarithm.
[0050] In some implementations, the elastic loss component is remapped to a more robust loss function. For example, although the human body is generally rigid, there are movements that can break our assumption of local rigidity, such as facial expressions involving localized stretching and compression of our skin. The elastic energy defined above can then be remapped using the following robust loss component:
[0051] L elastic-r (x i ) = w i ρ(||log∑ F ||,c), (3)
[0052]
[0053] Where ρ(·) is the Geman-McClure robust error function implemented using the hyperparameter c = 0.03, and w i It is a weight. At multiple points, the net robust loss component L elastic-r It is the weighted average of the robust loss components at each of those multiple points. For larger values of the independent variable, the robust loss components cause the gradient of the loss to decrease to zero, thereby reducing the impact of outliers during training.
[0054] Background loss data 144(2) represents the value of the background loss function used to determine the deformation model. The deformation field T is unconstrained, so everything can move freely. In some implementations, a regularization term is added to prevent background movement. Given a set of known static 3D points in the scene, any deformation at these points can be penalized. For example, camera registration from a moving structure produces a set of 3D feature points that behave rigidly on at least some observation sets. Given these static 3D points {x1, x2, ..., x3}... K}, the movement penalty is
[0055]
[0056] In addition to keeping the background point stationary, this regularization also has the benefit of aligning the observation coordinate system with the normal coordinate system.
[0057] The data 145 represents coarse-to-fine deformation regularization. The core component of the NeRF architecture is position encoding. The deformation field MLP uses a similar concept; defined as γ(x) = (x, ..., sin(2...)). k πx), cos(2 k Functions of πx, ... Used with k∈{0, ..., m-1}. This function projects a set of sine and cosine functions with increasing frequencies onto a position vector in a high-dimensional space. The hyperparameter m controls the number of frequency bands used in the mapping (and thus the highest frequency). This has been shown to control the smoothness of the network: higher m values allow for modeling higher frequency details, but may also lead to NeRF overfitting and modeling image noise as a 3D structure.
[0058] It was observed that jointly optimizing the NeRF along with the deformation field leads to an optimization problem prone to local minimization. Early in training, neither the NeRF nor the deformation field contains meaningful information. Using large values of m means the deformation field can overfit to an incomplete NeRF template. For example, if the subject rotates their head laterally, a network using a large m will often choose to keep the head in a forward position and use the view orientation component of the NeRF to encode the change in appearance. On the other hand, using small values of m means the network will fail to model deformations requiring high-frequency details—such as facial expressions or a moving strand of hair.
[0059] It has been shown that positional encoding used in NeRF has a convenient interpretation in terms of the neural tangential kernel (NTK) of NeRF's MLP: it results in a fixed interpolation kernel, where m controls the adjustable "bandwidth" of this interpolation kernel. A small number of frequencies result in a wide kernel, which leads to underfitting of the data, while a large number of frequencies result in a narrow kernel, which leads to overfitting of the data. For this reason, a method for smoothly annealing the bandwidth of the NTK is proposed by introducing a parameter α for windowing the frequency bands of positional encoding. The weights for each frequency band j used for positional encoding are defined as...
[0060]
[0061] In the case of linear annealing, the parameter α∈[0, m] can be interpreted as a sliding truncated Hann window over the frequency band (where the left side is clamped to 1 and the right side is clamped to 0). The positional encoding is then defined as γ. α (x)=(x,…,a k (α)sin(2 k πx), a k (α)cos(2 k πx), ...). During training, Where t is the current training iteration, and N is the maximum number of hyperparameters used to determine when α should reach frequency m.
[0062] The simplest version of the deformation uses a translation vector field V: (x, ω) i )→t, the deformation is defined as T(x, ω) i )=x+V(x,ω iThis formula is sufficient to represent all continuous deformations. However, rotating a set of points using a translation field requires different translations for each point, making it difficult to rotate blocks of the scene simultaneously. Therefore, this deformation uses a dense SE(3) field to formulate W:(x, ω) i → SE(3). The SE(3) transform encodes rigid motion, allowing rotation of the far point set using the same parameters. The SE(3) transform is encoded as a rotation q with pivot point s, followed by a translation t. This rotation is encoded as a pure logarithmic quaternion p = (0, v), whose exponent is guaranteed to be a unit quaternion, thus making it a valid rotation:
[0063]
[0064] Note that this can also be viewed as an axis-angle representation, where v / ||v|| is the unit rotation axis and 2||v|| is the rotation angle. The transformation using the SE(3) transformation is then given by a similarity transformation relative to the position of the pivot point s:
[0065] x′=q(xs)q-1+s+t. (8)
[0066] Transform fields are encoded in MLP
[0067] W: (x, ω) i )→(v, s, t), (9)
[0068] It uses an architecture similar to that used by the template NeRF manager 150. The transformation of each state i is determined by the latent code ω. i The adjustment is used to represent this. The latent code is optimized through the embedding layer. An important property of log-quaternions is that exp(0) is an identity transformation. Therefore, the weights of the last layer of the MLP are derived from... Initialize the deformation to approximate this identity.
[0069] Along these lines, the deformation field MLP data 143 also includes SE(3) transform data 146. SE(3) transform data 146 represents the transform field encoded in the MLP and as described above. SE(3) transform data 146 includes rotation data 147 representing rotation q, pivot data 148 representing pivot point s, and translation data representing translation t. In some embodiments, rotation data 147, pivot data 148, and translation data 149 are represented in quarter-ary form. In some embodiments, rotation data 147, pivot data 148, and translation data 149 are represented in another form, for example, matrix form.
[0070] The template NeRF manager 150 is configured to generate a five-dimensional representation of a canonical frame F: (x, d) → (c, σ) that maps 3D position x = (x, y, z) and viewing direction d = (φ, θ) to RGB color c = (r, g, b) and density σ. In some implementations, an appearance code ψ is provided for each image. i It modulates the color output to handle appearance changes between input frames, such as exposure and white balance. For example... Figure 2 As shown, the canonical framework 240 is visualized. Points along ray 212 in the observation framework 220 have been mapped to points along curve 242 in the canonical framework 240 using the deformation field MLP 230. Each location and ray / camera viewpoint, along with appearance codes, is mapped to color and density via the NeRF MLP 250.
[0071] like Figure 1 As shown, the template NeRF manager 150 is configured to generate canonical frame position data 151, orientation data 152, potential appearance code data 153, and template NeRF MLP data 154 to output output data 160 including color data 162 and density data 163. The canonical frame position data 152 represents the results using a deformation field MLP (e.g., Figure 2 The MLP (230) is located within the canonical frame mapped from the observation frame. Orientation data 152 represents the ray or camera angle or direction cosine passing through each point in the canonical frame. Latent appearance code data 153 represents the latent appearance code used for each image.
[0072] Template NeRF MLP data 154 represents the values defining a NeRF MLP. The example NeRF MLP in the context of the technical solution described herein has six layers (one input, one output, and four hidden layers). In this example, the size of the hidden layers (i.e., the number of nodes) is 128, there are skip connections in the fourth layer, and a ReLU activation function. Deformed field MLP data 154 also includes color loss function data 155 and coarse-to-fine data 156.
[0073] The color loss function data 155 represents the value of the color loss function as defined below. At each optimization iteration, a batch of camera rays is randomly sampled from the entire pixel set in the dataset; then, hierarchical sampling is performed to query N from the coarse network. c One sample and N from the fine network c +N f Each sample is used to render the color of each ray from both sets of samples. The color loss is the total mean square error between the rendered pixel color and the true pixel color used for coarse and fine rendering.
[0074]
[0075] in, It is a set of rays in a batch, and C(r), and These are the ground truth, the coarse volume prediction RGB color, and the fine volume prediction RGB color, respectively.
[0076] The coarse-to-fine data 156 represents a coarse-to-fine variation regularization similar to the coarse-to-fine data 145. However, for NeRF MLP, sine and cosine are not weighted.
[0077] Returning to the elastic loss function, the deformation field T is allowed to behave freely in empty space because the subject moving relative to the background requires non-rigid deformation at some point in space. Therefore, in some implementations, the elastic loss function is weighted at each point by its contribution to the rendered view, as described below.
[0078] The five-dimensional neural radiation field represents the scene as a volumetric density and directional radiation emitted at any point in space. The color of any ray passing through the scene is rendered using principles from classical volumetric rendering. The volumetric density σ(x) can be interpreted as the differential probability of a ray terminating at location x at an infinitesimal particle. It has a near-boundary t n and distant boundary t f The expected color C(r) of the camera ray r(t) = o + td is
[0079]
[0080] in
[0081]
[0082] The function T(t) represents the distance from t along the ray. n The cumulative transmittance up to t, i.e., the transmittance of rays from t n The probability of reaching t without colliding with any other particles. This integral needs to be estimated from our continuous neural radiation field rendering view for each pixel tracked by the desired virtual camera.
[0083] This continuous integral is estimated numerically using a quadrature method. The deterministic orthogonality typically used for rendering discrete voxel meshes would effectively limit the resolution of our representation, as the MLP would only be queried from a fixed, discrete set of locations. Instead, we use a hierarchical sampling method, where we will [t] n , t f Divide the space into M evenly spaced segments, and then randomly and evenly extract one sample from each segment:
[0084]
[0085] Although discrete sample sets are used for estimating the integral, stratified sampling allows for the representation of continuous scenes because it enables the evaluation of the MLP at continuous locations during optimization. We use these samples to estimate C(r) as follows:
[0086]
[0087] in
[0088]
[0089] δ i =t i+1 -t i It is the distance between adjacent samples, and c i =c(r(t) i ), d).
[0090] A rendering strategy that densely evaluates the neural radiation field network at M query points along each camera ray is inefficient: free space and occluded regions that do not contribute to the rendered image are still repeatedly sampled. Therefore, instead of using only a single network to represent the scene, two networks are optimized simultaneously: a "coarse" and a "fine" network. First, N c The set of locations is sampled using hierarchical sampling, and a “coarse” network is evaluated at these locations as described in equations (13), (14), and (15). Given the output of this “coarse” network, more informative sampling is performed along each ray-generating point where the samples are skewed toward the relevant portion of the volume. For this purpose, the coarse network in equation (14) is first used for sampling. Rewrite the alpha composited color as all sampled colors along the ray c i Weighted sum:
[0091]
[0092] in
[0093] w i =T i (1-exp(-σ i δ i (17)
[0094] Output data 160 indicates NeRF MLP (i.e., Figure 2The output of the MLP 250 is shown above. The output data 160 includes color data 162 and density data 164. As shown above (e.g., equation (11)), the color data 162 depends on the position and ray angle, while the density data 164 depends only on the ray angle. Based on the color data 162 and density data 164, a view of the scene can be obtained from any perspective.
[0095] Figure 2 The system 200 shown can be optimized into units, namely, MLP 230 and MLP 250 can be identified using a single loss function, which is the sum of the above loss components.
[0096] L = L rgb +λL elastic-r +μL bg (18)
[0097] Where λ and μ are weights. In some implementations, λ = μ = 10. -3 .
[0098] Figure 3 This is a flowchart depicting an example method 300 for generating deformable NeRF. Method 300 can be derived by combining... Figure 1 The software construct described is executed by residing in the memory 126 of computer 120 and being run by a set of processing units 124, or it may be executed by a software construct residing in the memory of a computing device different from (e.g., remote from) computer 120.
[0099] At 310, the image acquisition manager 130 acquires image data representing multiple images (e.g., images 134(1), 134(2), ..., 134(N)), each of the multiple images including an image of a scene within a viewing frame (e.g., viewing frame position data 141), the scene including non-rigid deformable objects viewed from a corresponding viewpoint (e.g., viewpoints 136(1), 136(2), ..., 136(N));
[0100] At 320, the deformation model manager 140 generates a deformation model (e.g., deformation field MLP data 143) based on the image data. This deformation model describes the movement performed by a non-rigid deformation object during the generation of the image data. The deformation model is represented by a mapping between positions in the observation frame and positions in the canonical frame (e.g., canonical frame position data 151).
[0101] At 330, the template NeRF manager 150 generates a deformable neural radiation field (D-NeRF) based on the viewing direction of the projected rays through the position in the canonical frame. This D-NeRF provides a mapping between the viewing direction and position to the color (e.g., color data 162) and optical density (e.g., density data 164) at each position in the viewing frame, where the color and optical density at each position in the viewing frame enable viewing of non-rigid deformable objects from new perspectives.
[0102] The influence of the deformation field on the main body Figure 4A , 4B The diagram is shown in 4C. Figure 4A This is an example pose 400 of a human subject within a viewing frame. Figure 4B This is a diagram of an example pose 450 of a human subject within the illustrated specification framework. Figure 4A and 4B Of the two, the main subject is shown along with an illustration, which shows an orthographic projection view in the forward and leftward directions. Figure 4A Within the observation frame, note the right-to-left and front-to-back displacements between the observation model and the canonical model, which are modeled by the deformation field used for this observation.
[0103] Figure 4C This is a diagram illustrating an example interpolation 470 between keyframes (boxed) with and without elastic regularization (top) and with elastic regularization (bottom). Figure 4C The illustration shows the new views synthesized with and without elastic regularization, observed through linear interpolation of the deformed code. Without elastic regularization, the intermediate states exhibit distortion; for example, the distances between facial features change from the original image.
[0104] Figure 5 The illustrations show examples of a general-purpose computer device 500 and a general-purpose mobile computer device 550 that can be used with the techniques described herein.
[0105] like Figure 5 As shown, computing device 500 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 550 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are intended to be exemplary only and are not intended to limit the embodiments of the invention described and / or claimed in this document.
[0106] Computing device 500 includes a processor 502, a memory 504, a storage device 506, a high-speed interface 508 connected to the memory 504 and a high-speed expansion port 510, and a low-speed interface 512 connected to a low-speed bus 514 and the storage device 506. Each of components 502, 504, 506, 508, 510, and 512 is interconnected using various buses and may be mounted on a common motherboard or otherwise mounted where appropriate. Processor 502 can process instructions for execution within computing device 500, including instructions stored in memory 504 or on storage device 506 to display graphical information on an internal input / output device such as a GUI coupled to a display 516 connected to the high-speed interface 508. In other embodiments, multiple processors and / or multiple buses may be used, along with multiple memories and various types of memory where appropriate. Furthermore, multiple computing devices 500 may be connected, with each device providing a portion of the necessary operation (e.g., as a server library, a group of blade servers, or a multiprocessor system).
[0107] Memory 504 stores information within computing device 500. In one embodiment, memory 504 is one or more volatile memory cells. In another embodiment, memory 504 is one or more non-volatile memory cells. Memory 504 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.
[0108] Storage device 506 provides large-capacity storage for computing device 500. In one embodiment, storage device 506 may be or contain computer-readable media such as floppy disk devices, hard disk devices, optical disk devices or tape devices, flash memory or other similar solid-state storage devices, or arrays of devices, including devices in storage area networks or other configurations. A computer program product may be tangibly embodied in an information carrier. The computer program product may also contain instructions that, when executed, perform one or more methods such as those described above. The information carrier is a computer or machine-readable medium, such as memory 504, storage device 506, or memory on processor 502.
[0109] High-speed controller 508 manages bandwidth-intensive operations of computing device 500, while low-speed controller 512 manages less bandwidth-intensive operations. This functional allocation is merely exemplary. In one embodiment, high-speed controller 508 is coupled to memory 504, display 516 (e.g., via a graphics processor or accelerator), and high-speed expansion port 510, which can accept various expansion cards (not shown). In another embodiment, low-speed controller 512 is coupled to storage device 506 and low-speed expansion port 514. The low-speed expansion port, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, Wireless Ethernet), can be coupled, for example, via a network adapter to one or more input / output devices such as a keyboard, pointing device, scanner, or network device such as a switch or router.
[0110] Computing device 500 can be implemented in many different forms, as shown in the figure. For example, it can be implemented as a standard server 520 or multiple times in a group of such servers. It can also be implemented as part of a rack server system 524. Alternatively, it can be implemented in a personal computer such as a laptop computer 522. Alternatively, components from computing device 500 can be combined with other components in a mobile device (not shown), such as device 550. Each of such devices can contain one or more of computing devices 500, 550, and the entire system can consist of multiple computing devices 500, 550 communicating with each other.
[0111] Among other components, computing device 550 includes processor 552, memory 564, input / output devices such as display 554, communication interface 566, and transceiver 568. Device 550 may also provide additional storage via storage devices such as microdrives or other devices. Each component 550, 552, 564, 554, 566, and 568 is interconnected using various buses, and some components may be mounted on a common motherboard or otherwise suitably mounted.
[0112] Processor 552 can execute instructions within computing device 450, including instructions stored in memory 564. The processor can be implemented as a chipset comprising individual and multiple analog and digital processor chips. For example, the processor can provide coordination with other components of device 550, such as controlling the user interface, applications running on device 550, and wireless communications performed by device 550.
[0113] Processor 552 can communicate with the user via control interface 558 and display interface 556 coupled to display 554. Display 554 may be, for example, a TFT LCD (Thin Film Transistor Liquid Crystal Display) or OLED (Organic Light Emitting Diode) display, or other suitable display technologies. Display interface 556 may include suitable circuitry for driving display 554 to present graphics and other information to the user. Control interface 558 can receive commands from the user and translate them for submission to processor 552. Furthermore, an external interface 562 may be provided to communicate with processor 552, enabling device 550 to perform local area communication with other devices. For example, external interface 562 may provide wired communication in some embodiments, wireless communication in others, and multiple interfaces may be used.
[0114] Memory 564 stores information within computing device 550. Memory 564 can be implemented as one or more computer-readable media, one or more volatile memory cells, or one or more non-volatile memory cells. Extended memory 574 may also be provided and connected to device 550 via an extended interface 572, which may include a SIMM (Single In-line Memory Module) card interface. Such extended memory 574 can provide additional storage space for device 550, or it can store applications or other information for device 550. In particular, extended memory 574 may include instructions to perform or supplement the processes described above, and may also include security information. Thus, for example, extended memory 574 can be provided as a security module of device 550 and can be programmed using instructions that allow secure use of device 550. Furthermore, secure applications can be provided via a SIMM card and additional information, such as setting identification information on the SIMM card in an indestructible manner.
[0115] For example, as discussed below, the memory may include flash memory and / or NVRAM memory. In one embodiment, the computer program product is tangibly embodied in an information carrier. The computer program product may also contain instructions that, when executed, perform one or more methods as described above. The information carrier is a computer or machine-readable medium, such as memory 564, extended memory 574, or memory on processor 552, which may be received, for example, on transceiver 568 or external interface 562.
[0116] Device 550 can communicate wirelessly via communication interface 566, which includes digital signal processing circuitry when necessary. Communication interface 566 can provide communication under various modes or protocols such as GSM voice calls, SMS, EMS or MMS message sending, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, and others. For example, such communication can be performed via radio frequency transceiver 568. Furthermore, short-range communication can be performed using transceivers such as Bluetooth, WiFi, or others (not shown). Additionally, GPS (Global Positioning System) receiver module 570 can provide device 550 with additional navigation and location-related wireless data, which can be appropriately used by applications running on device 550.
[0117] Device 550 also uses audio codec 560 for audible communication, which can receive voice information from the user and convert it into usable digital information. Audio codec 560 can also generate audible sounds for the user, such as through a speaker, for example, in the earpiece of device 550. Such sounds can include sounds from voice phone calls, recorded sounds (e.g., voice messages, music files, etc.), and sounds generated by applications running on device 550.
[0118] As shown in the figure, the computing device 550 can be implemented in a variety of different forms. For example, it can be implemented as a cellular phone 550. It can also be implemented as part of a smartphone 582, a personal digital assistant, or other similar mobile devices.
[0119] Various implementations of the systems and techniques described herein can be implemented as digital electronic circuits, integrated circuits, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be dedicated or general-purpose and coupled to receive and transmit data and instructions from and to a storage device, at least one input device, and at least one output device.
[0120] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level programming and / or object-oriented programming languages, and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0121] To provide interaction with the user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including sound, voice, or tactile input.
[0122] The systems and technologies described herein can be implemented in computing systems that include back-end components (e.g., data servers), middleware components (e.g., application servers), front-end components (e.g., client computers having a graphical user interface or web browser through which users can interact with implementations of the systems and technologies described herein), or any combination of these back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), and the Internet.
[0123] A computing system may include clients and servers. Clients and servers are typically located far apart and interact via communication networks. The client-server relationship is established through computer programs running on the respective computers and having a client-server relationship with each other.
[0124] Back Figure 1In some embodiments, memory 126 can be any type of memory, such as random access memory, disk drive memory, flash memory, etc. In some embodiments, memory 126 can be implemented as more than one memory component (e.g., more than one RAM component or disk drive memory) associated with a component of compression computer 120. In some embodiments, memory 126 can be database memory. In some embodiments, memory 126 can be non-local memory, or may include non-local memory. For example, memory 126 can be or may include memory shared by multiple devices (not shown). In some embodiments, memory 126 can be associated with a server device (not shown) within a network and is configured to serve a component of compression computer 120.
[0125] Components of the compression computer 120 (e.g., modules, processing unit 124) may be configured to operate on one or more platforms (e.g., one or more similar or different platforms), which may include one or more types of hardware, software, firmware, operating systems, runtime libraries, etc. In some embodiments, components of the compression computer 120 may be configured to operate within a cluster of devices (e.g., a server cluster). In such embodiments, the functionality and processing of the components of the compression computer 120 may be distributed to several devices within the cluster.
[0126] The components of computer 120 can be or may include any type of hardware and / or software configured to process attributes. In some embodiments, Figure 1 One or more portions of the components shown in the computer 120 may be hardware-based modules (e.g., digital signal processors (DSPs), field-programmable gate arrays (FPGAs), memory), firmware modules, and / or software-based modules (e.g., computer code modules, a set of computer-readable instructions executable at a computer), or may include hardware-based modules (e.g., digital signal processors (DSPs), field-programmable gate arrays (FPGAs), memory), firmware modules, and / or software-based modules (e.g., computer code modules, a set of computer-readable instructions executable at a computer). For example, in some embodiments, one or more portions of the components of computer 120 may be or may include software modules configured to be executed by at least one processor (not shown). In some embodiments, the functionality of the components may be included in... Figure 1 The different modules and / or different components shown.
[0127] Although not shown, in some embodiments, components (or portions thereof) of computer 120 may be configured to operate within, for example, a data center (e.g., a cloud computing environment), a computer system, one or more server / host devices, etc. In some embodiments, components (or portions thereof) of computer 120 may be configured to operate within a network. Therefore, components (or portions thereof) of computer 120 may be configured to operate within various types of network environments that may include one or more devices and / or one or more server devices. For example, the network may be or may include a local area network (LAN), a wide area network (WAN), etc. The network may be or may include a wireless network and / or a wireless network implemented using, for example, gateway devices, bridges, switches, etc. The network may include one or more segments and / or may have portions based on various protocols, such as Internet Protocol (IP) and / or proprietary protocols. The network may include at least a portion of the Internet.
[0128] In some embodiments, one or more components of computer 120 may be or may include a processor configured to process instructions stored in memory. For example, depth image manager 130 (and / or a portion thereof), viewpoint manager 140 (and / or a portion thereof), raycasting manager 150 (and / or a portion thereof), SDV manager 160 (and / or a portion thereof), aggregation manager 170 (and / or a portion thereof), root lookup manager 180 (and / or a portion thereof), and depth image generation manager 190 (and / or a portion thereof) may be a combination of processor and memory configured to execute instructions relating to processes for implementing one or more functions.
[0129] Many implementations have been described. However, it will be understood that various modifications can be made without departing from the spirit and scope of this specification.
[0130] It will also be understood that when an element is referred to as being on, connected to, electrically connected to, coupled to, or electrically coupled to another element, it may be directly on, connected to, or coupled to the other element, or one or more intervening elements may be present. Conversely, when an element is referred to as being directly on, directly connected to, or directly coupled to another element, no intervening element is present. Although the terms "directly on," "directly connected to," or "directly coupled to" may not be used throughout the specific embodiments, elements shown as being directly on, directly connected to, or directly coupled may be referred to as such. The claims of this application may be modified to describe the exemplary relationships described in this specification or shown in the figures.
[0131] While certain features of the described embodiments have been illustrated herein, many modifications, substitutions, variations, and equivalents will now occur to those skilled in the art. Therefore, it should be understood that the appended claims are intended to cover all such modifications and variations that fall within the scope of the embodiments. It should be understood that they have been presented by way of example only and not limitation, and various changes in form and detail may be made. Any part of the apparatus and / or method described herein can be combined in any combination other than mutually exclusive combinations. The factual manner described herein may include various combinations and / or sub-combinations of the functions, components, and / or features of the different embodiments described.
[0132] Furthermore, the logical flows depicted in the figures do not require a specific order or sequence to achieve the desired result. Additionally, other steps may be provided, or steps may be eliminated from the described flow, and other components may be added to or removed from the described system. Therefore, other embodiments are within the scope of the appended claims.
Claims
1. A method comprising: Acquire image data representing a plurality of images, each of which includes an image of a scene within an observation frame, the scene comprising non-rigidly deformable objects observed from a corresponding viewpoint; A deformation model is generated based on the image data. The deformation model describes the movement performed by the non-rigid deformation object at the time of generating the image data. The deformation model is represented by a differentiable nonlinear mapping between the position in the observation frame and the position in the canonical frame. as well as A neural radiation field is generated based on the position and viewing direction of the rays projected through the location in the specified frame. The neural radiation field provides a mapping between the position and viewing direction to the color and optical density at each location in the viewing frame, which enables viewing of the non-rigid deformable object from a new perspective.
2. The method according to claim 1, wherein, The deformation model is tuned based on latent codes that encode the state of the scene within the framework.
3. The method according to claim 1, wherein, The deformation model includes: rotation, pivot point corresponding to the rotation, and translation.
4. The method according to claim 3, wherein, The rotation is encoded as a pure logarithmic quaternion.
5. The method according to claim 3, wherein, The deformation model includes the sum of: (i) a similarity transformation of the difference between the position and the pivot point, (ii) the pivot point, and (iii) the translation.
6. The method according to claim 1, wherein, The deformable model includes a multilayer perceptron (MLP) within a neural network.
7. The method according to claim 6, wherein, The elastic loss function component of the MLP is based on the norm of the matrix representing the deformation model.
8. The method according to claim 7, wherein, The matrix is the Jacobian of the deformable model relative to its position in the observation frame.
9. The method according to claim 7, wherein, The elastic loss function components are based on the singular value decomposition of the matrix representing the deformation model.
10. The method according to claim 9, wherein, The elastic loss function component is based on the logarithm of the singular value matrix generated by the singular value decomposition.
11. The method according to claim 7, wherein, The elastic loss function components are composed of rational functions to produce a robust elastic loss function.
12. The method according to claim 6, wherein, The background loss function component involves designating points in the scene as static points with a penalty for movement.
13. The method according to claim 12, wherein, The background loss function component is based on the difference between a static point and the mapping from the static point in the observation frame to the canonical frame according to the deformation model.
14. The method according to claim 6, wherein, Generating the deformed model includes: Position encoding is applied to position coordinates within the scene to generate a periodic function of position, the periodic function having a frequency that increases with training iterations of the MLP.
15. The method according to claim 14, wherein, The periodic function of the position encoding is multiplied by a weight indicating whether the training iteration includes weights of a specific frequency.
16. A computer program product including a non-transferable storage medium, the computer program product including code, the code causing the processing circuitry of a computing device to perform a method when executed, the method comprising: Acquire image data representing a plurality of images, each of which includes an image of a scene within an observation frame, the scene comprising non-rigidly deformable objects observed from a corresponding viewpoint; A deformation model is generated based on the image data. The deformation model describes the movement performed by the non-rigid deformation object at the time of generating the image data. The deformation model is represented by a differentiable nonlinear mapping between the position in the observation frame and the position in the canonical frame. as well as A neural radiation field is generated based on the position and viewing direction of the rays projected through the location in the specified frame. The neural radiation field provides a mapping between the position and viewing direction to the color and optical density at each location in the viewing frame, which enables viewing of the non-rigid deformable object from a new perspective.
17. The computer program product according to claim 16, wherein, The deformable model includes a multilayer perceptron (MLP) within a neural network.
18. The computer program product according to claim 17, wherein, The elastic loss function component of the MLP is based on the norm of the matrix representing the deformation model.
19. The computer program product according to claim 18, wherein, The matrix is the Jacobian of the deformable model relative to its position in the observation frame.
20. An electronic device, the electronic device comprising: Memory; as well as A control circuit, coupled to the memory, is configured to: Acquire image data representing a plurality of images, each of which includes an image of a scene within an observation frame, the scene comprising non-rigidly deformable objects observed from a corresponding viewpoint; A deformation model is generated based on the image data. The deformation model describes the movement performed by the non-rigid deformation object at the time of generating the image data. The deformation model is represented by a differentiable nonlinear mapping between the position in the observation frame and the position in the canonical frame. as well as A neural radiation field is generated based on the position and viewing direction of the rays projected through the location in the specified frame. The neural radiation field provides a mapping between the position and viewing direction to the color and optical density at each location in the viewing frame, which enables viewing of the non-rigid deformable object from a new perspective.