Dynamic 3D Human Rendering and Synthesis Method Based on Intrinsic Coordinate Hash Encoding
Through hash encoding and implicit neural representation model based on intrinsic coordinates, the new perspective images of dynamic three-dimensional human bodies are quickly reconstructed and synthesized, which solves the problems of long computing time and device dependence in the existing technology, and realizes efficient dynamic three-dimensional human body rendering.
Patent Information
- Application Number
- CN202310084613.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-17
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2043-01-17
AI Technical Summary
The existing new dynamic three-dimensional human perspective synthesis method has a long calculation time and requires intensive image input or depth cameras, which are difficult to deploy and apply. The existing hash encoding based on external coordinates is only suitable for static scenes and cannot be extended to dynamic scenes.
The hash encoding and implicit neural representation model based on intrinsic coordinates are adopted. By optimizing the human body's pose and morphological parameters, the high-dimensional feature vector of query points is calculated using a multi-level hash encoder and perceptron. The model is self-supervised to quickly reconstruct the human body's geometry and texture, which is suitable for dynamic scenarios.
It realizes the training process within 20 minutes, can synthesize high-quality images from any perspective and human body shape, supports animation production and other applications, avoiding long-term training and equipment dependence.
Smart Images

Figure CN116109757B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a new perspective synthesis technology for dynamic three-dimensional human bodies, and in particular to a dynamic three-dimensional human body rendering synthesis method based on hash coding of intrinsic coordinates, a training method for an implicit neural representation model, an electronic device, and a storage medium. Background Art
[0002] Synthesizing new views of dynamic 3D human figures is an important research direction in computer vision. It has widespread applications in many fields, such as sports broadcasting, video conferencing, and VR / AR. Although this area has been studied for a long time, existing methods still require considerable computational time, making the technology difficult for general users to adopt.
[0003] Traditional novel view synthesis methods require either dense 2D image input for interpolation or high-fidelity 3D reconstruction from depth cameras to produce realistic results. Some model-based methods can reconstruct explicit 3D meshes from sparse-view videos, but these 3D meshes often lack geometric detail, resulting in sub-realistic rendered images. Some methods apply neural radiance fields to synthesize novel view images of dynamic human figures. By combining human body priors with neural radiance fields, these methods can reconstruct a rough set of human figures from sparse-view videos containing human figures. However, due to the expensive computational cost of neural radiance fields, these methods often require long training times to fit each input object. Furthermore, most methods still require calibrated multi-view camera systems, making them difficult to deploy and apply. In recent years, through carefully designed multi-resolution hash codes, the training speed of neural radiance fields has increased by several orders of magnitude. However, this strategy is currently based on extrinsic coordinates and is only applicable to static scenes, failing to scale to dynamic ones. Summary of the Invention
[0004] In view of the above problems, the present invention provides a dynamic three-dimensional human body rendering synthesis method based on hash coding of intrinsic coordinates, a training method of an implicit neural representation model, an electronic device and a storage medium in order to solve at least one of the above problems.
[0005] According to a first aspect of the present invention, a dynamic three-dimensional human body rendering and synthesis method based on hash coding of intrinsic coordinates is provided, comprising:
[0006] By optimizing the human body posture parameters and morphological parameters corresponding to the video frames in the motion video of the target dressed person, a parametric human body mesh of the target dressed person is obtained, and the parametric human body mesh is used as a rough explicit geometric proxy of the target dressed person.
[0007] Based on the camera parameters, the query point corresponding to the sampling ray of the pixel point in the space where the rough explicit geometric proxy of the target clothed person is located is calculated;
[0008] According to the preset mapping rules and the geometric information of the rough geometric proxy of the target clothed person, the intrinsic coordinates of the query point mapped to the radiation density cube grid are calculated;
[0009] The multi-layer first perceptron of the trained implicit neural representation model is used to predict the offset field of the radiation density cube grid, and the offset field of the radiation density cube grid is used to optimize the intrinsic coordinates of the query point, wherein the implicit neural representation model is used to represent the human body model of the target clothed person;
[0010] Utilize the multi-level hash encoder of the trained implicit neural representation model to calculate the high-dimensional feature vector of the intrinsic coordinates of the optimized query point;
[0011] The high-dimensional feature vector of the query point is processed using the multi-layer second perceptron of the trained implicit neural representation model to obtain the density and color of the query point;
[0012] The volume rendering formula is used to calculate the color of the pixel corresponding to the query point, and the video frame image of the target clothed person is obtained. The motion video of the target clothed person is synthesized based on the video frame image.
[0013] According to an embodiment of the present invention, the above-mentioned step of calculating the query point corresponding to the sampling ray of the pixel point in the space where the rough explicit geometric proxy of the target clothed person is located based on the camera parameters includes:
[0014] According to the camera optical center and light direction in the camera parameters, the sampling light of the pixel point is calculated;
[0015] According to the preset sampling depth, uniform sampling is performed on the sampling light to obtain the query point.
[0016] According to an embodiment of the present invention, the above-mentioned preset mapping rule is expressed by formula (1):
[0017] UVD(x|T t )=(UV(p|T t ),S(d)) (1),
[0018] Where x represents the query point, T t represents the rough explicit geometry proxy of the t-th frame, and d represents the query point x to T t The signed distance of the query point x in T t The nearest point on UV(p|T t ) represents the rough explicit geometry proxy T tThe corresponding texture coordinates in the texture expansion map, S(*) represents the Sigmoid function, and UVD(*) represents the mapping of the query point to the intrinsic coordinates in the radiation density cube grid.
[0019] According to an embodiment of the present invention, the above-mentioned method of calculating the high-dimensional feature vector of the intrinsic coordinates of the optimized query point using the multi-level hash encoder of the trained implicit neural representation model includes:
[0020] Divide the radiation density cube grid into multiple voxel grids with different resolutions from coarse to fine, where the resolution of the voxel grid is determined by the resolution of the motion video of the target clothed person;
[0021] According to a preset query formula and preset prime values, the multi-level hash encoder of the trained implicit neural representation model is used to calculate the feature vectors of the vertices of the voxel grid of the specific resolution, wherein the intrinsic coordinates of the optimized query point are located in the voxel grid of the specific resolution;
[0022] According to the coordinates of the vertices of the voxel grid of a specific resolution and the intrinsic coordinates of the optimized query point, the eigenvectors of the vertices are interpolated to obtain the eigenvectors of the intrinsic coordinates of the optimized query point in the voxel grid of the specific resolution;
[0023] Repeat the vertex feature vector calculation and interpolation operations to obtain the feature vectors of the optimized query point's intrinsic coordinates in voxel grids of different resolutions;
[0024] The feature vectors of the intrinsic coordinates of the optimized query point in voxel grids of different resolutions are vector-concatenated to obtain a high-dimensional feature vector of the intrinsic coordinates of the optimized query point.
[0025] According to an embodiment of the present invention, the above preset query formula is expressed by formula (2):
[0026]
[0027] Where z represents the coordinate of the vertex in the voxel grid, π i is a preset prime number, T represents the size of the hash table of the multi-level hash encoder of the trained implicit neural representation model, and ⊕ represents the exclusive-or operation.
[0028] According to an embodiment of the present invention, the multi-level first perceptron of the trained implicit neural representation model is expressed by formula (3):
[0029] Δr=F φ (r,e t ) (3),
[0030] Where r represents a point in the radiation density cube grid, Δr represents the offset corresponding to r, and e t is the conditional variable of the rough explicit geometry proxy at frame t, F Φ Multi-layer first perceptron representing the trained implicit neural representation model;
[0031] Among them, the multi-level second perceptron of the trained implicit neural representation model is expressed by formula (4):
[0032]
[0033] Among them, r is the intrinsic coordinate corresponding to the query point x, Δr represents the offset corresponding to r, σ t (x) represents the density of query point x in the tth frame, c t (x) represents the color of the query point x in the tth frame, represents the hash table of the multi-level hash encoder of the trained implicit neural representation model, h(*) represents the mapping of the optimized intrinsic coordinates of the query point to its high-dimensional feature vector, F ω Multi-layer second-order perceptron representing the trained implicit neural representation model.
[0034] According to an embodiment of the present invention, the above volume rendering formula is expressed by formulas (5) and (6):
[0035]
[0036] α(x i )=1-exp(-σ(x i )δ i ) (6),
[0037] Among them, C(γ) represents the color of light γ, c(x i ) and σ(x i ) represent the sampling points x i The color and density values, δ i represents the spacing between sampling points, α(*) represents the sampling point x i opacity.
[0038] According to a second aspect of the present invention, a method for training an implicit neural representation model is provided, comprising:
[0039] Extracting a real video frame image of the target dressed person based on the motion video of the target dressed person;
[0040] Obtaining a synthesized video frame image of a target clothed person using an implicit neural representation model, wherein the implicit neural representation model includes a multi-level first perceptron, a multi-level second perceptron, and a multi-level hash encoder;
[0041] Using a loss function to process the synthesized video frame images and the real video frame images, and optimizing the implicit neural representation model according to the loss value, wherein the loss function includes a photometric loss function and a regularization loss function;
[0042] The real video frame image and synthesized video frame image acquisition operations and model optimization operations are iteratively performed until the preset conditions are met to obtain a trained implicit neural representation model, wherein the trained implicit neural representation model is applied to a dynamic three-dimensional human body rendering synthesis method based on hash coding of intrinsic coordinates.
[0043] According to a third aspect of the present invention, there is provided an electronic device, comprising:
[0044] one or more processors;
[0045] a storage device for storing one or more programs,
[0046] When one or more programs are executed by one or more processors, the one or more processors execute a dynamic three-dimensional human body rendering synthesis method based on hash coding of intrinsic coordinates and a training method of an implicit neural representation model.
[0047] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which executable instructions are stored. When the instructions are executed by a processor, the processor executes a dynamic three-dimensional human body rendering and synthesis method based on hash coding of intrinsic coordinates and a training method for an implicit neural representation model.
[0048] The dynamic 3D human body rendering synthesis method based on hash coding of intrinsic coordinates provided by the present invention uses a motion video of a target clothed person, a hash encoder based on intrinsic coordinate representation, and neural rendering to self-supervise the reconstruction of the geometry and high-quality texture of the target clothed person in the video. This method can then synthesize rendered images from any perspective and body shape. Compared to various existing solutions, this method eliminates the need for training that can take up to ten hours and can complete the training process in just twenty minutes. Furthermore, the generated body shape can be edited using a rough geometric proxy, making it convenient for downstream applications such as animation production. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a flow chart of a dynamic three-dimensional human body rendering and synthesis method based on hash coding of intrinsic coordinates according to an embodiment of the present invention;
[0050] Figure 2 is a flow chart of obtaining a high-dimensional feature vector of the intrinsic coordinates of a query point according to an embodiment of the present invention;
[0051] Figure 3is a flowchart of a new perspective rendering and synthesis method for dynamic three-dimensional human motion video according to another embodiment of the present invention;
[0052] Figure 4 is a schematic diagram illustrating intrinsic coordinate representation according to another embodiment of the present invention;
[0053] Figure 5 is a schematic diagram illustrating an offset field according to another embodiment of the present invention;
[0054] Figure 6 is a flowchart of a method for training an implicit neural representation model according to an embodiment of the present invention;
[0055] Figure 7 The block diagram of an electronic device suitable for implementing a dynamic three-dimensional human body rendering and synthesis method based on hash coding of intrinsic coordinates and a training method of an implicit neural representation model according to an embodiment of the present invention is schematically shown. DETAILED DESCRIPTION
[0056] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0057] In the field of new perspective synthesis of dynamic three-dimensional human bodies, although the method based on implicit neural representation can show highly realistic results, it requires twelve hours or even longer training time for each object, which makes it difficult to deploy in practical applications. Although multi-level hash coding can accelerate the training of neural radiation fields, this strategy is currently only applicable to static scenes. To this end, an embodiment of the present invention provides a fast new perspective synthesis method for the human body based on hash coding and neural rendering of intrinsic coordinate representation, which is used to solve various technical problems in the prior art.
[0058] The purpose of the present invention is to provide a new perspective synthesis algorithm for human motion videos based on single-perspective or sparse perspectives, which can quickly reconstruct the approximate geometry and texture of the human body from the input video, and then perform synthetic rendering of the new perspective and new human form.
[0059] In the technical solution of the present invention, the acquisition, storage and application of the motion video and video frame images of the target dressed person have all been agreed by the target dressed person, comply with the provisions of relevant laws and regulations and public order and good morals, and necessary confidentiality measures have been taken.
[0060] The technical solution of the present invention utilizes an intrinsic coordinate representation based on the above-mentioned explicit human body geometric proxy to map corresponding points in different frames to the same intrinsic coordinates, and the value range of this mapping is called the UV-D grid (radiance density cube grid); the human body is represented as a rough explicit geometric proxy and a UV-D grid defined on it, and density and color are recorded in the UV-D grid; for any coordinate in the UV-D grid, multi-level hash coding is used to map it to a high-dimensional feature vector to achieve the effect of accelerated training; the above-mentioned high-dimensional feature vector is input into the neural radiation field network, and the density and color values recorded at that point in the UV-D grid are output; based on the above strategy, the density and color values of any point in space in any frame can be obtained, and then the light is sampled for each pixel according to the input camera pose parameters, and the color of each pixel is obtained using the volume rendering formula. By calculating the l1 norm of the difference between the generated color and the true color as the optimization target, the model can be self-supervised trained.
[0061] The technical solution provided by the present invention is described and explained in detail below through specific implementation methods in conjunction with the accompanying drawings.
[0062] Figure 1 The present invention is a flowchart of a dynamic three-dimensional human body rendering and synthesis method based on hash coding of intrinsic coordinates according to an embodiment of the present invention.
[0063] like Figure 1 As shown, the above-mentioned dynamic three-dimensional human body rendering and synthesis method based on hash coding of intrinsic coordinates includes operations S110 to S170.
[0064] In operation S110 , a human body parameterized mesh of the target clothed person is obtained by optimizing posture parameters and morphological parameters of the human body corresponding to video frames in a motion video of the target clothed person, and the human body parameterized mesh is used as a rough explicit geometric proxy of the target clothed person.
[0065] With the consent of the target dressed person, the present invention obtains the motion video of the target dressed person; and in various embodiments of the present invention, the processing of the video and the video frame images are all permitted by the target dressed person.
[0066] The above-mentioned method of calculating the query point corresponding to the sampling ray of the pixel point in the space where the rough explicit geometric proxy of the target clothed person is located based on the camera parameters includes: calculating the sampling ray of the pixel point based on the camera optical center and the ray direction in the camera parameters; and uniformly sampling on the sampling ray according to a preset sampling depth to obtain the query point.
[0067] The above-mentioned mapping rule is expressed by formula (1):
[0068] UVD(x|T t)=(UV(p|T t ),S(d)) (1),
[0069] Where x represents the query point, T t represents the rough explicit geometry proxy of the t-th frame, and d represents the query point x to T t The signed distance of the query point x in T t The nearest point on UV(p|T t ) represents the rough explicit geometry proxy T t The corresponding texture coordinates in the texture expansion map of , S(*) represents the Sigmoid function, and UVD(*) represents the mapping of the query point to the intrinsic coordinates in the radiation density cube grid.
[0070] In operation S120 , a query point corresponding to a sampling ray of a pixel point in a space where a rough explicit geometric proxy of the target clothed person is located is calculated based on the camera parameters.
[0071] In operation S130 , the intrinsic coordinates of the query point mapped onto the radiance density cube grid are calculated according to a preset mapping rule and geometric information of the rough geometric proxy of the target clothed person.
[0072] The above-mentioned radiation density cube grid is the UV-D grid, which records the density value and radiation value of the space near the surface of the rough geometric proxy grid through the multi-level second perceptron of the hash coding and implicit neural representation model.
[0073] In operation S140, a multi-level first perceptron of a trained implicit neural representation model is used to predict an offset field of a radiation density cube grid, and the offset field of the radiation density cube grid is used to optimize the intrinsic coordinates of the query point, wherein the implicit neural representation model is used to represent a human body model of a target clothed person.
[0074] In operation S150 , a high-dimensional feature vector of the optimized intrinsic coordinates of the query point is calculated using the multi-level hash encoder of the trained implicit neural representation model.
[0075] Figure 2 4 is a flowchart of obtaining a high-dimensional feature vector of the intrinsic coordinates of a query point according to an embodiment of the present invention.
[0076] like Figure 2 As shown, the above-mentioned calculation of the high-dimensional feature vector of the intrinsic coordinates of the optimized query point using the multi-level hash encoder of the trained implicit neural representation model includes operations S210 to S250.
[0077] In operation S210 , the radiation density cube grid is divided into a plurality of voxel grids with different resolutions from coarse to fine, wherein the resolution of the voxel grid is determined by the resolution of the motion video of the target clothed person.
[0078] A voxel grid is a cube grid made up of multiple small cubes, each of which has 8 vertices.
[0079] In operation S220, according to a preset query formula and a preset prime value, a multi-level hash encoder of a trained implicit neural representation model is used to calculate feature vectors of vertices of a voxel grid of a specific resolution, wherein the intrinsic coordinates of the optimized query point are located in the voxel grid of the specific resolution.
[0080] The above preset query formula is expressed by formula (2):
[0081]
[0082] Where z represents the coordinate of the vertex in the voxel grid, π i is a preset prime number, T represents the size of the hash table of the multi-level hash encoder of the trained implicit neural representation model, and ⊕ represents the XOR operation.
[0083] The eigenvector of each vertex of the voxel grid at this resolution is calculated.
[0084] In operation S230 , interpolation calculation is performed on the feature vectors of the vertices of the voxel grid of the specific resolution and the optimized intrinsic coordinates of the query point to obtain the feature vectors of the optimized intrinsic coordinates of the query point in the voxel grid of the specific resolution.
[0085] In operation S240 , the vertex feature vector calculation operation and the interpolation calculation operation are repeated to obtain feature vectors of the optimized intrinsic coordinates of the query point in voxel grids of different resolutions.
[0086] In operation S250 , feature vectors of the optimized intrinsic coordinates of the query point in voxel grids of different resolutions are concatenated to obtain a high-dimensional feature vector of the optimized intrinsic coordinates of the query point.
[0087] The UV-D grid (i.e., the radiation density cube grid) is divided into multiple voxel grids from coarse to fine. The resolution of the voxel grid is determined by the resolution of the input video. In this way, multiple voxel grids with different resolutions are obtained, wherein each voxel grid includes multiple small cubes, each small cube has 8 vertices, and the intrinsic coordinates of the query point are located in one of the voxel grids. For the subdivided voxel grids of a specific resolution, the feature vector corresponding to each vertex (or corner point) is queried in the hash table according to the preset query formula shown in formula (2). According to the position of the intrinsic coordinates in the voxel grid of the resolution, the feature vector at the resolution level is interpolated and calculated. The feature vectors at each level are combined to obtain the final feature vector.
[0088] First, the UV-D grid is divided into multiple voxel grids with different resolutions. For example, the UV-D grid is divided into voxel grids with 16 different resolutions. Then, the hash encoder of the trained implicit neural representation model is used to calculate the feature vectors of the 8 vertices of the voxel grid where the intrinsic coordinates of the query point are located. The feature vectors of these 8 vertices are then interpolated to obtain the 2-dimensional feature vector of the intrinsic coordinates of the query point at that resolution. Finally, the 2-dimensional feature vectors of the intrinsic coordinates of the query point at all resolutions are obtained, and these 2-dimensional feature vectors are concatenated to obtain the final high-dimensional feature vector of the intrinsic coordinates of the query point. For example, a 2-dimensional feature vector is obtained for each resolution, and a total of 16 2-dimensional feature vectors are obtained for 16 resolutions. These feature vectors are concatenated to obtain a 32-dimensional feature vector, which is used as the high-dimensional feature vector of the intrinsic coordinates of the query point.
[0089] In operation S160 , the high-dimensional feature vector of the query point is processed using the multi-level second perceptron of the trained implicit neural representation model to obtain the density and color of the query point.
[0090] According to an embodiment of the present invention, the multi-level first perceptron of the trained implicit neural representation model is expressed by formula (3):
[0091] Δr=f φ (r,e t ) (3),
[0092] Where r represents a point in the radiation density cube grid, Δr represents the offset corresponding to r, and e t is the conditional variable of the rough explicit geometry proxy at frame t, F Φ Multi-layer first perceptron representing the trained implicit neural representation model.
[0093] The multi-level second perceptron of the trained implicit neural representation model is expressed by formula (4):
[0094]
[0095] Among them, r is the intrinsic coordinate corresponding to the query point x, Δr represents the offset corresponding to r, σ t (x) represents the density of query point x in the tth frame, c t (x) represents the color of the query point x in the tth frame, represents the hash table of the multi-level hash encoder of the trained implicit neural representation model, h(*) represents the mapping of the optimized intrinsic coordinates of the query point to its high-dimensional feature vector, F ω Multi-layer second-order perceptron representing the trained implicit neural representation model.
[0096] In operation S170 , the color of the pixel corresponding to the query point is calculated using a volume rendering formula to obtain a video frame image of the target clothed person, and a motion video of the target clothed person is synthesized based on the video frame image.
[0097] The acquisition and synthesis processing of the above-mentioned video frame images have been authorized by the target person wearing the clothes, and the processing thereof has also been implemented within the scope of the target person wearing the clothes' permission.
[0098] According to an embodiment of the present invention, the above volume rendering formula is expressed by formulas (5) and (6):
[0099]
[0100] α(x i )=1-exp(-σ(x i )δ i ) (6),
[0101] Among them, C(γ) represents the color of light γ, c(x i ) and σ(x i ) represent the sampling points x i The color and density values, δ i represents the spacing between sampling points, α(*) represents the sampling point x i opacity.
[0102] The dynamic 3D human body rendering synthesis method based on hash coding of intrinsic coordinates provided by the present invention uses a motion video of a target clothed person, a hash encoder based on intrinsic coordinate representation, and neural rendering to self-supervise the reconstruction of the geometry and high-quality texture of the target clothed person in the video. This method can then synthesize rendered images from any perspective and body shape. Compared to various existing solutions, this method eliminates the need for training that can take up to ten hours and can complete the training process in just twenty minutes. Furthermore, the generated body shape can be edited using a rough geometric proxy, making it convenient for downstream applications such as animation production.
[0103] Next, another specific embodiment of the present invention and the combination Figures 3-5 The above technical solution provided by the present invention is further described in detail.
[0104] Figure 3 The present invention is a flowchart of a new perspective rendering and synthesis method for dynamic three-dimensional human motion video according to another embodiment of the present invention.
[0105] Figure 4 FIG. 4 is a schematic diagram illustrating intrinsic coordinate representation according to another embodiment of the present invention.
[0106] Figure 5 FIG. 4 is a schematic diagram illustrating an offset field according to another embodiment of the present invention.
[0107] like Figure 3 As shown in FIG, the new perspective rendering synthesis method for dynamic 3D human motion video can be divided into 7 steps. In step 1, a human body parameterized mesh is extracted from the motion video of the target dressed person (the video length can be ten to twenty seconds), that is, the human body parameterized mesh of each video frame is extracted as the rough geometric proxy of each frame. In the process of obtaining the human body parameterized mesh, it is optional to use the SMPL human body parameter mesh as the rough geometric proxy and obtain the rough geometric proxy T of each video frame by optimizing the posture parameters and morphological parameters of the human body in each video frame. i , based on the fact that these explicit grids share the same texture expansion map, formula (1) is used to map any point x in the t-frame space to the UV-D grid to ensure that the relevant points in different frames are mapped to the same intrinsic coordinates.
[0108] Step 2: Use the input camera parameters to calculate the query point in space. When rendering at any perspective, the volume rendering strategy is applied: sample a ray for each pixel, then sample several points on each ray, and calculate the color of the current pixel based on the density and color values of these sampled points. This is shown in formulas (7) and (8):
[0109]
[0110] x i =o+t i V (8),
[0111] Among them, γ is the sampled light, o is the optical center of the camera, V is the direction of the light, and xi is the sampling point on the light. and t i is the sampling depth.
[0112] Step 3: Use the geometric information of the rough geometry proxy to calculate the intrinsic coordinates of the query point in space. Figure 4 As shown, for any point x in the t-th frame space, the rough geometric proxy of the t-th frame is used to map it to the intrinsic coordinates in the UV-D grid. In this example, the query point x is mapped to the intrinsic coordinates in the UV-D grid using the geometric proxy T of the t-th frame. t The texture coordinates of the nearest point p on and the corresponding normalized signed distance d are expressed as intrinsic coordinates, as shown in formula (1):
[0113] UVD(x|T t )=(UV(p|T t ),S(d)) (1).
[0114] Step 4: Use a multi-layer perceptron to predict the offset field in the UV-D grid. Figure 5 As shown in Figure 3, the rough geometry proxy is not accurate enough to model details such as clothing wrinkles. In this example, the offset field shown in formula (3) is used to optimize the intrinsic coordinate mapping.
[0115] Δr=F φ (r,e t ) (3).
[0116] By using formula (3), we can find the intrinsic coordinates r+Δr of any query point in any frame space.
[0117] In step 5, the optimized intrinsic coordinates are input into a multi-level hash encoder to calculate the high-dimensional feature vector. The UV-D grid is divided into multiple voxel grids from coarse to fine. The resolution of the voxel grid is determined by the resolution of the input video. For a specific resolution of the subdivided voxel grid, the following query formula is used to query each corner point (i.e., the 8 vertices of the voxel grid) in the hash table. The corresponding feature vector is shown in formula (2):
[0118]
[0119] Then, we interpolate the feature vectors at that resolution level based on the position of the intrinsic coordinates in the voxel grid at that resolution. The feature vectors at each level are combined to get the final feature vector.
[0120] Step 6: Use a multi-layer perceptron to calculate the density and color at the query point. Apply the above steps to each query point, then input the obtained feature vector into the multi-layer perceptron to calculate the density and color value of the query point, as shown in formula (4):
[0121]
[0122] In step S7, the final composite image is calculated using the volume rendering formula. Based on the input camera pose parameters and the rough human proxy geometry, each pixel can be sampled with light, and then the color of each pixel can be obtained using the volume rendering formula, thereby obtaining a new perspective composite image under the specified camera pose parameters and human body shape and pose. For the sampling points of each light in step 2, steps 3-6 can be applied to calculate the density and color at these points, and then the color of each light can be calculated using formulas (5) and (6):
[0123]
[0124] α(x i )=1-exp(-σ(x i )δ i ) (6).
[0125] Figure 6 4 is a flowchart of a method for training an implicit neural representation model according to an embodiment of the present invention.
[0126] like Figure 6 As shown, the above-mentioned training method of an implicit neural representation model includes operations S610 to S640.
[0127] In operation S610 , a real video frame image of a target dressed person is extracted based on a motion video of the target dressed person.
[0128] In operation S620, a synthesized video frame image of a target clothed person is obtained using an implicit neural representation model, wherein the implicit neural representation model includes a multi-level first perceptron, a multi-level second perceptron, and a multi-level hash encoder.
[0129] In operation S630, the synthesized video frame image and the real video frame image are processed using a loss function, and the implicit neural representation model is optimized according to the loss value, wherein the loss function includes a photometric loss function and a regularization loss function.
[0130] The regularized loss function in the above loss function can be expressed by formula (9):
[0131]
[0132] in, Represents the ReLu function, β represents the hyperparameter, χ represents the sampling point set at each training, σ(x) represents the density of the sampling point set, and d(x) represents the signed distance from the sampling point set to the rough display geometry agent.
[0133] The photometric loss function in the above loss function can be expressed by formula (10):
[0134]
[0135] in, Represents the set of light samples sampled during each training, and C(r) represents the color of the pixel corresponding to the light r in the real picture.
[0136] The photometric loss and regularization loss are optimized to obtain the final model; the above loss functions can converge to the rough proxy geometry of the input more quickly.
[0137] In operation S640, the real video frame image and synthesized video frame image acquisition operation and the model optimization operation are iteratively performed until the preset conditions are reached to obtain a trained implicit neural representation model, wherein the trained implicit neural representation model is applied to a dynamic three-dimensional human body rendering synthesis method based on hash coding of intrinsic coordinates.
[0138] By comparing the synthesized images with the real input images, the entire model can be trained in a self-supervised manner.
[0139] Figure 7 The block diagram of an electronic device suitable for implementing a dynamic three-dimensional human body rendering and synthesis method based on hash coding of intrinsic coordinates and a training method of an implicit neural representation model according to an embodiment of the present invention is schematically shown.
[0140] like Figure 7 As shown, the electronic device 700 according to an embodiment of the present invention includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage portion 708 into a random access memory (RAM) 703. The processor 701 may, for example, include a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include an onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0141] Various programs and data required for the operation of the electronic device 700 are stored in the RAM 703. The processor 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The processor 701 performs various operations according to the method flow of the embodiment of the present invention by executing the programs in the ROM 702 and / or RAM 703. It should be noted that the programs may also be stored in one or more memories other than the ROM 702 and RAM 703. The processor 701 may also perform various operations according to the method flow of the embodiment of the present invention by executing the programs stored in the one or more memories.
[0142] According to an embodiment of the present invention, electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to bus 704. Electronic device 700 may further include one or more of the following components connected to I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 708 including a hard disk; and a communication section 709 including a network interface card such as a LAN card or a modem. Communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to I / O interface 705 as needed. Removable media 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed in drive 710 as needed, so that computer programs read from the removable media can be installed into storage section 708 as needed.
[0143] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.
[0144] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, such as, but not limited to, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, a computer-readable storage medium may include the ROM 702 and / or RAM 703 described above and / or one or more memories other than ROM 702 and RAM 703.
[0145] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0146] Those skilled in the art will appreciate that various combinations and / or combinations of features described in the various embodiments and / or claims of the present invention may be made, even if such combinations and / or combinations are not explicitly described in the present invention. In particular, various combinations and / or combinations of features described in the various embodiments and / or claims of the present invention may be made, without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.
[0147] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A dynamic 3D human body rendering synthesis method based on hash coding of intrinsic coordinates, comprising: Obtaining a human body parameterized mesh of the target dressed person by optimizing the human body posture parameters and morphological parameters corresponding to video frames in a motion video of the target dressed person, and using the human body parameterized mesh as a rough explicit geometric proxy of the target dressed person; Calculating, based on camera parameters, a query point in the space where a rough explicit geometric proxy of the target clothed person is located that corresponds to a sampling ray of a pixel point; Calculating the intrinsic coordinates of the query point mapped onto the radiation density cube grid according to a preset mapping rule and geometric information of the rough geometric proxy of the target clothed person; Predicting an offset field of the radiation density cube grid using a multi-layer first perceptron of a trained implicit neural representation model, and optimizing the intrinsic coordinates of the query point using the offset field of the radiation density cube grid, wherein the implicit neural representation model is used to represent a human body model of the target clothed person; Calculating an optimized high-dimensional feature vector of the intrinsic coordinates of the query point using a multi-level hash encoder of the trained implicit neural representation model; Processing the high-dimensional feature vector of the query point using the multi-layer second perceptron of the trained implicit neural representation model to obtain the density and color of the query point; The color of the pixel corresponding to the query point is calculated using a volume rendering formula to obtain a video frame image of the target dressed person, and a motion video of the target dressed person is synthesized based on the video frame image.
2. The method according to claim 1, wherein Calculating the query point corresponding to the sampling ray of the pixel point in the space where the rough explicit geometric proxy of the target clothed person is located according to the camera parameters includes: Calculate the sampling light of the pixel point according to the camera optical center and the light direction in the camera parameters; According to a preset sampling depth, uniform sampling is performed on the sampling light to obtain the query point.
3. The method according to claim 1, wherein The preset mapping rule is expressed by formula (1): UVD(x|Tt)=(UV(p|Tt),S(d)) (1), Wherein, x represents the query point, Tt represents the coarse explicit geometry proxy of the tth frame, d represents the signed distance from the query point x to Tt, p represents the nearest point of the query point x on Tt, UV(p|Tt) represents the texture coordinate corresponding to p in the texture unfolding map of the coarse explicit geometry proxy Tt, S(*) represents the Sigmoid function, and UVD(*) represents the mapping of the query point to the intrinsic coordinates in the radiation density cube mesh.
4. The method according to claim 1, wherein The method of calculating the optimized high-dimensional feature vector of the intrinsic coordinates of the query point using the multi-level hash encoder of the trained implicit neural representation model includes: Dividing the radiation density cube grid into a plurality of voxel grids with different resolutions from coarse to fine, wherein the resolution of the voxel grid is determined by the resolution of the motion video of the target dressed person; Calculating feature vectors of vertices of a voxel grid of a specific resolution using a multi-level hash encoder of the trained implicit neural representation model according to a preset query formula and preset prime values, wherein the intrinsic coordinates of the optimized query point are located in the voxel grid of the specific resolution; performing interpolation calculation on the eigenvectors of the vertices of the voxel grid of the specific resolution and the intrinsic coordinates of the optimized query point according to the coordinates of the vertices of the voxel grid of the specific resolution to obtain the eigenvectors of the intrinsic coordinates of the optimized query point in the voxel grid of the specific resolution; Repeating the vertex feature vector calculation operation and the interpolation calculation operation to obtain the feature vectors of the intrinsic coordinates of the optimized query point in voxel grids of different resolutions; The feature vectors of the intrinsic coordinates of the optimized query point in the voxel grids of different resolutions are vector-concatenated to obtain a high-dimensional feature vector of the intrinsic coordinates of the optimized query point.
5. The method according to claim 4, wherein The preset query formula is expressed by formula (2): Where z represents the coordinate of the vertex in the voxel grid, π i is a preset prime number, T represents the size of the hash table of the multi-level hash encoder of the trained implicit neural representation model, Represents the exclusive OR operation.
6. The method according to claim 1, wherein The multi-level first perceptron of the trained implicit neural representation model is expressed by formula (3): Δr=F φ (re t ) (3), Where r represents a point in the radiation density cube grid, Δr represents the offset corresponding to r, and e t is the conditional variable of the rough explicit geometry proxy at frame t, F Φ A multi-level first perceptron representing the trained implicit neural representation model; The multi-level second perceptron of the trained implicit neural representation model is expressed by formula (4): Among them, r is the intrinsic coordinate corresponding to the query point x, Δr represents the offset corresponding to r, σ t (x) represents the density of query point x in the tth frame, c t (x) represents the color of the query point x in the tth frame, represents the hash table of the multi-level hash encoder of the trained implicit neural representation model, h(*) represents the mapping of the intrinsic coordinates of the optimized query point to its high-dimensional feature vector, F ω A multi-layer second perceptron representing the trained implicit neural representation model.
7. The method according to claim 1, wherein The volume rendering formula is shown by formulas (5) and (6): a(x i )=1-exp(-σ(x i )d i ) (6), Among them, C(γ) represents the color of light γ, c(x i ) and σ(x i ) represent the sampling points x i The color and density values, δ i represents the spacing between sampling points, α(*) represents the sampling point x i opacity.
8. A method for training an implicit neural representation model, comprising: Extracting a real video frame image of the target dressed person based on a motion video of the target dressed person; Obtaining a synthesized video frame image of the target dressed person using the implicit neural representation model, wherein the implicit neural representation model includes a multi-level first perceptron, a multi-level second perceptron, and a multi-level hash encoder; Processing the synthesized video frame image and the real video frame image using a loss function, and optimizing the implicit neural representation model according to the loss value, wherein the loss function includes a photometric loss function and a regularization loss function; Iteratively perform real video frame image and synthesized video frame image acquisition operations and model optimization operations until preset conditions are reached to obtain a trained implicit neural representation model, wherein the trained implicit neural representation model is applied to any of the methods described in claims 1-7.
9. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to perform the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 8.