Method and device for reconstructing 360-degree head of monocular video based on 3DGS
By using monocular video for GAN inversion and multi-view data processing in the 3DGS model, the problem of generating full-view video in a single-view video was successfully solved, and high-definition details reconstruction and stability improvement of 360-degree human heads were achieved.
Patent Information
- Application Number
- CN202510130116.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-09
AI Technical Summary
In the absence of other perspective information, how to generate full-view videos using single-view videos, especially reconstructing high-definition details textures at full-view and solving the crash problem.
The method of reconstructing a 360-degree head based on 3DGS is adopted to reconstruct the 360-degree head by inputting the front face data of the head into the full head generation model SphereHead for GAN inversion, obtaining 360° full head data, using face 3D key point detection and affine transformation to process face data from multiple perspectives, and processing it through the key tuning inversion technology PTI. Finally, the data is input into the 3DGS model for model training, and a 360-degree head model is generated.
The 360° full head of high-definition details is realized based on 3DGS, which improves the clarity and stability of the reconstruction and solves the error problem caused by spatial conversion.
Smart Images

Figure CN119963741A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital human, and specifically provides a method and device for reconstructing a 360-degree human head based on 3DGS using a monocular video. Background Art
[0002] The emergence of 3D Gaussian Splatting (3DGS) has greatly accelerated the rendering speed of new view synthesis. Unlike neural implicit representations such as Neural Radiance Field (NeRF) that represent 3D scenes with position and viewpoint-conditioned neural networks, 3DGS uses a set of Gaussian ellipsoids to model the scene, thereby achieving efficient rendering by rasterizing the Gaussian ellipsoids into images.
[0003] For head modeling using 3DGS, MonoGaussianAvatar first applies 3DGS for dynamic head reconstruction, using canonical space modeling and deformation prediction. In addition, PSAvatar introduces an explicit Flame facial model to initialize the Gaussian, which can capture high-fidelity facial geometry and even complex volumetric objects (such as glasses). GaussianHead uses a three-plane representation and motion field to simulate a geometrically changing head in continuous motion and render rich textures, including skin and hair. In order to more easily control head expressions, GaussianAvatars introduces geometric priors (Flame parameterized facial model) into 3DGS, binds the Gaussian to an explicit mesh, and optimizes the parameters of the Gaussian ellipsoid.
[0004] Rig3DGS uses learnable deformations to provide stability and generalization to novel expressions, head poses, and viewing directions, enabling controllable portraits on portable devices. HeadGas endows 3DGS with a set of latent feature-based foundations weighted by the expression vectors of 3DMMs, enabling real-time animated head reconstruction.
[0005] FlashAvatar further embeds the uniform 3DGS field into the parameterized facial model and learns additional spatial offsets to capture facial details, successfully pushing the rendering speed to 300FPS. In order to synthesize high-resolution results, GaussianHead Avatar uses a super-resolution network to achieve high-fidelity avatar learning. GaussianHair combines the Marschner hair model with UE4's real-time hair rendering for the first time to create a Gaussian hair scattering model. It captures complex hair geometry and appearance for fast rasterization and volume rendering, enabling applications including editing and relighting.
[0006] However, when the input data only has frontal face data, a full-face digital human is needed. This situation often occurs in situations where higher fidelity and flexibility of the digital human are required, or in scenes where the digital human has a large range of movement. Most of the existing historical data are shot from the front, and there is no data from other perspectives.
[0007] The solution of the present invention is to solve the problem of generating a full-view video using a single-view video in the absence of other view information; at the same time, it focuses on solving the high-definition detail texture problem and the crash problem of reconstructing the full-view.
[0008] In view of this, this application is hereby filed. Summary of the invention
[0009] In view of the above technical problems, the present invention proposes a method and device for reconstructing a 360-degree human head based on 3DGS using a monocular video, and realizes a technical solution that can reconstruct a 360-degree full human head with high-definition details using a monocular video based on 3DGS. Specifically, the following technical solutions are adopted:
[0010] In a first aspect, the present invention provides a method for reconstructing a 360-degree human head based on a monocular video based on 3DGS, comprising:
[0011] Input the front face data of the human head into the full head generation model SphereHead for GAN inversion to generate 360° full head data;
[0012] Use the deep learning model for facial 3D key point detection to obtain the facial 3D key points in the 360° full head data;
[0013] Using the 3D facial key points to perform affine transformation, cropping and alignment, to obtain multi-view facial data;
[0014] The multi-view face data is processed using the key tuning inversion technology PTI, and the transformation matrix of the affine transformation is used for inverse transformation to obtain the multi-view face training data;
[0015] The multi-view face training data is used as pseudo data, and the front face data of the human head is used as real data to input into the 3DGS model for model training to generate a 360-degree head model.
[0016] As an optional embodiment of the present invention, in a method for reconstructing a 360-degree human head based on a monocular video of the present invention based on 3DGS, the method of using the 3D key points of the face to perform affine transformation, cropping and aligning, and obtaining multi-view face data includes:
[0017] The GFPGAN model is used to process multi-view facial data to reduce the regional deviation between facial data from side angles and back of the head and frontal face data.
[0018] As an optional embodiment of the present invention, in a method for reconstructing a 360-degree human head based on a monocular video based on 3DGS of the present invention, the method of processing multi-view face data using the key tuning inversion technology PTI includes:
[0019] Render a horizontal camera track and increase the camera radius so that the entire head is within the field of view;
[0020] Then, the transformation matrix of the affine transformation is used to perform an inverse transformation, and the transformation is transformed back to the original position to obtain multi-view face training data.
[0021] As an optional embodiment of the present invention, in a method for reconstructing a 360-degree human head based on a monocular video based on 3DGS of the present invention, the step of obtaining multi-view face training data includes:
[0022] Use the MODNet network to perform image segmentation processing on multi-view face training data to obtain head mask images of multi-view face training data;
[0023] The head mask image of the multi-view face training data is used as pseudo data, and the front face data of the head is used as real data to input into the 3DGS model for model training to generate a 360-degree head model.
[0024] As an optional embodiment of the present invention, in a method of reconstructing a 360-degree human head based on 3DGS from a monocular video of the present invention, the multi-perspective face training data is used as pseudo data, and the front face data of the human head is input into the 3DGS model as real data for model training, and the background color is randomized to alleviate the alignment defect deviation problem.
[0025] As an optional embodiment of the present invention, in a method for reconstructing a 360-degree human head based on a monocular video based on 3DGS of the present invention, inputting the front face data of the human head into the full head generation model SphereHead for GAN inversion includes:
[0026] The frontal facial data of human heads with calm expressions are selected and input into the full human head generation model SphereHead for GAN inversion.
[0027] As an optional embodiment of the present invention, in a method for reconstructing a 360-degree human head based on a monocular video based on 3DGS of the present invention, the frontal face data of the human head is input into a full head generation model SphereHead for GAN inversion, and after generating 360° full head data, the method includes:
[0028] Filter out low-confidence data in the 360° full head data, and then use the deep learning model for facial 3D key point detection to obtain the facial 3D key points in the 360° full head data.
[0029] In a second aspect, the present invention provides a device for reconstructing a 360-degree human head based on a monocular video based on 3DGS, comprising:
[0030] The full head generation module has a full head generation model SphereHead. The front face data of the head is input into the full head generation model SphereHead for GAN inversion to generate 360° full head data.
[0031] The face 3D key point detection module uses the deep learning model of face 3D key point detection to obtain the face 3D key points in 360° full head data;
[0032] An affine transformation module uses the 3D facial key points to perform affine transformation, crop and align, and obtain multi-view facial data;
[0033] Key tuning inversion module, which uses key tuning inversion technology PTI to process multi-view face data;
[0034] An affine inverse transformation module, for processing multi-view face data using the key tuning inversion technology PTI, uses the transformation matrix of the affine transformation to perform inverse transformation to obtain multi-view face training data;
[0035] The 3DGS module has a 3DGS model. It uses multi-view face training data as pseudo data and the front face data of the human head as real data to input into the 3DGS model for model training to generate a 360-degree human head model.
[0036] In a third aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory is used to store a computer executable program, and is characterized in that when the computer program is executed by the processor, the processor executes the method for reconstructing a 360-degree human head based on 3DGS using a monocular video.
[0037] In a fourth aspect, the present invention provides a computer-readable recording medium storing a computer executable program, characterized in that when the computer executable program is executed, a method for reconstructing a 360-degree human head based on a monocular video based on 3DGS is implemented.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] The present invention provides a method for reconstructing a 360-degree human head based on a monocular video based on 3DGS, uses a full human head generation model SphereHead instead of panohead, and proposes a method for reconstructing a 360-degree human head based on a monocular video based on 3DGS to connect generated data and 3D space, and optimize data at different angles so that the reconstruction process will not bring more errors due to space conversion, thereby obtaining a higher-definition and more stable 360° human head reconstruction structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 A flowchart of a method for reconstructing a 360-degree human head through 3DGS based on a monocular video according to a first embodiment of the present invention;
[0041] Figure 2 An example of a training strategy for a method of reconstructing a 360-degree human head through 3DGS based on a monocular video in Embodiment 1 of the present invention;
[0042] Figure 3 Embodiment 1 of the present invention is a method for reconstructing a 360-degree human head through 3DGS based on a monocular video to reconstruct a 360-degree human head. Figure 1 ;
[0043] Figure 4 Embodiment 1 of the present invention is a method for reconstructing a 360-degree human head through 3DGS based on a monocular video to reconstruct a 360-degree human head. Figure 2 ;
[0044] Figure 5 Embodiment 1 of the present invention is a method for reconstructing a 360-degree human head through 3DGS based on a monocular video to reconstruct a 360-degree human head. Figure 3 ;
[0045] Figure 6 Embodiment 1 of the present invention is a method for reconstructing a 360-degree human head through 3DGS based on a monocular video to reconstruct a 360-degree human head. Figure 4 ;
[0046] Figure 7 A flowchart of a densification method for single video 3DGS head reconstruction according to a second embodiment of the present invention;
[0047] Figure 8 An experimental table of a densification method for single video 3DGS head reconstruction according to the second embodiment of the present invention;
[0048] Fig. 9 A densified visualization effect diagram of a densification method for single video 3DGS head reconstruction according to the second embodiment of the present invention;
[0049] Fig.10A rendering effect diagram of a densification method for single video 3DGS head reconstruction according to the second embodiment of the present invention;
[0050] Fig.11 A flowchart of a Gaussian sphere updating method for improving single-video 3DGS head reconstruction according to a third embodiment of the present invention;
[0051] Fig.12 An experimental comparison table of a Gaussian sphere update method for improving single-video 3DGS head reconstruction according to Embodiment 3 of the present invention;
[0052] Fig.13 A diagram showing the experimental effect of a Gaussian sphere updating method for improving single-video 3DGS head reconstruction according to the third embodiment of the present invention;
[0053] Fig.14 A flowchart of a method for reconstructing a 360-degree human head based on 3DGS using a monocular video according to a fourth embodiment of the present invention;
[0054] Fig.15 Comparison of the effects of a method for reconstructing a 360-degree human head based on 3DGS using a monocular video in Example 4 of the present invention Figure 1 ;
[0055] Fig.16 Comparison of the effects of a method for reconstructing a 360-degree human head based on 3DGS using a monocular video in Example 4 of the present invention Figure 2 ;
[0056] Fig.17 Comparison of the effects of a method for reconstructing a 360-degree human head based on 3DGS using a monocular video in Example 4 of the present invention Figure 2 ;
[0057] Fig.18 A schematic structural diagram of an electronic device according to a fifth embodiment of the present invention;
[0058] Fig.19 A schematic diagram of a computer-readable recording medium according to a fifth embodiment of the present invention. DETAILED DESCRIPTION
[0059] To make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be described clearly and completely in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them.
[0060] Therefore, the following detailed description of the embodiments of the present invention is not intended to limit the scope of the invention claimed for protection, but merely represents some embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0061] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features and technical solutions in the embodiments may be combined with each other.
[0062] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0063] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the invention product is usually placed when in use, or the orientation or positional relationship commonly understood by those skilled in the art. Such terms are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.
[0064] Embodiment 1
[0065] See also Figure 1 As shown, this embodiment provides a method for reconstructing a 360-degree human head through 3DGS based on a monocular video, including:
[0066] The monocular video is used as input to the 3D perception GAN model for GAN inversion to obtain the side image data and back image data of 360-degree head reconstruction;
[0067] The side image data and back image data are used as pseudo data, and the monocular video is used as real data to input into the 3DGS model for model training to generate a 360-degree human head model.
[0068] In this embodiment, a method for reconstructing a 360-degree human head through 3DGS based on a monocular video is used. The side image data and the back image data are obtained by a generation method, and are absorbed as pseudo data to participate in the model training of real data through a training method to generate a 360-degree human head model.
[0069] Therefore, the method for reconstructing a 360-degree human head through 3DGS based on a monocular video of this embodiment can generate a 3D reconstruction result of the whole human head from a single-view video, and can cross-drive various motion amplitudes. The method for reconstructing a 360-degree human head through 3DGS based on a monocular video of this embodiment is the first solution that proposes to use a 3DGS solution to reconstruct 360° data using a single-view video.
[0070] Full head reconstruction: To reconstruct a 360° renderable head image from a monocular video with only the front view, a more powerful prior model must be used. We introduced PanoHead as a prior model, which is a 3D-aware GAN model that contains a lot of prior information about the side view and back view of the head. It mainly uses GAN inversion technology, and PTI is used for GAN inversion because it does not rely on additional encoder training and works well in the inversion of PanoHead.
[0071] GAN inversion refers to the process of mapping the target image into the latent space vector of the pre-trained GAN model, and then inputting the vector into the pre-trained generator to complete the image reconstruction. This process enables the pre-trained GAN model to be applied to real image editing and has strong interpretability.
[0072] Pivotal Tuning Inversion (PTI) is a technique for mapping real images to the latent space of StyleGAN, allowing for advanced editing. By fine-tuning the latent space of the generative model, PTI enables editing of real images, maintaining the accuracy of their identity features while reducing distortion.
[0073] Through experiments, we find that PanoHead performs well in generating static FFHQ aligned head images with small facial expression variations. In addition, we can also obtain head images with high-quality frontal rendering after the monocular reconstruction stage.
[0074] Therefore, we use images rendered with neutral expressions and poses as input for our analysis of avatars.
[0075] As an optional implementation of this embodiment, in a method of reconstructing a 360-degree human head through 3DGS based on a monocular video in this embodiment, before the monocular video is used as input and input into the 3D perception GAN model, the method includes:
[0076] The monocular video is input into a low-level vision pre-trained model for preprocessing, which injects more image-specific details into the wild input image of the monocular video to obtain high-quality input data.
[0077] Furthermore, in a method for reconstructing a 360-degree human head through 3DGS based on a monocular video in this embodiment, the low-level visual pre-training model is pre-trained based on a high-quality face dataset.
[0078] Face Restoration: The input images used for inversion are essentially in-the-wild images, which have a large domain gap from the high-quality dataset used for PanoHead training. Directly using rendered images for inversion, the optimized latent codes tend to break away from the GAN manifold, resulting in suboptimal images. The classic approach to address this problem is to apply domain adaptation or transfer learning. However, we can achieve the same goal more concisely. By processing the input images with a low-level vision model pre-trained on a high-quality face dataset, more image-specific details can be injected into the in-the-wild input images. This narrows the domain gap. We call this pre-trained model MR. In our experiments, we use the state-of-the-art face restoration model GFPGAN. It contains a generative prior on the FFHQ dataset that meets our needs.
[0079] As an optional implementation of this embodiment, a method for reconstructing a 360-degree human head through 3DGS based on a monocular video in this embodiment, wherein the 3D perception GAN model uses the Pivotal Tuning Inversion technology to perform GAN inversion;
[0080] Neutral face images of different angles are rendered with arbitrary camera postures, and neutral face images of different angles are filtered out through landmark detection to reconstruct good neutral face images. The filtered neutral face images are used to extend the Pivotal Tuning Inversion technology to multi-view situations.
[0081] Standard PTI takes as input a single image, then searches for latent codes and fine-tunes generator parameters. However, this approach tends to overfit the input image. In PanoHead, the performance of this naive approach drops dramatically in new views. It is worth noting that we can render neutral faces in arbitrary camera poses, and landmark detection allows us to filter out images with well-reconstructed faces. Therefore, we can use the filtered images to extend PTI to the multi-view case.
[0082] Specifically, in a method of reconstructing a 360-degree human head through 3DGS based on a monocular video in this embodiment, the use of a filtered neutral face image to extend the Pivotal Tuning Inversion technology to a multi-view situation includes:
[0083] Search key potential code wp:
[0084]
[0085] M is the number of valid multi-view images, LLPIPS is the perceptual loss, IMR is the face image after face restoration, GPano is the frozen full head generator, θ is the generator parameter, and c is the camera pose;
[0086] Adjust the parameters of the full head generator, and the loss setting is as follows
[0087]
[0088] θ* is the optimized parameter, initialized using the pre-trained model θ.
[0089] After adding the side image data and the back image data, we can use this additional data to reconstruct the human head, but training directly on these data will lead to the degradation of the frontal perspective, so we developed a training strategy as follows.
[0090] In a method for reconstructing a 360-degree human head through 3DGS based on a monocular video of this embodiment, the side image data and the back image data are used as pseudo data, and the monocular video is used as real data to input into the 3DGS model for model training to generate a 360-degree human head model, which includes:
[0091] Sampling from real datasets
[0092] pass and c i Rendering
[0093] pass And back propagation calculates the loss L1;
[0094] Update the 3D point cloud {G};
[0095] Where, I = {Ii} video frame sequence;
[0096] Ψ={Ψi} tracked task expression;
[0097] Θ={Θi}, character posture;
[0098] c i It's the camera posture.
[0099] In a method for reconstructing a 360-degree human head through 3DGS based on a monocular video of this embodiment, the side image data and the back image data are used as pseudo data, and the monocular video is used as real data to input into the 3DGS model for model training to generate a 360-degree human head model, which includes:
[0100] Sampling from a pseudo dataset Replace with Ψ i ;
[0101] pass and c j Rendering
[0102] pass And back propagation calculates the loss L2;
[0103] Update the 3D point cloud {G}.
[0104] A method of reconstructing a 360-degree human head through 3DGS based on a monocular video in this embodiment cross-trains between pseudo data and real data, similar to the training of an adversarial network. Because the Gaussian sphere behind the human head will be cropped during the training process, an additional Gaussian sphere set is initialized on the basis of the existing Gaussian field to model the back of the human head. Although each pseudo data image is generated using a neural network expression, we observe that most tables do not affect the appearance of the side and back. Therefore, when training pseudo data, expression Ψi is used instead of an all-zero vector.
[0105] Experiments show that the training strategy of this embodiment improves the effect on the sides and back of the head.
[0106] Specifically, in this embodiment, a method for reconstructing a 360-degree human head through 3DGS based on a monocular video is used. For example, the training strategy adopted is Figure 2 shown.
[0107] The effect of reconstructing a 360-degree human head through 3DGS based on a monocular video in this embodiment is shown as follows:
[0108] See also Figure 3 and Figure 4 As shown in the figure, the reconstruction effect of the head from the side and back is shown. It can be seen that compared with the best solutions on the market, the solution of this embodiment (the "ours" column in the figure) is better. Both multi-view and additional auxiliary training bring obvious positive benefits.
[0109] See also Figure 5 and Figure 6 As shown, the effects of self-driving and cross-driving are demonstrated. The scheme of this embodiment ( Figure 5 The "ours" column and Figure 6 The rightmost column in the figure all work well in expressing details.
[0110] This embodiment also provides a device for reconstructing a 360-degree human head through 3DGS based on a monocular video, including:
[0111] The GAN inversion module takes the monocular video as input and inputs it into the 3D perception GAN model of the GAN inversion module for GAN inversion to obtain the side image data and back image data of the 360-degree head reconstruction;
[0112] The 3DGS model module uses the side image data and the back image data as pseudo data and the monocular video as real data to input the 3DGS model of the 3DGS model module for model training to generate a 360-degree human head model.
[0113] In this embodiment, a device for reconstructing a 360-degree human head through 3DGS based on a monocular video is used. The GAN inversion module uses a generation method to obtain side image data and back image data, and absorbs them as pseudo data to participate in the model training of the 3DGS model module of real data through training, thereby generating a 360-degree human head model.
[0114] Therefore, the device for reconstructing a 360-degree human head based on a monocular video through 3DGS of this embodiment can generate a 3D reconstruction result of the whole human head from a single-view video, and can cross-drive various motion amplitudes. The method for reconstructing a 360-degree human head based on a monocular video through 3DGS of this embodiment is the first solution that proposes to use a 3DGS solution to reconstruct 360° data using a single-view video.
[0115] Embodiment 2
[0116] See also Figure 7 As shown, this embodiment provides a method for single video 3DGS head reconstruction, including a densification method:
[0117] Based on the video samples, a sparse point cloud containing three-dimensional position information is calculated;
[0118] Initialize the sparse point cloud to get a 3D Gaussian sphere;
[0119] The 3DGS model is trained based on the 3D Gaussian sphere. In the adaptive densification control process of the 3DGS model training, the densification index is introduced into the sampling probability function. The best densification path is automatically updated during the 3DGS model training process, and the 3DGS model is obtained after the training is completed.
[0120] A densification method for single-video 3DGS head reconstruction in this embodiment introduces a densification index into the sampling probability function during the adaptive densification control process of 3DGS model training, and upgrades the process of adding and deleting 3D Gaussian balls to a continuous and differentiable fully automatic process. This can more keenly perceive the changes in data during the training process, replace the processing method that requires manual setting of thresholds and is non-differentiable, and avoid human interference.
[0121] Furthermore, in a densification method for single video 3DGS head reconstruction in this embodiment, the sampling probability function is:
[0122]
[0123] in, is the importance index of a certain 3D Gaussian sphere, and the denominator represents the sum of the importance indexes of all 3D Gaussian spheres of the triangle face where the certain 3D Gaussian sphere is located.
[0124] It can be seen from the above sampling probability function formula that the larger the importance index value of a 3D Gaussian ball is, the more important it is, the greater the sampling probability is, and vice versa. This rule is consistent with the expectations of the densification process.
[0125] So what kind of importance index works best? As an optional implementation of this embodiment, in a densification method for single-video 3DGS head reconstruction in this embodiment, the importance index of the sampling probability function includes: maximum value index max(s), maximum value ratio index max(s) / min(s), parameter gradient index, transparency index; where s represents position information, that is, the gradient relative to the position coordinates (x, y, z) of the 3D Gaussian sphere.
[0126] Specifically, using the sampling probability functions of the above four different indicators, the 3D head reconstruction results are as follows: Figure 8 As shown in the table, overall, the gradient-based head reconstruction has the best effect, while the transparency-based reconstruction has the worst effect.
[0127] In addition, the effect comparison on the test set is Fig. 9 The densification visualization shown shows the distribution of splash positions during the densification process under different environments, focusing on the reconstruction effects of facial features and hair. It can be seen that the best reconstruction effect is the head reconstruction based on the gradient indicator, and the head reconstruction effect based on the maximum value indicator is the worst.
[0128] like Fig.10 As shown in the figure, the difference in effect between the "gradient-based densification strategy" and the "gradient-based densification strategy" can be observed from four angles: clarity, edge differentiation, whether hair is differentiated, and the distribution density of 3D Gaussian balls. In high-frequency areas, such as hair, eyes, nose, mouth, and eyebrows, after adding gradient-based probability sampling densification on the left, the Gaussian balls in these areas are obviously denser, which is in line with our expectations. More Gaussian balls are needed in key areas, while a small number of Gaussian balls can be used to represent smooth areas.
[0129] From the above comparison results, we can see that the 3DGS head reconstruction effect is better using the automatic continuously differentiable densification scheme.
[0130] It should be noted that this embodiment exemplifies the above four different importance indicators. Of course, the importance indicators in the sampling probability function include but are not limited to the above four indicators.
[0131] As an optional implementation of this embodiment, in a densification method for single video 3DGS head reconstruction of this embodiment, the sampling probability function adds different weights to different designated important areas of head reconstruction. In this way, the weight of the sampling probability function can be set based on each key area in the head reconstruction, and more attention is paid to the key areas in the 3D Gaussian sphere densification project of 3DGS, so that the densification effect is better and meets the needs of head reconstruction.
[0132] Specifically, in a single video 3DGS head reconstruction densification method of this embodiment, the 3DGS model training based on the 3D Gaussian sphere includes:
[0133] For every N iterations of the 3D Gaussian sphere, the order of the spherical harmonic coefficients is increased, where N is the set value;
[0134] Randomly select a camera perspective;
[0135] Render the image, obtaining the view space point, visibility filter, and radius information;
[0136] Calculate the loss L = (1-λ)L 1 +λL D-SSIM (L1 loss and L D-SSIM The weighted sum of the losses, L1 loss helps ensure pixel-level accuracy, and L D-SSIM The loss helps to maintain the overall structure and visual quality of the image), λ is the preset weight value, and back propagation is performed;
[0137] Subsequent operations via gradient-free contextual learning include:
[0138] Perform point cloud density operations based on the number of iterations:
[0139] Update the maximum radius information;
[0140] Increase and trim point cloud density based on conditions;
[0141] Update the optimizer parameters;
[0142] During the point cloud density operation, the densification index is introduced into the sampling probability function, which is:
[0143]
[0144] in, is the importance index of a certain 3D Gaussian sphere, and the denominator represents the sum of the importance indexes of all 3D Gaussian spheres of the triangle face where the certain 3D Gaussian sphere is located.
[0145] A densification method for single video 3DGS head reconstruction in this embodiment includes:
[0146] Shoot an image or video of the target person, input the image or video into the 3DGS model, and generate a 3D reconstruction model of the target person.
[0147] Embodiment 3
[0148] See also Fig.11 As shown, this embodiment provides a method for improving single video 3DGS head reconstruction, including a Gaussian sphere update method:
[0149] Based on the captured video, a sparse point cloud containing three-dimensional position information is calculated;
[0150] Initialize the sparse point cloud to get a 3D Gaussian sphere, and reconstruct the 3DGS head based on the 3D Gaussian sphere;
[0151] Initialize the 3D Gaussian sphere position to obtain the human head point cloud;
[0152] A new 3D Gaussian sphere is expanded outside the human head point cloud. Through a predefined loop scheduling program, the offset δ of the new 3D Gaussian sphere is set based on the 3D Gaussian sphere position of the human head point cloud, and a shell-like structure is constructed to obtain the human head hair point cloud.
[0153] During 3DGS training, the 3D Gaussian sphere will continue to grow or shrink. When the task of head reconstruction is heavy, we will put the 3D Gaussian sphere on the head model during initialization, such as the surface of 3DMM. We then hope that the 3D Gaussian sphere can be generated outside the head to reconstruct the hair, but we do not want the Gaussian sphere to grow inside the head, because the latter is meaningless.
[0154] The present embodiment provides a Gaussian sphere updating method for improving the efficiency of single-video 3DGS head reconstruction, which is to solve the problem of effectively generating a Gaussian sphere in a directional outward direction, expand a new 3D Gaussian sphere outside the point cloud of the human head skull, and set the offset δ of the new 3D Gaussian sphere based on the 3D Gaussian sphere position of the human head skull point cloud through a predefined loop scheduling program to construct a shell-like structure and obtain a point cloud of human head hair. Therefore, the present embodiment provides a Gaussian sphere updating method for improving the efficiency of single-video 3DGS head reconstruction, which achieves the effect of spreading the 3D Gaussian sphere outward along the hair, and does not waste computing power in areas without hair.
[0155] Furthermore, in a Gaussian sphere updating method for improving single-video 3DGS head reconstruction of this embodiment, a new 3D Gaussian sphere is expanded outside the head point cloud, and an offset δ of the new 3D Gaussian sphere is set based on the 3D Gaussian sphere position of the head point cloud through a predefined loop scheduling program to construct a shell-like structure, and the head hair point cloud is obtained, including:
[0156] p(i) represents the position j of the new 3D Gaussian sphere on the i-th triangle patch, and δ(i) is the offset assigned to the new 3D Gaussian sphere. The growth formula of the new 3D Gaussian sphere is:
[0157]
[0158] in, is the 3D Gaussian ball growth formula of the human head point cloud, n(i) is the expansion direction parameter from the i-th triangle patch, It is the sampling point Coordinates on the i-th triangle patch; are the coordinates of three points on the i-th triangle patch.
[0159] Since the 3D Gaussian sphere is pruned and densified during the 3DGS model training process, we can eventually form a shell-like structure in a progressive and unsupervised manner, which can effectively capture fine-scale details.
[0160] As an optional implementation of this embodiment, in a Gaussian sphere update method for improving single-video 3DGS head reconstruction of this embodiment, the n(i) is the normal vector from the vertex of the triangle patch whose outward expansion direction parameter is the i-th triangle patch.
[0161] As an optional implementation of this embodiment, in a Gaussian sphere update method for improving single-video 3DGS head reconstruction of this embodiment, the n(i) is the midline of the triangle patch whose outward expansion direction parameter is the i-th triangle patch.
[0162] As an optional implementation of this embodiment, in a Gaussian sphere update method for improving single-video 3DGS head reconstruction of this embodiment, the n(i) is a perpendicular line from the i-th triangle patch whose outward expansion direction parameter is a triangle patch.
[0163] Specifically, in a Gaussian sphere updating method for improving single-video 3DGS head reconstruction of this embodiment, the offset δ of setting a new 3D Gaussian sphere based on the 3D Gaussian sphere position of the head point cloud includes:
[0164] The value range of the preset offset δ is [0, δmax], where δmax is the preset maximum offset;
[0165] Based on the 3D Gaussian sphere position of the human head point cloud, the new 3D Gaussian sphere offset δ is set in the value range [0, δmax].
[0166] This embodiment ensures that the 3D Gaussian sphere is expanded within the numerical range [0, δmax] by presetting the maximum offset δmax, constructs a shell-like structure, obtains the point cloud of the hair on the human head, and avoids wasting computing power in areas without hair.
[0167] As an optional implementation of this embodiment, a Gaussian sphere update method for improving the efficiency of single-video 3DGS head reconstruction in this embodiment, initializing the 3D Gaussian sphere position to obtain a human head skull point cloud includes: initializing the 3D Gaussian sphere position and placing it on the surface of the 3DMM head model to obtain a human head skull point cloud.
[0168] Experiments on eight open-source human heads on the market show that the new 3D Gaussian sphere growth formula of this embodiment can reduce the 3D Gaussian sphere calculation by half compared with the random point scattering method.
[0169] In terms of effect, see Fig.12 As shown in the comparison table, the "w / oδ" column is the experimental result data of the new 3D Gaussian sphere growth formula of this embodiment. The accurate calculation of the 3D Gaussian sphere position can improve the reconstruction effect. For specific visualization effects, see Fig.13 shown.
[0170] Embodiment 4
[0171] See also Fig.14 As shown, a method for reconstructing a 360-degree human head based on a monocular video based on 3DGS in this embodiment includes:
[0172] Input the front face data of the human head into the full head generation model SphereHead for GAN inversion to generate 360° full head data;
[0173] Use the deep learning model for facial 3D key point detection to obtain the facial 3D key points in the 360° full head data;
[0174] Using the 3D facial key points to perform affine transformation, cropping and alignment, to obtain multi-view facial data;
[0175] The multi-view face data is processed using the key tuning inversion technology PTI, and the transformation matrix of the affine transformation is used for inverse transformation to obtain the multi-view face training data;
[0176] The multi-view face training data is used as pseudo data, and the front face data of the human head is used as real data to input into the 3DGS model for model training to generate a 360-degree head model.
[0177] Embodiment 1 proposes a method for reconstructing a 360-degree human head through 3DGS based on monocular video, but the solution still has flaws in actual use, mainly reflected in the fact that the clarity of the reconstructed image is not very high. The direct cause of this problem is that the generative model PanoHead is used in the solution. The generative model itself has some problems. For example, due to the lack of spatial perception ability, the generation effect of the back of the head is unstable. In addition, after using the generated back of the head, subsequent operations in 3DGS are required, so alignment is very important. Misaligned data will introduce new errors, resulting in poor reconstruction effects. These are the pain points of the previous solutions.
[0178] In this embodiment, a method for reconstructing a 360-degree human head from a monocular video based on 3DGS is used. The whole head generation model SphereHead is used instead of panohead, and a method for reconstructing a 360-degree human head from a monocular video based on 3DGS is proposed to connect the generated data and the 3D space, and optimize the data at different angles so that the reconstruction process will not bring more errors due to space conversion, thereby obtaining a higher-definition and more stable 360° head reconstruction structure.
[0179] The full head generation model SphereHead is a novel three-plane representation in a spherical coordinate system, which conforms to the geometric characteristics of the human head and effectively alleviates many generated artifacts. A view-image consistency loss is further introduced for the discriminator to emphasize the correspondence between camera parameters and images. The combination of these efforts brings visually superior results and greatly reduces artifacts.
[0180] In this embodiment, a method for reconstructing a 360-degree human head from a monocular video based on 3DGS is used. Multi-view face training data is used as pseudo data, and the front face data of the human head is used as real data to be input into a 3DGS model for model training. The degradation of the front face data is avoided by cross-training with pseudo data and real data.
[0181] This embodiment uses a deep learning model for facial 3D key point detection to obtain the 3D facial key points in 360° full head data, and specifically uses the TDDFA model. The TDDFA model is a deep learning model for facial 3D key point detection. To use TDDFA to obtain facial 3D key points, you need to first install the model and dependent libraries, then load the model and perform inference on the face image.
[0182] As an optional implementation of this embodiment, in a method for reconstructing a 360-degree human head based on 3DGS using a monocular video in this embodiment, the method uses the 3D key points of the face to perform affine transformation, and then crops and aligns the face data from multiple perspectives to obtain the following steps:
[0183] The GFPGAN model is used to process multi-view facial data to reduce the regional deviation between facial data from side angles and back of the head and frontal face data.
[0184] GFPGAN has a wide range of applications in face image generation and restoration. It can be used to convert low-resolution face images into high-resolution images to improve image clarity; it can also be used to restore high-quality face images from low-quality images and repair blur, noise and other problems in the image.
[0185] As an optional implementation of this embodiment, in a method for reconstructing a 360-degree human head based on 3DGS from a monocular video of this embodiment, the method of processing multi-view face data using the key tuning inversion technology PTI includes:
[0186] Render a horizontal camera track and increase the camera radius so that the entire head is within the field of view;
[0187] Then, the transformation matrix of the affine transformation is used to perform an inverse transformation, transforming back to the original position to obtain multi-view face training data.
[0188] As an optional implementation of this embodiment, in a method for reconstructing a 360-degree human head based on a monocular video based on 3DGS in this embodiment, the method of obtaining multi-view face training data includes:
[0189] Use the MODNet network to perform image segmentation processing on multi-view face training data to obtain head mask images of multi-view face training data;
[0190] The head mask image of the multi-view face training data is used as pseudo data, and the front face data of the head is used as real data to input into the 3DGS model for model training to generate a 360-degree head model.
[0191] MODNET is a deep learning model designed specifically for image segmentation and processing tasks. Its main purpose is to quickly and accurately extract regions of interest (ROI) from images in an automated way to meet the needs of different application scenarios.
[0192] As an optional implementation of this embodiment, in a method of reconstructing a 360-degree human head based on 3DGS from a monocular video of this embodiment, the multi-perspective face training data is used as pseudo data, and the front face data of the human head is used as real data to input into the 3DGS model for model training, and the background color is randomized to alleviate the alignment defect deviation problem.
[0193] As an optional implementation of this embodiment, in a method for reconstructing a 360-degree human head based on 3DGS from a monocular video of this embodiment, the inputting of the front face data of the human head into the full head generation model SphereHead for GAN inversion includes:
[0194] The frontal facial data of human heads with calm expressions are selected and input into the full human head generation model SphereHead for GAN inversion.
[0195] As an optional implementation of this embodiment, in a method for reconstructing a 360-degree human head based on 3DGS from a monocular video of this embodiment, the frontal face data of the human head is input into a full head generation model SphereHead for GAN inversion, and after generating 360° full head data, the method includes:
[0196] Filter out low-confidence data in the 360° full head data, and then use the deep learning model for facial 3D key point detection to obtain the facial 3D key points in the 360° full head data.
[0197] See also Fig.15 As shown, Fig.15 The first row is the effect diagram generated by the method for reconstructing a 360-degree human head based on 3DGS using a monocular video of this embodiment, and the second row is the effect diagram generated by the method for reconstructing a 360-degree human head based on 3DGS using a monocular video of this embodiment. It can be seen that after the method for reconstructing a 360-degree human head based on 3DGS using a monocular video of this embodiment is adopted, the reconstruction effect of the side and back sides is significantly improved.
[0198] Depend on Fig.15 By comparison, we can see that:
[0199] 1. The method of reconstructing a 360-degree human head based on 3DGS using a monocular video in this embodiment has a significant effect gain in reconstructing a full human head;
[0200] 2. Compared with SOTA, in the task of full head reconstruction, the method of reconstructing a 360-degree human head from a monocular video based on 3DGS proposed in this embodiment has the best effect.
[0201] See also Fig.16 As shown in the figure, a method for reconstructing a 360-degree human head based on 3DGS using a monocular video of this embodiment (the ours column in the figure) is compared with SOTA. Since SOTA does not do a full human head, only the detailed effects of the front and small-angle side faces are compared here.
[0202] See also Fig.17As shown, more effects are displayed, namely, reconstruction effects from different perspectives. The odd-numbered rows are effect diagrams generated by the method for reconstructing a 360-degree human head based on 3DGS using a monocular video of the present embodiment, and the even-numbered rows are effect diagrams generated by the method for reconstructing a 360-degree human head based on 3DGS using a monocular video of the present embodiment [row numbers start from 1].
[0203] This embodiment also provides a device for reconstructing a 360-degree human head based on a monocular video based on 3DGS, including:
[0204] The full head generation module has a full head generation model SphereHead. The front face data of the head is input into the full head generation model SphereHead for GAN inversion to generate 360° full head data.
[0205] The face 3D key point detection module uses the deep learning model of face 3D key point detection to obtain the face 3D key points in 360° full head data;
[0206] An affine transformation module uses the 3D facial key points to perform affine transformation, crop and align, and obtain multi-view facial data;
[0207] Key tuning inversion module, which uses key tuning inversion technology PTI to process multi-view face data;
[0208] An affine inverse transformation module, for processing multi-view face data using the key tuning inversion technology PTI, uses the transformation matrix of the affine transformation to perform inverse transformation to obtain multi-view face training data;
[0209] The 3DGS module has a 3DGS model. It uses multi-view face training data as pseudo data and the front face data of the human head as real data to input into the 3DGS model for model training to generate a 360-degree human head model.
[0210] Embodiment 5
[0211] The following describes an electronic device embodiment of the present invention, which can be regarded as a specific physical implementation of the method and device embodiments of the present invention. The details described in the electronic device embodiment of the present invention should be regarded as a supplement to the above method or device embodiments; details not disclosed in the electronic device embodiment of the present invention can be implemented with reference to the above method or device embodiments.
[0212] Fig.18It is a structural schematic diagram of an electronic device of an embodiment of the present invention, the electronic device includes a processor and a memory, the memory is used to store a computer executable program, when the computer program is executed by the processor, the processor executes a single video 3DGS head reconstruction method of embodiment one, or two, or three, or four.
[0213] like Fig.18 As shown, the electronic device is presented in the form of a general computing device. The processor may be one or more and work in coordination. The present invention does not exclude distributed processing, that is, the processor may be dispersed in different physical devices. The electronic device of the present invention is not limited to a single entity, but may also be the sum of multiple physical devices.
[0214] The memory stores a computer executable program, which is usually a machine-readable code. The computer-readable program can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least part of the steps in the method.
[0215] The memory includes a volatile memory, such as a random access memory unit (RAM) and / or a cache memory unit, and may also be a non-volatile memory, such as a read-only memory unit (ROM).
[0216] Optionally, in this embodiment, the electronic device further includes an I / O interface, which is used for the electronic device to exchange data with an external device. The I / O interface can represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.
[0217] It should be understood that Fig.18 The electronic device shown is only an example of the present invention, and the electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as display screens, and some electronic devices also include human-computer interaction elements such as buttons, keyboards, etc. As long as the electronic device can execute the computer-readable program in the memory to implement the method of the present invention or at least part of the steps of the method, it can be considered as an electronic device covered by the present invention.
[0218] Fig.19 Schematic diagram of a computer readable recording medium according to an embodiment of the present invention. Fig.19As shown, a computer executable program is stored in a computer-readable recording medium, and when the computer executable program is executed, a single-video 3DGS head reconstruction method of Embodiment 1, or 2, or 3, or 4 of the present invention is implemented. The computer-readable recording medium may include a data signal propagated in a baseband or as part of a carrier, which carries a readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable recording medium may also be any readable medium other than a readable recording medium, which may send, propagate, or transmit a program for use by or in combination with an instruction execution system, device, or device. The program code contained on the readable recording medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0219] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0220] Through the above description of the implementation mode, it is easy for those skilled in the art to understand that the present invention can be implemented by hardware capable of executing a specific computer program, such as the system of the present invention, and the electronic processing unit, server, client, mobile phone, control unit, processor, etc. contained in the system. The present invention can also be implemented by computer software that executes the method of the present invention, such as control software executed by a microprocessor, an electronic control unit, a client, a server, etc. However, it should be noted that the computer software that executes the method of the present invention is not limited to being executed by one or a specific hardware entity, and it can also be implemented in a distributed manner by unspecified specific hardware. For computer software, the software product can be stored in a computer-readable recording medium (which can be a CD-ROM, a USB flash drive, a mobile disk, etc.), and can also be distributed and stored on the network, as long as it enables the electronic device to execute the method according to the present invention.
[0221] The above embodiments are only used to illustrate the present invention and are not intended to limit the technical solutions described in the present invention. Although the present invention has been described in detail with reference to the above embodiments, the present invention is not limited to the above specific implementation methods. Therefore, any modification or equivalent replacement of the present invention; and all technical solutions and improvements thereof that do not depart from the spirit and scope of the invention are included in the scope of the claims of the present invention.
Claims
1. A method for reconstructing a 360-degree human head from a monocular video based on 3DGS, characterized in that: include: Input the front face data of the human head into the full head generation model SphereHead for GAN inversion to generate 360° full head data; Use the deep learning model for facial 3D key point detection to obtain the facial 3D key points in the 360° full head data; Using the 3D facial key points to perform affine transformation, cropping and alignment, to obtain multi-view facial data; The multi-view face data is processed using the key tuning inversion technology PTI, and the transformation matrix of the affine transformation is used for inverse transformation to obtain the multi-view face training data; The multi-view face training data is used as pseudo data, and the front face data of the human head is used as real data to input into the 3DGS model for model training to generate a 360-degree head model.
2. The method for reconstructing a 360-degree human head from a monocular video based on 3DGS according to claim 1, characterized in that: The method of using the 3D facial key points to perform affine transformation, cropping and aligning to obtain multi-view facial data includes: The GFPGAN model is used to process multi-view facial data to reduce the regional deviation between facial data from side angles and back of the head and frontal face data.
3. The method for reconstructing a 360-degree human head based on a monocular video based on 3DGS according to claim 1, characterized in that: The method of processing multi-view face data using the key tuning inversion technology PTI includes: Render a horizontal camera track and increase the camera radius so that the entire head is within the field of view; Then, the transformation matrix of the affine transformation is used to perform an inverse transformation, and the transformation is transformed back to the original position to obtain multi-view face training data.
4. The method for reconstructing a 360-degree human head based on a monocular video based on 3DGS according to claim 3, characterized in that: The step of obtaining multi-view face training data includes: Use the MODNet network to perform image segmentation processing on multi-view face training data to obtain head mask images of multi-view face training data; The head mask image of the multi-view face training data is used as pseudo data, and the front face data of the head is used as real data to input into the 3DGS model for model training to generate a 360-degree head model.
5. The method for reconstructing a 360-degree human head based on a monocular video based on 3DGS according to claim 1, characterized in that: In the process of inputting the multi-view face training data as pseudo data and the front face data of the human head as real data into the 3DGS model for model training, the background color is randomized to alleviate the alignment defect deviation problem.
6. The method for reconstructing a 360-degree human head based on a monocular video based on 3DGS according to claim 1, characterized in that: The step of inputting the front face data of a human head into the full human head generation model SphereHead for GAN inversion includes: The frontal facial data of human heads with calm expressions are selected and input into the full human head generation model SphereHead for GAN inversion.
7. The method for reconstructing a 360-degree human head based on a monocular video based on 3DGS according to claim 1, characterized in that: The front face data of the human head is input into the full head generation model SphereHead for GAN inversion, and the 360° full head data is generated, which includes: Filter out low-confidence data in the 360° full head data, and then use the deep learning model for facial 3D key point detection to obtain the facial 3D key points in the 360° full head data.
8. A device for reconstructing a 360-degree human head from a monocular video based on 3DGS, characterized in that: include: The full head generation module has a full head generation model SphereHead. The front face data of the head is input into the full head generation model SphereHead for GAN inversion to generate 360° full head data. The face 3D key point detection module uses the deep learning model of face 3D key point detection to obtain the face 3D key points in 360° full head data; An affine transformation module uses the 3D facial key points to perform affine transformation, crop and align, and obtain multi-view facial data; Key tuning inversion module, which uses key tuning inversion technology PTI to process multi-view face data; An affine inverse transformation module, for processing multi-view face data using the key tuning inversion technology PTI, uses the transformation matrix of the affine transformation to perform inverse transformation to obtain multi-view face training data; The 3DGS module has a 3DGS model. It uses multi-view face training data as pseudo data and the front face data of the human head as real data to input into the 3DGS model for model training to generate a 360-degree human head model.
9. An electronic device comprising a processor and a memory, wherein the memory is used to store a computer executable program, characterized in that: When the computer program is executed by the processor, the processor executes the method for reconstructing a 360-degree human head based on 3DGS from a monocular video as described in any one of claims 1 to 7.
10. A computer-readable recording medium storing a computer-executable program, characterized in that: When the computer executable program is executed, a method for reconstructing a 360-degree human head based on 3DGS using a monocular video as described in any one of claims 1 to 7 is implemented.