Method and device for reconstructing 360-degree head through 3DGS based on monocular video
Through the combination of the 3DGS model and the 3D perceptual GAN model, a 360-degree human head model is generated using monocular video, which solves the problem of generating a full-view video in a single-view video, and achieves a high fidelity and flexibility in digital human reconstruction effect.
Patent Information
- Application Number
- CN202510130119.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-06
AI Technical Summary
In the absence of other perspective information, how to use single-view videos to generate full-view videos, especially in situations where there is higher demand for digital people's fidelity and flexibility.
Through the 3DGS model, monocular video is input to the 3D-aware GAN model for GAN inversion, and the side and back image data of the 360-degree human head are generated, and these data are input as pseudo-data to the 3DGS model for model training, and a 360-degree human head model is generated.
It realizes the 3D reconstruction results of the whole head from a single-view video, and supports cross-driven of various action amplitudes, improving the performance of digital people in fidelity and flexibility.
Smart Images

Figure CN119941999A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital humans, and specifically provides a method and device for reconstructing a 360-degree human head through 3DGS based on a monocular video. Background Art
[0002] The emergence of 3D Gaussian Splatting (3DGS) has greatly accelerated the rendering speed of new view synthesis. Unlike neural implicit representations such as Neural Radiance Field (NeRF) that represent 3D scenes with position and viewpoint-conditioned neural networks, 3DGS uses a set of Gaussian ellipsoids to model the scene, thereby achieving efficient rendering by rasterizing the Gaussian ellipsoids into images.
[0003] For head modeling using 3DGS, MonoGaussianAvatar first applies 3DGS for dynamic head reconstruction, using canonical space modeling and deformation prediction. In addition, PSAvatar introduces an explicit Flame facial model to initialize the Gaussian, which can capture high-fidelity facial geometry and even complex volumetric objects (such as glasses). GaussianHead uses a three-plane representation and motion field to simulate a geometrically changing head in continuous motion and render rich textures, including skin and hair. In order to more easily control head expressions, GaussianAvatars introduces geometric priors (Flame parameterized facial model) into 3DGS, binds the Gaussian to an explicit mesh, and optimizes the parameters of the Gaussian ellipsoid.
[0004] Rig3DGS uses learnable deformations to provide stability and generalization to novel expressions, head poses, and viewing directions, enabling controllable portraits on portable devices. HeadGas endows 3DGS with a set of latent feature-based foundations weighted by the expression vectors of 3DMMs, enabling real-time animated head reconstruction.
[0005] FlashAvatar further embeds the uniform 3DGS field into the parameterized facial model and learns additional spatial offsets to capture facial details, successfully pushing the rendering speed to 300FPS. In order to synthesize high-resolution results, GaussianHead Avatar uses a super-resolution network to achieve high-fidelity head portrait learning. GaussianHai r combines the Marschner hair model with UE4's real-time hair rendering for the first time to create a Gaussian hair scattering model. It captures complex hair geometry and appearance for fast rasterization and volume rendering, enabling applications including editing and relighting.
[0006] However, when the input data only has frontal face data, a full-face digital human is needed. This situation often occurs in situations where higher fidelity and flexibility of the digital human are required, or in scenes where the digital human has a large range of movement. Most of the existing historical data are shot from the front, and there is no data from other perspectives.
[0007] The solution of the present invention is to solve the problem of generating a full-view video using a single-view video in the absence of other view information.
[0008] In view of this, this application is hereby filed. Summary of the invention
[0009] In view of the above technical problems, the present invention proposes a method and device for reconstructing a 360-degree human head through 3DGS based on a monocular video, so as to realize a technical solution for reconstructing a 360-degree high-fidelity video from a single video based on 3DGS.
[0010] Specifically, the following technical solutions are adopted:
[0011] In a first aspect, the present invention provides a method for reconstructing a 360-degree human head through 3DGS based on a monocular video, comprising:
[0012] The monocular video is used as input to the 3D perception GAN model for GAN inversion to obtain the side image data and back image data of 360-degree head reconstruction;
[0013] The side image data and back image data are used as pseudo data, and the monocular video is used as real data to input into the 3DGS model for model training to generate a 360-degree human head model.
[0014] As an optional embodiment of the present invention, in a method of reconstructing a 360-degree human head through 3DGS based on a monocular video of the present invention, before the monocular video is input as an input to the 3D perception GAN model, the method includes:
[0015] The monocular video is input into a low-level vision pre-trained model for preprocessing, which injects more image-specific details into the wild input image of the monocular video to obtain high-quality input data.
[0016] As an optional embodiment of the present invention, in a method of reconstructing a 360-degree human head through 3DGS based on a monocular video of the present invention, the low-level visual pre-training model is pre-trained based on a high-quality face dataset.
[0017] As an optional embodiment of the present invention, in a method of reconstructing a 360-degree human head through 3DGS based on a monocular video of the present invention, the 3D perception GAN model uses a pivotal tuning inversion technique to perform GAN inversion;
[0018] Neutral face images of different angles are rendered with arbitrary camera postures, and neutral face images of different angles are filtered out through landmark detection to reconstruct good neutral face images. The filtered neutral face images are used to extend the Pivotal Tuning Inversion technology to multi-view situations.
[0019] As an optional embodiment of the present invention, in a method of reconstructing a 360-degree human head through 3DGS based on a monocular video of the present invention, the use of a filtered neutral face image to extend the Pivotal Tuning Inversion technology to a multi-view situation includes:
[0020] Search key potential code wp:
[0021]
[0022] M is the number of valid multi-view graphs, L LPIPS is the perceptual loss, I MR is the face image after face restoration, G Pano is the frozen full head generator, θ is the generator parameter, and c is the camera pose;
[0023] Adjust the parameters of the full head generator, and the loss setting is as follows
[0024]
[0025] θ* is the optimized parameter, initialized using the pre-trained model θ.
[0026] As an optional embodiment of the present invention, in a method of reconstructing a 360-degree human head through 3DGS based on a monocular video of the present invention, the side image data and the back image data are used as pseudo data, and the monocular video is used as real data to input into the 3DGS model for model training, and generating a 360-degree human head model includes:
[0027] Sampling from real datasets
[0028] pass and c i Rendering
[0029] pass And back propagation calculates the loss L1;
[0030] Update the 3D point cloud {G};
[0031] Where, I = {Ii} video frame sequence;
[0032] Ψ={Ψi} tracked task expression;
[0033] Θ={Θi}, character posture;
[0034] c i It's the camera posture.
[0035] As an optional embodiment of the present invention, in a method of reconstructing a 360-degree human head through 3DGS based on a monocular video of the present invention, the side image data and the back image data are used as pseudo data, and the monocular video is used as real data to input into the 3DGS model for model training, and generating a 360-degree human head model includes:
[0036] Sampling from a pseudo dataset Replace with Ψ i ;
[0037] pass and c j Rendering
[0038] pass And back propagation calculates the loss L2;
[0039] Update the 3D point cloud {G}.
[0040] In a second aspect, the present invention provides a device for reconstructing a 360-degree human head through 3DGS based on a monocular video, comprising:
[0041] The GAN inversion module takes the monocular video as input and inputs it into the 3D perception GAN model of the GAN inversion module for GAN inversion to obtain the side image data and back image data of the 360-degree head reconstruction;
[0042] The 3DGS model module uses the side image data and the back image data as pseudo data and the monocular video as real data to input the 3DGS model of the 3DGS model module for model training to generate a 360-degree human head model.
[0043] In a third aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory is used to store a computer executable program. When the computer program is executed by the processor, the processor executes the method of reconstructing a 360-degree human head based on a monocular video through 3DGS.
[0044] In a fourth aspect, the present invention provides a computer-readable recording medium storing a computer-executable program, which, when executed, implements the method of reconstructing a 360-degree human head through 3DGS based on a monocular video.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] The present invention provides a method for reconstructing a 360-degree human head through 3DGS based on a monocular video, which uses a generation method to obtain side image data and back image data, and absorbs them as pseudo data to participate in the model training of real data through a training method to generate a 360-degree human head model.
[0047] Therefore, the method of reconstructing a 360-degree human head through 3DGS based on a monocular video of the present invention can generate a 3D reconstruction result of the whole human head from a single-view video, and can cross-drive various motion amplitudes. The method of reconstructing a 360-degree human head through 3DGS based on a monocular video of this embodiment is the first solution that proposes to use a 3DGS solution to reconstruct 360° data using a single-view video. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A flowchart of a method for reconstructing a 360-degree human head through 3DGS based on a monocular video according to a first embodiment of the present invention;
[0049] Figure 2 An example of a training strategy for a method of reconstructing a 360-degree human head through 3DGS based on a monocular video in Embodiment 1 of the present invention;
[0050] Figure 3 Embodiment 1 of the present invention is a method for reconstructing a 360-degree human head through 3DGS based on a monocular video to reconstruct a 360-degree human head. Figure 1 ;
[0051] Figure 4 Embodiment 1 of the present invention is a method for reconstructing a 360-degree human head through 3DGS based on a monocular video to reconstruct a 360-degree human head. Figure 2 ;
[0052] Figure 5 Embodiment 1 of the present invention is a method for reconstructing a 360-degree human head through 3DGS based on a monocular video to reconstruct a 360-degree human head. Figure 3 ;
[0053] Figure 6 Embodiment 1 of the present invention is a method for reconstructing a 360-degree human head through 3DGS based on a monocular video to reconstruct a 360-degree human head. Figure 4 ;
[0054] Figure 7 A flowchart of a densification method for single video 3DGS head reconstruction according to a second embodiment of the present invention;
[0055] Figure 8 An experimental table of a densification method for single video 3DGS head reconstruction according to the second embodiment of the present invention;
[0056] Fig. 9 A densified visualization effect diagram of a densification method for single video 3DGS head reconstruction according to the second embodiment of the present invention;
[0057] Fig.10 A rendering effect diagram of a densification method for single video 3DGS head reconstruction according to the second embodiment of the present invention;
[0058] Fig.11 A flowchart of a Gaussian sphere updating method for improving single-video 3DGS head reconstruction according to a third embodiment of the present invention;
[0059] Fig.12 An experimental comparison table of a Gaussian sphere update method for improving single-video 3DGS head reconstruction according to Embodiment 3 of the present invention;
[0060] Fig.13 A diagram showing the experimental effect of a Gaussian sphere updating method for improving single-video 3DGS head reconstruction according to the third embodiment of the present invention;
[0061] Fig.14 A schematic structural diagram of an electronic device according to a fourth embodiment of the present invention;
[0062] Fig.15 A schematic diagram of a computer-readable recording medium according to a fourth embodiment of the present invention. DETAILED DESCRIPTION
[0063] To make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be described clearly and completely in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them.
[0064] Therefore, the following detailed description of the embodiments of the present invention is not intended to limit the scope of the invention claimed for protection, but merely represents some embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0065] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features and technical solutions in the embodiments may be combined with each other.
[0066] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0067] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the invention product is usually placed when in use, or the orientation or positional relationship commonly understood by those skilled in the art. Such terms are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.
[0068] Embodiment 1
[0069] See also Figure 1 As shown, this embodiment provides a method for reconstructing a 360-degree human head through 3DGS based on a monocular video, including:
[0070] The monocular video is used as input to the 3D perception GAN model for GAN inversion to obtain the side image data and back image data of 360-degree head reconstruction;
[0071] The side image data and back image data are used as pseudo data, and the monocular video is used as real data to input into the 3DGS model for model training to generate a 360-degree human head model.
[0072] In this embodiment, a method for reconstructing a 360-degree human head through 3DGS based on a monocular video is used. The side image data and the back image data are obtained by a generation method, and are absorbed as pseudo data to participate in the model training of real data through a training method to generate a 360-degree human head model.
[0073] Therefore, the method for reconstructing a 360-degree human head through 3DGS based on a monocular video of this embodiment can generate a 3D reconstruction result of the whole human head from a single-view video, and can cross-drive various motion amplitudes. The method for reconstructing a 360-degree human head through 3DGS based on a monocular video of this embodiment is the first solution that proposes to use a 3DGS solution to reconstruct 360° data using a single-view video.
[0074] Full head reconstruction: To reconstruct a 360° renderable head image from a monocular video with only the front view, a more powerful prior model must be used. We introduced PanoHead as a prior model, which is a 3D-aware GAN model that contains a lot of prior information about the side view and back view of the head. It mainly uses GAN inversion technology, and PTI is used for GAN inversion because it does not rely on additional encoder training and works well in the inversion of PanoHead.
[0075] GAN inversion refers to the process of mapping the target image into the latent space vector of the pre-trained GAN model, and then inputting the vector into the pre-trained generator to complete the image reconstruction. This process enables the pre-trained GAN model to be applied to real image editing and has strong interpretability.
[0076] Pivotal Tuning Inversion (PTI) is a technique for mapping real images to the latent space of StyleGAN, allowing for advanced editing. By fine-tuning the latent space of the generative model, PTI enables editing of real images, maintaining the accuracy of their identity features while reducing distortion.
[0077] Through experiments, we find that PanoHead performs well in generating static FFHQ aligned head images with small facial expression variations. In addition, we can also obtain head images with high-quality frontal rendering after the monocular reconstruction stage.
[0078] Therefore, we use images rendered with neutral expressions and poses as input for our analysis of avatars.
[0079] As an optional implementation of this embodiment, in a method of reconstructing a 360-degree human head through 3DGS based on a monocular video in this embodiment, before the monocular video is used as input and input into the 3D perception GAN model, the method includes:
[0080] The monocular video is input into a low-level vision pre-trained model for preprocessing, which injects more image-specific details into the wild input image of the monocular video to obtain high-quality input data.
[0081] Furthermore, in a method of reconstructing a 360-degree human head through 3DGS based on a monocular video in this embodiment, the low-level visual pre-training model is pre-trained based on a high-quality face dataset.
[0082] Face Restoration: The input images used for inversion are essentially in-the-wild images, which have a large domain gap from the high-quality dataset used for PanoHead training. Directly using rendered images for inversion, the optimized latent codes tend to break away from the GAN manifold, resulting in suboptimal images. The classic approach to address this problem is to apply domain adaptation or transfer learning. However, we can achieve the same goal more concisely. By processing the input images with a low-level vision model pre-trained on a high-quality face dataset, more image-specific details can be injected into the in-the-wild input images. This narrows the domain gap. We call this pre-trained model MR. In our experiments, we use the state-of-the-art face restoration model GFPGAN. It contains a generative prior on the FFHQ dataset that meets our needs.
[0083] As an optional implementation of this embodiment, a method for reconstructing a 360-degree human head through 3DGS based on a monocular video in this embodiment, wherein the 3D perception GAN model uses the Pivotal Tuning Inversion technology to perform GAN inversion;
[0084] Neutral face images of different angles are rendered with arbitrary camera postures, and neutral face images of different angles are filtered out through landmark detection to reconstruct good neutral face images. The filtered neutral face images are used to extend the Pivotal Tuning Inversion technology to multi-view situations.
[0085] Standard PTI takes as input a single image, then searches for latent codes and fine-tunes generator parameters. However, this approach tends to overfit the input image. In PanoHead, the performance of this naive approach drops dramatically in new views. It is worth noting that we can render neutral faces in arbitrary camera poses, and landmark detection allows us to filter out images with well-reconstructed faces. Therefore, we can use the filtered images to extend PTI to the multi-view case.
[0086] Specifically, in a method of reconstructing a 360-degree human head through 3DGS based on a monocular video in this embodiment, the use of a filtered neutral face image to extend the Pivotal Tuning Inversion technology to a multi-view situation includes:
[0087] Search key potential code wp:
[0088]
[0089] M is the number of valid multi-view graphs, L LPIPS is the perceived loss, is the face image after face restoration, G Pano is the frozen full head generator, θ is the generator parameter, and c is the camera pose;
[0090] Adjust the parameters of the full head generator, and the loss setting is as follows
[0091]
[0092] θ* is the optimized parameter, initialized using the pre-trained model θ.
[0093] After adding the side image data and the back image data, we can use this additional data to reconstruct the human head, but training directly on these data will lead to the degradation of the frontal perspective, so we developed a training strategy as follows.
[0094] In a method for reconstructing a 360-degree human head through 3DGS based on a monocular video of this embodiment, the side image data and the back image data are used as pseudo data, and the monocular video is used as real data to input into the 3DGS model for model training to generate a 360-degree human head model, which includes:
[0095] Sampling from real datasets
[0096] pass and c i Rendering
[0097] pass And back propagation calculates the loss L1;
[0098] Update the 3D point cloud {G};
[0099] Where, I = {Ii} video frame sequence;
[0100] Ψ={Ψi} tracked task expression;
[0101] Θ={Θi}, character posture;
[0102] c i It's the camera posture.
[0103] In a method for reconstructing a 360-degree human head through 3DGS based on a monocular video of this embodiment, the side image data and the back image data are used as pseudo data, and the monocular video is used as real data to input into the 3DGS model for model training to generate a 360-degree human head model, which includes:
[0104] Sampling from a pseudo dataset Replace with Ψ i ;
[0105] pass and c j Rendering
[0106] pass And back propagation calculates the loss L2;
[0107] Update the 3D point cloud {G}.
[0108] The present embodiment is a method for reconstructing a 360-degree human head through 3DGS based on a monocular video, cross-training between pseudo data and real data, similar to the training of an adversarial network. Because the Gaussian sphere behind the human head will be cropped during the training process, an additional Gaussian sphere set is initialized on the basis of the existing Gaussian field to model the back of the human head. Although each pseudo data image is generated using a neural network expression, we observe that most tables do not affect the appearance of the side and back. Therefore, when training pseudo data, expression Ψi is used instead of an all-zero vector.
[0109] Experiments show that the training strategy of this embodiment improves the effect on the sides and back of the head.
[0110] Specifically, in this embodiment, a method for reconstructing a 360-degree human head through 3DGS based on a monocular video is used. For example, the training strategy adopted is Figure 2 shown.
[0111] The effect of reconstructing a 360-degree human head through 3DGS based on a monocular video in this embodiment is shown as follows:
[0112] See also Figure 3 and Figure 4 As shown in the figure, the reconstruction effect of the head from the side and back is shown. It can be seen that compared with the best solutions on the market, the solution of this embodiment (the "ours" column in the figure) is better. Both multi-view and additional auxiliary training bring obvious positive benefits.
[0113] See also Figure 5 and Figure 6 As shown, the effects of self-driving and cross-driving are demonstrated. The scheme of this embodiment ( Figure 5 The "ours" column and Figure 6 The rightmost column in the figure all work well in expressing details.
[0114] This embodiment also provides a device for reconstructing a 360-degree human head through 3DGS based on a monocular video, including:
[0115] The GAN inversion module takes the monocular video as input and inputs it into the 3D perception GAN model of the GAN inversion module for GAN inversion to obtain the side image data and back image data of the 360-degree head reconstruction;
[0116] The 3DGS model module uses the side image data and the back image data as pseudo data and the monocular video as real data to input the 3DGS model of the 3DGS model module for model training to generate a 360-degree human head model.
[0117] In this embodiment, a device for reconstructing a 360-degree human head through 3DGS based on a monocular video is used. The GAN inversion module uses a generation method to obtain side image data and back image data, and absorbs them as pseudo data to participate in the model training of the 3DGS model module of real data through training, thereby generating a 360-degree human head model.
[0118] Therefore, the device for reconstructing a 360-degree human head based on a monocular video through 3DGS of this embodiment can generate a 3D reconstruction result of the whole human head from a single-view video, and can cross-drive various motion amplitudes. The method for reconstructing a 360-degree human head based on a monocular video through 3DGS of this embodiment is the first solution that proposes to use a 3DGS solution to reconstruct 360° data using a single-view video.
[0119] Embodiment 2
[0120] See also Figure 7 As shown, this embodiment provides a method for single video 3DGS head reconstruction, including a densification method:
[0121] Based on the video samples, a sparse point cloud containing three-dimensional position information is calculated;
[0122] Initialize the sparse point cloud to get a 3D Gaussian sphere;
[0123] The 3DGS model is trained based on the 3D Gaussian sphere. In the adaptive densification control process of the 3DGS model training, the densification index is introduced into the sampling probability function. The best densification path is automatically updated during the 3DGS model training process, and the 3DGS model is obtained after the training is completed.
[0124] A densification method for single-video 3DGS head reconstruction in this embodiment introduces a densification index into the sampling probability function during the adaptive densification control process of 3DGS model training, and upgrades the process of adding and deleting 3D Gaussian balls to a continuous and differentiable fully automatic process. This can more keenly perceive the changes in data during the training process, replace the processing method that requires manual setting of thresholds and is non-differentiable, and avoid human interference.
[0125] Furthermore, in a densification method for single video 3DGS head reconstruction in this embodiment, the sampling probability function is:
[0126]
[0127] in, is the importance index of a certain 3D Gaussian sphere, and the denominator represents the sum of the importance indexes of all 3D Gaussian spheres of the triangle face where the certain 3D Gaussian sphere is located.
[0128] It can be seen from the above sampling probability function formula that the larger the importance index value of a 3D Gaussian ball, that is, the more important it is, the greater the sampling probability, and vice versa. This rule is consistent with the expectations of the densification process.
[0129] So what kind of importance index works best? As an optional implementation of this embodiment, in a densification method for single-video 3DGS head reconstruction in this embodiment, the importance index of the sampling probability function includes: maximum value index max(s), maximum value ratio index max(s) / min(s), parameter gradient index, transparency index; where s represents position information, that is, the gradient relative to the position coordinates (x, y, z) of the 3D Gaussian sphere.
[0130] Specifically, using the sampling probability functions of the above four different indicators, the 3D head reconstruction results are as follows: Figure 8 As shown in the table, overall, the gradient-based head reconstruction has the best effect, while the transparency-based reconstruction has the worst effect.
[0131] In addition, the effect comparison on the test set is Fig. 9 The densification visualization shown shows the distribution of splash positions during the densification process under different environments, focusing on the reconstruction effects of facial features and hair. It can be seen that the best reconstruction effect is the head reconstruction based on the gradient indicator, and the head reconstruction effect based on the maximum value indicator is the worst.
[0132] like Fig.10 As shown in the figure, the difference in effect between the "gradient-based densification strategy" and the "gradient-based densification strategy" can be observed from four angles: clarity, edge differentiation, whether hair is differentiated, and the distribution density of 3D Gaussian balls. In high-frequency areas, such as hair, eyes, nose, mouth, and eyebrows, after adding gradient-based probability sampling densification on the left, the Gaussian balls in these areas are obviously denser, which is in line with our expectations. More Gaussian balls are needed in key areas, while a small number of Gaussian balls can be used to represent smooth areas.
[0133] From the above comparison results, we can see that the 3DGS head reconstruction effect is better using the automatic continuously differentiable densification scheme.
[0134] It should be noted that this embodiment exemplifies the above four different importance indicators. Of course, the importance indicators in the sampling probability function include but are not limited to the above four indicators.
[0135] As an optional implementation of this embodiment, in a densification method for single video 3DGS head reconstruction of this embodiment, the sampling probability function adds different weights to different designated important areas of head reconstruction. In this way, the weight of the sampling probability function can be set based on each key area in the head reconstruction, and more attention is paid to the key areas in the 3D Gaussian sphere densification project of 3DGS, so that the densification effect is better and meets the needs of head reconstruction.
[0136] Specifically, in a single video 3DGS head reconstruction densification method of this embodiment, the 3DGS model training based on the 3D Gaussian sphere includes:
[0137] For every N iterations of the 3D Gaussian sphere, the order of the spherical harmonic coefficients is increased, where N is the set value;
[0138] Randomly select a camera perspective;
[0139] Render the image, obtaining the view space point, visibility filter, and radius information;
[0140] Calculate the loss L = (1-λ)L1 + λL D-SSIM (L1 loss and L D-SSIM The weighted sum of the losses, L1 loss helps ensure pixel-level accuracy, and L D-SSIM The loss helps to maintain the overall structure and visual quality of the image), λ is the preset weight value, and back propagation is performed;
[0141] Subsequent operations via gradient-free contextual learning include:
[0142] Perform point cloud density operations based on the number of iterations:
[0143] Update the maximum radius information;
[0144] Increase and trim point cloud density based on conditions;
[0145] Update the optimizer parameters;
[0146] During the point cloud density operation, the densification index is introduced into the sampling probability function, which is:
[0147]
[0148] in, is the importance index of a certain 3D Gaussian sphere, and the denominator represents the sum of the importance indexes of all 3D Gaussian spheres of the triangle face where the certain 3D Gaussian sphere is located.
[0149] A densification method for single video 3DGS head reconstruction in this embodiment includes:
[0150] Shoot an image or video of the target person, input the image or video into the 3DGS model, and generate a 3D reconstruction model of the target person.
[0151] Embodiment 3
[0152] See also Fig.11 As shown, this embodiment provides a method for improving single video 3DGS head reconstruction, including a Gaussian sphere update method:
[0153] Based on the captured video, a sparse point cloud containing three-dimensional position information is calculated;
[0154] Initialize the sparse point cloud to get a 3D Gaussian sphere, and reconstruct the 3DGS head based on the 3D Gaussian sphere;
[0155] Initialize the 3D Gaussian sphere position to obtain the human head point cloud;
[0156] A new 3D Gaussian sphere is expanded outside the human head point cloud. Through a predefined loop scheduling program, the offset δ of the new 3D Gaussian sphere is set based on the 3D Gaussian sphere position of the human head point cloud, and a shell-like structure is constructed to obtain the human head hair point cloud.
[0157] During 3DGS training, the 3D Gaussian sphere will continue to grow or shrink. When the task of head reconstruction is heavy, we will put the 3D Gaussian sphere on the head model during initialization, such as the surface of 3DMM. We then hope that the 3D Gaussian sphere can be generated outside the head to reconstruct the hair, but we do not want the Gaussian sphere to grow inside the head, because the latter is meaningless.
[0158] The present embodiment provides a Gaussian sphere updating method for improving the efficiency of single-video 3DGS head reconstruction, which is to solve the problem of effectively generating a Gaussian sphere in a directional outward direction, expand a new 3D Gaussian sphere outside the point cloud of the human head skull, and set the offset δ of the new 3D Gaussian sphere based on the 3D Gaussian sphere position of the human head skull point cloud through a predefined loop scheduling program to construct a shell-like structure and obtain a point cloud of human head hair. Therefore, the present embodiment provides a Gaussian sphere updating method for improving the efficiency of single-video 3DGS head reconstruction, which achieves the effect of spreading the 3D Gaussian sphere outward along the hair, and does not waste computing power in areas without hair.
[0159] Furthermore, in a Gaussian sphere updating method for improving single-video 3DGS head reconstruction of this embodiment, a new 3D Gaussian sphere is expanded outside the head point cloud, and an offset δ of the new 3D Gaussian sphere is set based on the 3D Gaussian sphere position of the head point cloud through a predefined loop scheduling program to construct a shell-like structure, and the head hair point cloud is obtained, including:
[0160] p(i) represents the position j of the new 3D Gaussian sphere on the i-th triangle patch, and δ(i) is the offset assigned to the new 3D Gaussian sphere. The growth formula of the new 3D Gaussian sphere is:
[0161]
[0162] in, is the 3D Gaussian ball growth formula of the human head point cloud, n(i) is the expansion direction parameter from the i-th triangle patch, It is the sampling point Coordinates on the i-th triangle patch; are the coordinates of three points on the i-th triangle patch.
[0163] Since the 3D Gaussian spheres are pruned and densified during the 3DGS model training process, we can eventually form a shell-like structure in a progressive and unsupervised manner, which can effectively capture fine-scale details.
[0164] As an optional implementation of this embodiment, in a Gaussian sphere update method for improving single-video 3DGS head reconstruction of this embodiment, the n(i) is the normal vector from the vertex of the triangle patch whose outward expansion direction parameter is the i-th triangle patch.
[0165] As an optional implementation of this embodiment, in a Gaussian sphere update method for improving single-video 3DGS head reconstruction of this embodiment, the n(i) is the midline of the triangle patch whose outward expansion direction parameter is the i-th triangle patch.
[0166] As an optional implementation of this embodiment, in a Gaussian sphere update method for improving single-video 3DGS head reconstruction of this embodiment, the n(i) is a perpendicular line from the i-th triangle patch whose outward expansion direction parameter is a triangle patch.
[0167] Specifically, in a Gaussian sphere updating method for improving single-video 3DGS head reconstruction of this embodiment, the offset δ of setting a new 3D Gaussian sphere based on the 3D Gaussian sphere position of the head point cloud includes:
[0168] The value range of the preset offset δ is [0, δmax], where δmax is the preset maximum offset;
[0169] Based on the 3D Gaussian sphere position of the human head point cloud, the new 3D Gaussian sphere offset δ is set in the value range [0, δmax].
[0170] This embodiment ensures that the 3D Gaussian sphere is expanded within the numerical range [0, δmax] by presetting the maximum offset δmax, constructs a shell-like structure, obtains the point cloud of the human head hair, and avoids wasting computing power in areas without hair.
[0171] As an optional implementation of this embodiment, a Gaussian sphere update method for improving the efficiency of single-video 3DGS head reconstruction in this embodiment, initializing the 3D Gaussian sphere position to obtain a human head skull point cloud includes: initializing the 3D Gaussian sphere position and placing it on the surface of the 3DMM head model to obtain a human head skull point cloud.
[0172] Experiments on eight open-source human heads on the market show that the new 3D Gaussian sphere growth formula of this embodiment can reduce the 3D Gaussian sphere calculation by half compared with the random point scattering method.
[0173] In terms of effect, see Fig.12 As shown in the comparison table, the "w / oδ" column is the experimental result data of the new 3D Gaussian sphere growth formula of this embodiment. The accurate calculation of the 3D Gaussian sphere position can improve the reconstruction effect. For specific visualization effects, see Fig.13 shown.
[0174] Embodiment 4
[0175] The following describes an electronic device embodiment of the present invention, which can be regarded as a specific physical implementation of the method and device embodiments of the present invention. The details described in the electronic device embodiment of the present invention should be regarded as a supplement to the above method or device embodiments; details not disclosed in the electronic device embodiment of the present invention can be implemented with reference to the above method or device embodiments.
[0176] Fig.14 It is a structural schematic diagram of an electronic device of an embodiment of the present invention, the electronic device includes a processor and a memory, the memory is used to store a computer executable program, when the computer program is executed by the processor, the processor executes a single video 3DGS head reconstruction method of embodiment one, or two, or three.
[0177] like Fig.14 As shown, the electronic device is presented in the form of a general computing device. The processor may be one or more and work in coordination. The present invention does not exclude distributed processing, that is, the processor may be dispersed in different physical devices. The electronic device of the present invention is not limited to a single entity, but may also be the sum of multiple physical devices.
[0178] The memory stores a computer executable program, which is usually a machine-readable code. The computer-readable program can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least part of the steps in the method.
[0179] The memory includes a volatile memory, such as a random access memory unit (RAM) and / or a cache memory unit, and may also be a non-volatile memory, such as a read-only memory unit (ROM).
[0180] Optionally, in this embodiment, the electronic device further includes an I / O interface, which is used for the electronic device to exchange data with an external device. The I / O interface can represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.
[0181] It should be understood that Fig.14 The electronic device shown is only an example of the present invention, and the electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as display screens, and some electronic devices also include human-computer interaction elements such as buttons, keyboards, etc. As long as the electronic device can execute the computer-readable program in the memory to implement the method of the present invention or at least part of the steps of the method, it can be considered as an electronic device covered by the present invention.
[0182] Fig.15 Schematic diagram of a computer readable recording medium according to an embodiment of the present invention. Fig.15 As shown, a computer executable program is stored in a computer-readable recording medium, and when the computer executable program is executed, a single-video 3DGS head reconstruction method of Embodiment 1, or 2, or 3 of the present invention is implemented. The computer-readable recording medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable recording medium may also be any readable medium other than a readable recording medium, which may send, propagate, or transmit a program for use by or in combination with an instruction execution system, device, or device. The program code contained on the readable recording medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0183] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0184] Through the above description of the implementation mode, it is easy for those skilled in the art to understand that the present invention can be implemented by hardware capable of executing a specific computer program, such as the system of the present invention, and the electronic processing unit, server, client, mobile phone, control unit, processor, etc. contained in the system. The present invention can also be implemented by computer software that executes the method of the present invention, such as control software executed by a microprocessor, an electronic control unit, a client, a server, etc. However, it should be noted that the computer software that executes the method of the present invention is not limited to being executed by one or a specific hardware entity, and it can also be implemented in a distributed manner by unspecified specific hardware. For computer software, the software product can be stored in a computer-readable recording medium (which can be a CD-ROM, a USB flash drive, a mobile disk, etc.), and can also be distributed and stored on the network, as long as it enables the electronic device to execute the method according to the present invention.
[0185] The above embodiments are only used to illustrate the present invention and are not intended to limit the technical solutions described in the present invention. Although the present invention has been described in detail with reference to the above embodiments, the present invention is not limited to the above specific implementation methods. Therefore, any modification or equivalent replacement of the present invention; and all technical solutions and improvements thereof that do not depart from the spirit and scope of the invention are included in the scope of the claims of the present invention.
Claims
1. A method for reconstructing a 360-degree human head based on a monocular video through 3DGS, characterized in that: include: The monocular video is used as input to the 3D perception GAN model for GAN inversion to obtain the side image data and back image data of 360-degree head reconstruction; The side image data and back image data are used as pseudo data, and the monocular video is used as real data to input into the 3DGS model for model training to generate a 360-degree human head model.
2. The method for reconstructing a 360-degree human head through 3DGS based on a monocular video according to claim 1, characterized in that: Before taking the monocular video as input and feeding it into the 3D perception GAN model, it includes: The monocular video is input into a low-level vision pre-trained model for preprocessing, which injects more image-specific details into the wild input image of the monocular video to obtain high-quality input data.
3. The method for reconstructing a 360-degree human head through 3DGS based on a monocular video according to claim 2, characterized in that: The low-level visual pre-training model is pre-trained based on a high-quality face dataset.
4. The method for reconstructing a 360-degree human head through 3DGS based on a monocular video according to claim 2, characterized in that: The 3D perception GAN model uses the Pivotal Tuning Inversion technique to perform GAN inversion; Neutral face images of different angles are rendered with arbitrary camera postures, and neutral face images of different angles are filtered out through landmark detection to reconstruct good neutral face images. The filtered neutral face images are used to extend the Pivotal Tuning Inversion technology to multi-view situations.
5. The method for reconstructing a 360-degree human head through 3DGS based on a monocular video according to claim 4, characterized in that: The extension of the Pivotal Tuning Inversion technique to a multi-view situation using a filtered neutral face image includes: Search key potential code wp: M is the number of valid multi-view graphs, L LPIPS is the perceived loss, is the face image after face restoration, G Pano is the frozen full head generator, θ is the generator parameter, and c is the camera pose; Adjust the parameters of the full head generator, and the loss setting is as follows θ* is the optimized parameter, initialized using the pre-trained model θ.
6. The method for reconstructing a 360-degree human head through 3DGS based on a monocular video according to claim 5, characterized in that: The side image data and the back image data are used as pseudo data, and the monocular video is used as real data to input into the 3DGS model for model training to generate a 360-degree head model, which includes: Sampling from real datasets pass and c i Rendering pass And back propagation calculates the loss L1; Update the 3D point cloud {G}; Where, I = {Ii} video frame sequence; Ψ={Ψi} tracked task expression; Θ={Θi}, character posture; c i It's the camera posture.
7. The method for reconstructing a 360-degree human head through 3DGS based on a monocular video according to claim 6, characterized in that: The side image data and the back image data are used as pseudo data, and the monocular video is used as real data to input into the 3DGS model for model training to generate a 360-degree head model, which includes: Sampling from a pseudo dataset Replace with Ψ i ; pass and c j Rendering pass And back propagation calculates the loss L2; Update the 3D point cloud {G}.
8. A device for reconstructing a 360-degree human head based on a monocular video through 3DGS, characterized in that: include: The GAN inversion module takes the monocular video as input and inputs it into the 3D perception GAN model of the GAN inversion module for GAN inversion to obtain the side image data and back image data of the 360-degree head reconstruction; The 3DGS model module uses the side image data and the back image data as pseudo data and the monocular video as real data to input the 3DGS model of the 3DGS model module for model training to generate a 360-degree human head model.
9. An electronic device comprising a processor and a memory, wherein the memory is used to store a computer executable program, characterized in that: When the computer program is executed by the processor, the processor executes a method for reconstructing a 360-degree human head through 3DGS based on a monocular video as described in any one of claims 1 to 7.
10. A computer-readable recording medium storing a computer-executable program, characterized in that: When the computer executable program is executed, a method for reconstructing a 360-degree human head through 3DGS based on a monocular video as described in any one of claims 1 to 7 is implemented.