Head reconstruction method based on Gaussian sputtering
By optimizing the 3D head model using Gaussian sputtering and convolutional neural networks, the problems of low head reconstruction quality, complex input, and high acquisition cost in existing technologies are solved, achieving high-fidelity and real-time rendering under sparse input.
Patent Information
- Application Number
- CN202511109146.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-12-05
AI Technical Summary
Existing head reconstruction technologies suffer from low quality, complex input, high acquisition costs, and lack of generalization, especially in cases of sparse input where it is difficult to achieve high fidelity and real-time rendering.
A Gaussian sputtering-based method is adopted to align and clone the initial point cloud model using supervised images. By combining convolutional neural networks and multilayer perceptrons, the Gaussian model of the 3D head is optimized. Feature extraction and rendering are performed using sparse input images to achieve high-fidelity 3D head reconstruction.
It achieves high-quality head reconstruction under sparse input, with generalization and real-time rendering capabilities, and can handle more facial details while reducing acquisition costs.
Smart Images

Figure CN121074243A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision, specifically to a head reconstruction method based on Gaussian sputtering. Background Technology
[0002] 3D reconstruction has always been an important research topic in the field of computer vision, and head reconstruction, as a branch of it, remains a challenging problem. Despite significant technological advancements, reconstructing a high-fidelity 3D head model solely from sparse images remains extremely difficult.
[0003] In recent years, many excellent results have emerged, and significant progress has been made in sparse input head reconstruction. Among these, meshes and point clouds are the most common 3D representations of heads because they are explicit and well-suited for fast CUDA-based rasterization. However, they also have some drawbacks. First, they typically require a large amount of memory to store vertex coordinates and connectivity information. Second, because points are discrete, they cannot capture continuous shape or surface features well, especially when smooth or subtle changes are needed. Furthermore, for scenarios with frequent modifications or dynamic updates, their representation methods may not be flexible enough, making them difficult to extend or modify.
[0004] Subsequently, more structured mesh structures for the head were developed, such as in reference 1: Deng Y, Yang J, Xu S, et al. Accurate 3D Face Reconstruction With Weakly-Supervised Learning: From Single Image to Image Set [C] / / 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW).IEEE,2020.DOI:10.1109 / CVPRW.2019.00038, which describes a more structured mesh structure for the head. This mesh structure consists of vertices and faces and has strong prior influence, such as the flame head model (i.e., the FLAME average head model). The FLAME average head model consists of 5023 vertices and approximately 10000 faces, encompassing a complete head and neck. These models only require obtaining the corresponding facial expressions, shape, and other coefficients from the image to fit a relatively accurate head. Due to the complexity of human facial structure, relying solely on these coefficients for fitting is far from sufficient. During the fitting process, the fitting results may have errors compared to the real head. Secondly, there will be occlusion issues on the side of the face, and the treatment of hair is also an important problem.
[0005] With the development of deep learning and neural rendering, the Neural Radiance Field (NeRF) method has begun to emerge. For example, in reference 2: Mildenhall B, Srinivasan PP, Tancik M, et al. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis[C] / / 2020.DOI:10.48550 / arXiv.2003.08934, a NeRF method is mentioned. This method is based on continuous scene representation. It expresses the scene as a continuous latent function through a fully connected multilayer perceptron (MLP). During the training phase, it learns how to map any three-dimensional coordinates to the corresponding color and density values to form a continuous radiation field. In particular, the effect is even more amazing when combined with three-plane position encoding.
[0006] Reference 3: Chan ER, Lin CZ, Chan MA, et al. Efficient Geometry-aware 3D Generative Adversarial Networks[J]. 2021. DOI:10.48550 / arXiv.2112.07945 mentions an adversarial generative model (GAN, specifically named EG3D), which requires a large number of viewpoints for training. Not only does it have high data requirements and cannot perform high-fidelity head modeling, but it also has a very slow rendering speed and cannot be accelerated by CUDA. It is not suitable for completing such complex head modeling and rendering under sparse input conditions.
[0007] Reference 4: Kerbl B, Kopanas G, Leimkühler T, et al. 3d gaussian splatting for real-time radiance field rendering[J]. ACM Transactions on Graphics, 2023, 42(4): 1-14 involves Gaussian splatting. Gaussian splatting is a scene representation model based on Gaussian distribution. It uses a large number of Gaussian spheres to represent the position of points. Each Gaussian sphere has four attributes: color c, opacity α, and covariance matrix Σ. The covariance matrix Σ can be decomposed into rotation r and scale s.
[0008] Most existing methods rely on video input or simultaneous acquisition from a large number of cameras for head reconstruction, such as in reference 5: Wang J, Xie JC, Li X, et al. Gaussian Head: Impressive Head Avatars with Learnable Gaussian Diffusion[J]. arXiv preprint arXiv:2312.01632,2023. These methods require a large amount of data input, have high acquisition costs, and are redundant in terms of available information, making them overly idealistic for practical applications.
[0009] This invention utilizes sparse input to achieve head reconstruction, balancing the issues of redundant video input information and missing information in a single image. Furthermore, thanks to the superiority of Gaussian sputtering representation, it exhibits good generalization ability and supports CUDA acceleration, enabling real-time rendering (above 24 FPS) and producing high-quality head reconstruction results. Summary of the Invention
[0010] To address the problems of low quality, complex input, high acquisition cost, and lack of generalization in existing head reconstruction techniques, this invention proposes a head reconstruction method based on Gaussian sputtering, belonging to the field of computer vision. The method involves aligning and cloning an initial point cloud model using a supervised image to obtain a precise head point cloud from the supervised image, with the precise head point cloud positions used as Gaussian point cloud positions. The Gaussian point cloud is projected onto the image coordinate system, and the supervised image is used as input to a convolutional neural network for feature extraction to obtain a convolutional feature map. Feature vectors are obtained by sampling the precise head point cloud and the convolutional feature map in the image coordinate system. These feature vectors are then decoded to obtain Gaussian point cloud parameters, thus obtaining a 3D head Gaussian model. The 3D head Gaussian model is then used to quickly render a head image from any viewpoint. Pixel-by-pixel loss calculations are performed on the rendered image and the real image to optimize the 3D head Gaussian model. This invention, under sparse input conditions, fits a convolutional neural network and Gaussian sputtering to obtain a high-fidelity 3D head Gaussian model that can be rendered in real time and has generalization capabilities.
[0011] A head reconstruction method based on Gaussian sputtering, such as Figure 1 and Figure 2 As shown, the technical solution specifically includes the following steps:
[0012] S1, Determine the training and test sets: Obtain the dataset through adversarial generative networks. The dataset consists of 1,000 to 2,000 head images of different identities. The dataset includes a training set and a test set, with a ratio of 4:1.
[0013] S2, using the FLMAE average head model as the initial point cloud model of 3D Gaussian, the initial point cloud model is aligned and cloned and split by the supervised image to obtain the position of Gaussian point cloud in 3D Gaussian in world coordinate system.
[0014] S3 projects the Gaussian point cloud positions from the world coordinate system to the image coordinate system; simultaneously, it extracts features through a convolutional neural network to obtain a convolutional feature map.
[0015] S4, sample the Gaussian point cloud positions and convolutional feature maps projected onto the image coordinate system to obtain feature vectors; the feature vectors are decoded by the decoder to obtain complete Gaussian point cloud parameters, thereby obtaining a three-dimensional head Gaussian model;
[0016] S5 updates parameters using a loss function to optimize the 3D head Gaussian model.
[0017] Furthermore, in step S2, the process of obtaining the position of the Gaussian point cloud in the 3D Gaussian coordinate system is as follows:
[0018] S21, use the FLMAE average head model as the initial point cloud model of the 3D Gaussian; there are 120 viewpoint images for the head of each identity in the training set. Select 2-4 viewpoint images of one identity in the training set as the supervision images of the initial point cloud model. The 2-4 viewpoint images are also the sparse input images of the convolutional neural network. The remaining viewpoint images are used as training viewpoint images.
[0019] S22, the initial point cloud model is aligned and cloned using the supervised image to obtain the precise point cloud position of the head corresponding to the supervised image; cloned splitting is to create points and move them towards the position gradient of the points, and remove redundant and less influential points; alignment and cloned splitting are regularization processes.
[0020] S23, take the precise point cloud position of the head as the position of the 3D Gaussian point cloud in the world coordinate system;
[0021] The pixel differences between images obtained by rendering the supervised image using Gaussian point clouds are used as the loss function. align Loss function align This includes absolute value L1 loss and structural similarity SSIM loss;
[0022] By calculating the loss function Loss align Make the shape of the Gaussian sphere match the head of the corresponding identity, and finally obtain the position of the Gaussian point cloud in the world coordinate system;
[0023] Loss function align The calculation process is as follows:
[0024] Loss align=a×Loss l1 +b×Loss ssim
[0025] Among them, Loss l1 The L1 loss is the absolute value of the difference between the Gaussian point cloud locations rendered to the supervised image during the initial point cloud model optimization process, representing the absolute difference between the two supervised images; Loss ssim Structural similarity (SSIM) between two supervised images when the absolute value L1 loss is the same. SSIM represents the degree of similarity between the two supervised images, which conforms to the intuitive perception of the human eye; a and b are the loss values, respectively. l1 and Loss ssim The proportion it accounts for;
[0026] Loss l1 =|I pr -I gt |
[0027]
[0028] Among them, I pr This represents the image rendered by the predicted Gaussian sphere; I gt Represents the ground truth image before Gaussian sphere rendering; μ pr Indicates in I pr Data in the image window; μ gt Indicates in I gt Data in the image window; σ pr_gt Indicate I pr and I gt The covariance between them; C1 and C2 are used to calculate the loss. ssim Constants in the process.
[0029] Furthermore, in step S3, the process of obtaining the convolutional feature map through feature extraction using a convolutional neural network is as follows:
[0030] The convolutional neural network selects the feature extraction network feature_net;
[0031] S31, The convolutional neural network receives a three-channel input image; the convolutional neural network includes 19 convolutional layers, and each convolutional layer connects and fuses features through residual blocks;
[0032] First, the sparse input image undergoes preliminary feature extraction through the first convolutional layer;
[0033] Then, feature channels are transformed and size adjusted sequentially through three stages of residual blocks, finally outputting three feature maps of different scales; each stage contains 3 residual blocks, and each residual block includes 2 convolutional layers;
[0034] S32: Input the supervised image into the convolutional neural network to obtain the convolutional feature map corresponding to the sparse input image.
[0035] Furthermore, in step S4, the process of obtaining the three-dimensional Gaussian head model is as follows:
[0036] S41, Obtain the fused feature vector of the Gaussian point cloud position of the head:
[0037] The Gaussian point cloud positions in the image coordinate system are sampled by bilinear interpolation on the convolutional feature map to obtain the feature vectors corresponding to the Gaussian point cloud positions in the image coordinate system; the feature vectors of the Gaussian point cloud positions in the sparse input image are merged to obtain the fused feature vector of the Gaussian point cloud positions in the head.
[0038] S42, setting up the decoding network (decoder):
[0039] The decoding network is a multilayer perceptron (MLP) model used for decoding features; the multilayer perceptron model includes a preprocessor, a decoder, and an output layer;
[0040] The processor includes three identical MLP_pre modules; the decoder includes three stages of ResidualBlock1D modules; the output layer includes three sub-modules: covariance, opacity, and RGB.
[0041] S43, decodes the fused feature vector through a decoding network:
[0042] The fused feature vector obtained in S41 is input into the decoding network. After processing, the fused feature vector is used to obtain feature vectors feat1, feat2, and feat3. Feature vectors feat3, feat2, and feat1 are then processed by the feature fusion and upsampling of the decoder, and then passed through the three sub-modules of the output layer to obtain the covariance, opacity, and RGB values, respectively. Combined with the Gaussian point cloud positions obtained in S23, a preliminary 3D Gaussian head model is obtained.
[0043] S44, the multilayer perceptron in the decoding network decodes the three parameters of covariance, opacity and RGB value to obtain the attributes of the head Gaussian sphere; the attributes of the head Gaussian sphere include color c, opacity α and covariance matrix Σ; finally, a three-dimensional head Gaussian model corresponding to the sparse input image of the convolutional neural network is generated.
[0044] Furthermore, in step S5, the optimization process of the three-dimensional head Gaussian model is as follows:
[0045] S51, given the camera pose and camera parameters of the rendering viewpoint, rasterize the obtained 3D head Gaussian model to obtain the head rendering image of any viewpoint.
[0046] S52, the pixel difference between the predicted image and the training viewpoint image is the loss function. net ;
[0047] Loss function net Including absolute value L1 loss, structural similarity SSIM loss, and VGG perceptual loss;
[0048] Loss function net This serves as the supervision term for both the convolutional neural network and the decoding network; Adam is chosen as the optimizer to adjust the weights and biases of the convolutional neural network to minimize the loss function. net ;
[0049] First, the loss function Loss is calculated using the backpropagation algorithm. net For each weight and bias, the gradient;
[0050] Then, the parameters are updated using gradient descent, the training data is input into the convolutional neural network, and the predicted value is calculated through forward propagation.
[0051] Then, the parameters are updated through backpropagation, and the process is repeated until the convolutional neural network and the decoding network converge. The number of iterations is 10 cycles, and each cycle consists of training all data in the training set once. After training, the trained convolutional neural network and the decoding network are obtained.
[0052] Loss function net The calculation is as follows:
[0053] Loss net =α×Loss align +β×Loss vgg
[0054] Among them, Loss vgg This represents the VGG perceptual loss between the predicted and ground truth images. The VGG perceptual loss is a perceptual loss function for the pre-trained VGG network. It minimizes the differences between the feature maps of the generated and target images in the VGG network layers, ensuring that the content information of the generated image does not deviate too much. The VGG network layers include conv1_1, conv2_1, conv3_1, conv4_1, and conv5_1 in VGG19. The differences between feature maps include the mean squared error; α and β are the loss values, respectively. align and Loss vgg The proportion of the loss function:
[0055] Loss vgg =||vgg(I pr )-vgg(I gt )|| 2
[0056] Among them, vgg(I pr ) represents the feature values of the input VGG pre-trained network layer to the rendered image of the 3D head Gaussian model; vgg(I gt ) represents the feature values of the 3D head Gaussian model image excluding the supervised image.
[0057] Compared with existing technologies, the advantages and positive effects of this invention are as follows:
[0058] Compared to existing model-based (3DMM, FLAME) head reconstruction methods, this invention can handle more facial details, such as hair or profile, and obtain higher quality images.
[0059] Compared to Neural Radiation Field (NeRF), this invention not only guarantees quality but also allows for real-time rendering like traditional rendering pipelines; this invention is generalizable and can model heads of different identities.
[0060] Compared to existing image-based head reconstruction techniques, this invention requires only a small number of images to achieve accurate reconstruction results, and it features generalization and fast rendering. Attached Figure Description
[0061] Figure 1 This is an overall flowchart of the present invention;
[0062] Figure 2 This is a schematic diagram of the forward propagation of the three-dimensional Gaussian head model of the present invention;
[0063] Figure 3 This is a schematic diagram of the FLAME average head model alignment process of the present invention.
[0064] Figure 4 This is a new viewpoint image rendered after the Gaussian point cloud position and attributes have been improved according to the present invention;
[0065] Figure 4 In the image, (a) is the rendered top view; (b) is the rendered right view; (c) is the rendered left view; and (d) is the rendered front view. Detailed Implementation
[0066] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0067] Using 1500 data sets as the dataset, with 1200 sets as the training set and 300 sets as the test set, the implementation process is as follows:
[0068] FLAME average head model alignment based on sparse input with Gaussian sputtering:
[0069] Step 2.1, the alignment process iterates a total of 9000 times. The initial FLAME average head model is a smooth head, lacking information such as the position of head details and the correct color.
[0070] Step 2.2: In the first 5000 iterations, the positions of the original points are updated iteratively, and clone splitting is performed. At the same time, unwanted points are pruned to obtain dense points, and other Gaussian properties are also updated iteratively.
[0071] Step 2.3: When the iteration count reaches 5000, delete the index corresponding to the original FLAME of the current point. The reason for deleting the original point is that the original point has not changed during the update process and is redundant in the entire Gaussian sphere set. Repeat step 2.2 4000 times to obtain the accurate and dense Gaussian sphere positions. Since there are few supervised viewpoints, only the Gaussian sphere positions (x, y, z) are reliable; other Gaussian parameters are inaccurate. This invention only requires accurate x, y, z, such as... Figure 3 As shown, Figure 3 Point cloud image after aligning the average head model of FLAME;
[0072] Step 2.4: During the iteration process, a loss function is established by rendering the 3D head Gaussian model onto the supervised image and comparing it with the ground truth value of the supervised image. The loss function used is as follows:
[0073] Loss align =a×Loss l1 +b×Loss ssim
[0074] Among them, Loss l1 The L1 loss is the absolute value of the difference between the Gaussian point cloud locations rendered to the supervised image during the initial point cloud model optimization process, representing the absolute difference between the two supervised images; Loss ssim Structural similarity (SSIM) between two supervised images when the absolute value L1 loss is the same. SSIM represents the degree of similarity between the two supervised images, which conforms to the intuitive perception of the human eye; a and b are the loss values, respectively. l1 and Loss ssim The proportions; a takes a value of 0.2, b takes a value of 0.8;
[0075] Loss l1 =|I pr -I gt |
[0076]
[0077] Among them, I pr This represents the image rendered by the predicted Gaussian sphere; Igt Represents the ground truth image before Gaussian sphere rendering; μ pr Indicates in I pr Data in the image window; μ gt Indicates in I gt Data in the image window; σ pr_gt Indicate I pr and I gt The covariance between them; C1 and C2 are used to calculate the loss. ssim The constants in the process, where C1 is 6.5025 and C2 is 58.5225.
[0078] Feature extraction using convolutional neural network feature_net:
[0079] Step 3.1: The convolutional neural network feature_net is used for feature extraction, including an encoder and a decoder. The encoder uses a ResNet model as a feature extractor to extract features according to different scales of the input image. The ResNet model includes ResNet18 and ResNet34, which determine the number of feature channels in different layers. The decoder restores the features of the encoder through upsampling and convolution operations. The ResNet model is used to extract features from the input viewpoint image, obtaining a feature vector for each pixel position. The feature vector contains higher-dimensional information. The input image size is 4×3×512×512, and the extracted feature vector is 4×32×512×512. The feature vector is concatenated with the input image to form a high-dimensional feature vector of size 4×35×512×512.
[0080] Step 3.2: Project the aligned Gaussian spheres onto the input viewpoint and sample the feature vectors obtained in step 3.1 to obtain the features of all points of size N×35, where N represents the number of Gaussian spheres in space.
[0081] Step 3.3: Input the N×35 feature values into the decoder. The output is color c, opacity α, rotation r and scale s in the covariance Σ. The predicted values are N×3, N×1, N×1×4 and N×1×3, respectively. The decoder is a multilayer perceptron decoder used to input features and generate the predicted results. The decoder contains multiple fully connected layers and activation functions. During forward propagation, different fully connected layers process the different input features to calculate the Gaussian parameters.
[0082] Step 3.4: During the training process on the training set, the Gaussian head is rendered onto 116 viewpoints (excluding the input viewpoint) using the Gaussian parameters obtained in step 3.3, and supervised by ground truth. The loss function used is Loss. net for:
[0083] Loss net =α×Loss align +β×Loss vgg
[0084] Where α takes the value of 1 and β takes the value of 0.01.
[0085] Loss vgg =||vgg(I pr )-vgg(I gt )|| 2
[0086] Among them, vgg(I pr ) represents the feature value of the input VGG pre-trained layer for the predicted image, vgg(I) gt ) represents the eigenvalue corresponding to the true value.
[0087] Combining and rendering a 3D Gaussian head model:
[0088] Step 5.1: Iteratively update the 3D head Gaussian model on the training set to obtain the optimized 3D head Gaussian model.
[0089] Step 5.2: In the test set, using the trained 3D head Gaussian model, after obtaining the accurate Gaussian positions x, y, z and parameters c, α, r, and s as described in steps 5.1 and 5.2 respectively, render the result as shown below. Figure 4 The images shown are from four viewpoints. In addition to using c to represent color information, spherical coefficients are also used for color conversion, resulting in better head reconstruction.
[0090] During the experiment, the equipment used was an NVIDIA TITAN RTX with 24GB of video memory, running on a 64-bit Linux system.
[0091] This invention uses the same device to capture sparse (four viewpoints in this invention) head images from different directions or four devices from different locations. Using deep learning, it starts with the FLAME average head model and aligns it with the input sparse viewpoints to obtain accurate and dense Gaussian sphere positions (x, y, z). Then, by projecting spatial points onto the input viewpoints, it samples feature information extracted by a feature extraction network and decodes it using an MLP to obtain other parameter properties of the Gaussian sphere. Based on the obtained complete 3D Gaussian head model, rendering results at any position can be obtained. Compared to existing reconstruction methods, this invention not only uses less input but also guarantees realistic reconstruction quality. Furthermore, it enables real-time rendering, allowing for richer applications of 3D head reconstruction.
[0092] The above description is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any person skilled in the art may use the above-disclosed technical content to make changes or modifications to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method of head reconstruction based on Gaussian sputtering, characterized in that, The method comprises the following steps: S1, determining a training set and a test set: obtaining a data set through a generative adversarial network, the data set being 1000-2000 head images of different identities, the data set comprising the training set and the test set, and the ratio of the training set to the test set being 4:1; S2, taking the FLMAE average head model as an initial point cloud model of a 3D Gaussian, aligning and splitting the initial point cloud model through a supervised image to obtain a Gaussian point cloud position in the 3D Gaussian in a world coordinate system; S3, projecting the Gaussian point cloud position from the world coordinate system to an image coordinate system; meanwhile, performing feature extraction through a convolutional neural network to obtain a convolutional feature map; S4, sampling the Gaussian point cloud position projected to the image coordinate system and the convolutional feature map to obtain a feature vector; decoding the feature vector through a decoder to obtain complete Gaussian point cloud parameters, thereby obtaining a three-dimensional head Gaussian model; S5, updating parameters through a loss function to optimize the three-dimensional head Gaussian model.
2. The head reconstruction method based on Gaussian sputtering according to claim 1, characterized in that, In step S2, the Gaussian point cloud position in the 3D Gaussian in the world coordinate system is obtained in the following manner: S21, taking the FLMAE average head model as an initial point cloud model of a 3D Gaussian; there are 120 view point images of the head of each identity in the training set, 2-4 view point images of an identity in the training set are selected as supervised images of the initial point cloud model, the 2-4 view point images are also sparse input images of the convolutional neural network, and the remaining view point images are training view point images; S22, aligning and splitting the initial point cloud model through the supervised images to obtain accurate head point cloud positions corresponding to the supervised images; performing dense point and deleting redundant points through the splitting to create points and move them towards the position gradient of the points, and remove redundant points and points with less influence; the alignment and splitting are a regularization process; S23, taking the accurate head point cloud positions as the Gaussian point cloud positions in the 3D Gaussian in the world coordinate system; The pixel difference between the supervised image and the image rendered by the Gaussian point cloud is the loss function Loss align ; Loss function Loss align including an absolute value L1 loss and a structural similarity SSIM loss; By calculating the loss function Loss align Make the shape of the Gaussian ball consistent with the head of the corresponding identity, and finally get the position of the Gaussian point cloud in the world coordinate system; Loss function Loss align The calculation process is as follows: Loss align = a x Loss ll + b x Loss ssim Lossa ll is the absolute value L1 loss between the rendered Gaussian point cloud position and the supervised image, and represents the absolute value difference between the two supervised images; Loss ssim is the absolute value L1 loss between the rendered Gaussian point cloud position and the supervised image, and represents the absolute value difference between the two supervised images; Loss ll and Loss ssim are the proportions of Lossa and Loss Loss ll = |I pr -I gt | where I pr represents the predicted image by Gaussian sphere rendering; I gt represents the ground truth image before Gaussian sphere rendering; μ pr represents the data of the image window in I pr ; μ gt represents the data of the image window in I gt ; σ prgt represents the covariance between I pt and I gt ; C1 and C2 are both constants in the calculation of Loss ssim .
3. The method of claim 1, wherein the Gaussian sputtering-based head reconstruction method is characterized by, In step S3, the convolutional feature map obtained through the convolutional neural network is obtained in the following manner: The convolutional neural network selects a feature extraction network feature_net; S31, the convolutional neural network receives a three-channel input image; the convolutional neural network comprises 19 convolutional layers, and each convolutional layer is connected and fused through a residual block; First, the sparse input image is preliminarily extracted through the first convolutional layer; Then, the feature channel conversion and size adjustment are sequentially performed through three stages of residual blocks, and finally three feature maps of different scales are output; each stage comprises 3 residual blocks, and each residual block comprises 2 convolutional layers; S32: inputting the supervised image into the convolutional neural network to obtain a convolutional feature map corresponding to the sparse input image.
4. The method of claim 1, wherein the Gaussian sputtering-based head reconstruction method is characterized by, In step S4, the three-dimensional head Gaussian model is obtained in the following manner: S41, obtaining a fusion feature vector of the Gaussian point cloud position of the head: The Gaussian point cloud position in the image coordinate system is bilinearly interpolated and sampled on the convolutional feature map to obtain a feature vector corresponding to the Gaussian point cloud position in the image coordinate system; the feature vectors of the Gaussian point cloud positions of the sparse input images are combined to obtain a fusion feature vector of the Gaussian point cloud position of the head; S42, build a decoding network: The decoding network is a multi-layer perception model for decoding features; the multi-layer perception model includes a preprocessor, a decoder, and an output layer; The processor includes three identical MLP_pre modules; the decoder includes three stages of ResidualBlock1D modules; and the output layer includes three sub-modules of covariance, opacity, and RGB; S43, decode the fusion feature vector through the decoding network: The fusion feature vector obtained in S41 is input into the decoding network, and the fusion feature vector is processed to obtain feature vectors feat1, feat2, and feat3; the feature vectors feat3, feat2, and feat1 are sequentially subjected to feature fusion and up-sampling processing of the decoder, and then are subjected to RGB three sub-modules of the output layer to obtain covariance, opacity, and RGB values, respectively; in combination with the obtained Gaussian point cloud position in S23, a preliminary three-dimensional head Gaussian model is obtained; S44, the multi-layer perception in the decoding network decodes three parameters of covariance, opacity, and RGB values to obtain attributes of the head Gaussian sphere; the attributes of the head Gaussian sphere include color c, opacity a, and covariance matrix Σ; and finally, a three-dimensional head Gaussian model corresponding to the sparse input image of the convolutional neural network is generated.
5. The method of claim 1, wherein the Gaussian sputtering-based head reconstruction method is characterized by, In step S5, the three-dimensional head Gaussian model optimization process is as follows: S51, given the camera pose and camera parameters of the rendering viewpoint, rasterize and render the obtained three-dimensional head Gaussian model to obtain a head rendering image at an arbitrary viewpoint; S52, predict the pixel difference between the image and the training viewpoint image as a loss function Loss net ; Loss function Loss net including absolute value LI loss, structural similarity SSIM loss, and VGG perception loss; Loss net is the supervised term for the convolutional neural network and the decoding network; Adam is selected as the optimizer to adjust the weights and biases of the convolutional neural network to minimize the loss function Loss net ; First, the loss function Loss is calculated by the backpropagation algorithm net For each weight and bias gradient; Then, the gradient descent method is used to update the parameters, the training data is input into the convolutional neural network, and the predicted value is calculated through forward propagation; Then, the parameters are updated through back propagation, and the iteration is continuously performed until the convolutional neural network and the decoding network converge; the iteration number is 10 cycles, all data in the training set are trained once for one cycle, and after the training is completed, the trained convolutional neural network and decoding network are obtained; Loss function Loss net is calculated as follows: Loss net = a x Loss align + b x Loss vgg wherein, Loss vgg represents the VGG perceptual loss between the predicted image and the real image, the VGG perceptual loss is a perceptual loss function for a pre-trained VGG network, by minimizing the difference between the generated image and the target image in the VGG network layer feature map, it is ensured that the content information of the generated image will not deviate too much, the VGG network layer includes conv1_1, conv2_1, conv3_1, conv4_1 and conv5_1 in VGG19, and the difference between the feature maps includes mean square error; and α and β are the proportions of Loss align and Loss vgg loss function respectively. Loss vgg = ||vgg(I px )-vgg(I gt || 2 wherein vgg(I pr ) represents the eigenvalues of the input VGG pre-training network layer of the three-dimensional head Gaussian model rendering image; and vgg(I gt ) represents the eigenvalues of the image of the three-dimensional head Gaussian model except the supervised image.
6. An electronic device, comprising: Comprise: One or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to perform the method of any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that, The computer readable storage medium stores program code, which can be called and executed by the processor to perform the method of any one of claims 1 to 5.
Citation Information
Cited By
Sparse viewpoint 3D Gaussian sputtering reconstruction method based on adaptive viewpoint sampling
CN121527328A
A sparse-view 3D Gaussian sintering reconstruction method based on adaptive view sampling
CN121527328B