A multi-view three-dimensional human head model reconstruction method based on occipital complementation
Through the multi-view three-dimensional human head model reconstruction method based on back-head completion, the three-dimensional generative adversarial network is used to generate back-head images and train high-quality models, which solves the problem of missing or distortion of back-head region data in the prior art, and realizes high-quality three-dimensional human head model reconstruction.
Patent Information
- Application Number
- CN202510260085.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The existing three-dimensional human head model reconstruction method ignores the difficult-to-acquire perspective data such as the back of the head, resulting in obvious shortcomings in the integrity and detail restoration of the reconstruction model, especially in the geometric shape and texture details of the back of the head.
A multi-view three-dimensional human head model reconstruction method based on back-head completion was adopted. By acquiring multi-view data sets, the three-dimensional generation adversarial network iterative optimization is used to generate images of the back-head part, and a high-quality three-dimensional human head model is trained in combination with the back-head view data set.
It realizes a high-quality, high-fidelity three-dimensional human head model with geometric texture consistency under the premise of multi-view input images, which significantly improves the reconstruction integrity of the back head area and can generate rich details back head area data.
Smart Images

Figure CN119762372B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of three-dimensional human head model reconstruction, and particularly to a multi-view three-dimensional human head model reconstruction method based on occipital complementation. Background Art
[0002] Three-dimensional human head model reconstruction is one of the core tasks in the research of digital human three-dimensional reconstruction, and its accuracy and detail restoration ability directly affect the realism and naturalness of digital humans. Traditional three-dimensional human head model reconstruction methods mainly rely on devices such as multi-view images, depth sensors, or laser scanners to obtain three-dimensional data. Although these methods can generate high-precision three-dimensional models, they have limitations such as high equipment costs, complex operations, and strict environmental requirements, making it difficult to meet the convenience and universality requirements in practical applications.
[0003] In recent years, three-dimensional reconstruction methods based on single or a small number of images have gradually become a research hotspot, especially using deep learning technology for model reconstruction. These methods can infer three-dimensional geometric information from limited input images by learning the head shape and texture features in large-scale datasets. However, existing methods mainly rely on frontal or side image information, ignoring perspective data such as the occipital region that is difficult to obtain, resulting in obvious deficiencies in the integrity and detail restoration of the reconstructed model, especially in the geometric shape and texture details of the occipital region.
[0004] In practical applications, the integrity of the occipital region is crucial for the reconstruction of three-dimensional human head models. For example, in virtual reality or film and television production, users or viewers may observe the model from any angle. If the occipital region is missing or distorted, it will seriously affect the visual effect and user experience. In addition, the details of the hair, ears, etc. in the occipital region are complex and diverse, further increasing the reconstruction difficulty. The occipital complementation technology can effectively fill the missing parts in multi-view data by generating or inferring the geometric and texture information of the occipital region. Therefore, how to use the occipital complementation technology to improve the reconstruction effect of three-dimensional human head models has become a current research hotspot.
[0005] Therefore, there is a need for a method for high-precision three-dimensional human head model reconstruction based on occipital complementation. Summary of the Invention
[0006] The main purpose of the present invention is to provide a multi-view three-dimensional human head model reconstruction method based on occipital complementation to solve the problem that in the prior art, perspective data such as the occipital region that is difficult to obtain is ignored, resulting in obvious deficiencies in the integrity and detail restoration of the reconstructed model, especially in the geometric shape and texture details of the occipital region.
[0007] To achieve the above object, the present invention provides a multi-view three-dimensional human head model reconstruction method based on occipital completion, which specifically includes the following steps:
[0008] S1. Obtain a multi-view data set of the target person, and use a pose estimation model to estimate the camera pose parameters of the frontal face orientation in the multi-view data set images.
[0009] S2. Use the multi-view images and the corresponding camera pose parameters as supervision information to iteratively optimize the three-dimensional generative adversarial network.
[0010] S3. Render the images of the occipital part based on the optimized three-dimensional generative adversarial network, and use the images and the corresponding camera parameters as the occipital view data set.
[0011] S4. Use the occipital view data set and the original multi-view data set together as supervision information to train a high-quality three-dimensional human head model.
[0012] Further, step S1 specifically includes the following steps:
[0013] S1.1. Obtain a multi-view data set of the target person, perform face detection on the images, screen and retain the images containing valid face regions, and use the face alignment model 3DDFA to crop the multi-view images so that the coordinate origin of the multi-view images is aligned with the center position of the face.
[0014] S1.2. Based on the face orientation information in the multi-view images , use the face alignment model 3DDFA to estimate the camera parameters corresponding to the faces:
[0015] ;
[0016] wherein, is the face alignment model 3DDFA.
[0017] Further, step S2 specifically includes the following steps:
[0018] S2.1. Randomly sample Gaussian noise from the latent space containing arbitrary Gaussian noise , and combine the Gaussian noise with the camera parameters under the input view , and map the randomly sampled Gaussian noise to an intermediate latent vector through the mapping network , and use the three-dimensional generative adversarial network to decode the intermediate latent vector , obtain the generated image from the input perspective :
[0019] .
[0020] S2.2, calculate the mean squared error loss between the multi-view image and the generated image and the perceptual loss :
[0021] ;
[0022] ;
[0023] wherein, is the input image index, is the VGG16 neural network for extracting deep features of the image.
[0024] S2.3, jointly optimize the intermediate latent vector according to the mean squared error loss and the perceptual loss , and the regularization term, and update the intermediate latent vector in each iteration, so that the generated image gradually approaches the multi-view image .
[0025] S2.4, after the optimization of the intermediate latent vector is completed, fix the intermediate latent vector and optimize the parameters of the three-dimensional generative adversarial network .
[0026] Furthermore, step S3 specifically includes the following steps:
[0027] S3.1, determine the camera internal parameters and the k camera positions and orientations for rendering the back of the head, and calculate the corresponding camera external parameter matrices according to the camera positions and the camera orientations :
[0028] ;
[0029] wherein, , is the focal length, is the principal point coordinate of the image;
[0030] ;
[0031] wherein, , is the camera facing direction; , is the right direction of the camera; , is the upward direction of the camera; is the norm, , .
[0032] S3.2, Combine the optimized intermediate latent vector with the k camera parameters corresponding to the back-of-the-head view , and use the fine-tuned 3D generative adversarial network to decode and obtain the generated image from the back-of-the-head view :
[0033] .
[0034] S3.3, Combine the generated image , the extrinsic camera parameters and the intrinsic camera parameters corresponding to the generated image in the format of the original multi-view dataset to obtain the back-of-the-head view dataset.
[0035] Further, step S4 specifically includes the following steps:
[0036] S4.1, Bind the initialized Gaussian point cloud to the parameterized human head model FLAME as the Gaussian character 3D human head model, and use Gaussian rendering combined with the camera parameters of the input view to obtain the rendered image corresponding to the view .
[0037] S4.2, Construct a coordinate transformation matrix, including the rotation matrix , the scaling matrix and the translation matrix ; When the input is the back-of-the-head dataset, transform the coordinates of all mesh vertices of the Gaussian character 3D human head model to the same position as the back-of-the-head dataset and then perform rendering:
[0038] ;
[0039] Among them, , , , , is the number of vertices, is the after rendering.
[0040] S4.3, Use the back-of-the-head view dataset and the multi-view dataset as the real images , and calculate the structural similarity loss between the real images and the rendered images:
[0041] ;
[0042] Among them, is the structural similarity index, is the weight coefficient of the structural similarity loss.
[0043] S4.4. In each iteration, use the structural similarity loss between the real image and the rendered image and the geometric constraints between the Gaussian points to adjust the parameters of the parameterized head model and the Gaussian point cloud. For the back-of-the-head view dataset, additionally optimize the coordinate transformation matrix to make the rendered image and the real image gradually approach.
[0044] The present invention has the following beneficial effects:
[0045] The present invention solves the problem of data loss or distortion in the back-of-the-head region of traditional multi-view reconstruction methods. Then, by constructing a coordinate transformation matrix, the three-dimensional space offset problem between the back-of-the-head view dataset and the parameterized head model is solved, ensuring the consistency between the rendered view and the real view, so as to reconstruct a high-quality and high-fidelity three-dimensional head model with geometric and texture consistency under the premise of multi-view input images. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0047] Figure 1 Shows a flowchart of a multi-view three-dimensional head model reconstruction method based on back-of-the-head completion according to the present invention.
[0048] Figure 2 Shows an image of the back-of-the-head part of a Gaussian character three-dimensional head model trained from a multi-view dataset.
[0049] Figure 3 Shows an image of the back-of-the-head part of a Gaussian character three-dimensional head model trained by the method provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0051] As Figure 1 shown, a multi-view three-dimensional human head model reconstruction method based on back-of-head completion specifically includes the following steps:
[0052] S1. Obtain a multi-view data set of the target person, and use a pose estimation model to estimate the camera pose parameters of the front face orientation in the multi-view data set images.
[0053] S2. Use the multi-view images and the corresponding camera pose parameters as supervision information to iteratively optimize the three-dimensional generative adversarial network.
[0054] S3. Render the images of the back-of-head part based on the optimized three-dimensional generative adversarial network, and use the images and the corresponding camera parameters as the back-of-head view data set.
[0055] S4. Use the back-of-head view data set and the original multi-view data set together as supervision information to train a high-quality three-dimensional human head model.
[0056] Specifically, step S1 specifically includes the following steps:
[0057] S1.1. Obtain a multi-view data set of the target person, perform face detection on the images, screen and retain the images containing valid face regions, and use the face alignment model 3DDFA to crop the multi-view images so that the origin of the coordinates of the multi-view images is aligned with the center position of the human face.
[0058] S1.2. Based on the face orientation information in the multi-view images , use the face alignment model 3DDFA to estimate the camera parameters (including camera intrinsics and camera extrinsics) corresponding to the human faces:
[0059] ;
[0060] wherein, is the face alignment model 3DDFA.
[0061] Specifically, step S2 specifically includes the following steps:
[0062] S2.1, Randomly sample Gaussian noise from the latent space containing arbitrary Gaussian noise , and combine the Gaussian noise with the camera parameters (including the camera intrinsics and extrinsics) in the input view. Through the mapping network , map the randomly sampled Gaussian noise to an intermediate latent vector , and use the 3D generative adversarial network to decode the intermediate latent vector to obtain the generated image in the input view:
[0063] .
[0064] S2.2, Calculate the mean squared error loss between the multi-view images and the generated images and the perceptual loss :
[0065] ;
[0066] ;
[0067] where is the input image index, and is the VGG16 neural network used to extract the deep features of the image.
[0068] S2.3, According to the mean squared error loss and the perceptual loss , and the regularization term, jointly optimize the intermediate latent vector , and update the intermediate latent vector in each iteration, so that the generated image gradually approaches the multi-view image .
[0069] S2.4, After the optimization of the intermediate latent vector is completed, fix the intermediate latent vector and optimize the parameters of the 3D generative adversarial network .
[0070] Specifically, step S3 specifically includes the following steps:
[0071] S3.1, Determine the camera intrinsics and the k camera positions and orientations for rendering the back of the head, and calculate the corresponding camera extrinsic matrix according to the camera positions and the camera orientations :
[0072] ;
[0073] Among them, , is the focal length, is the principal point coordinate of the image;
[0074] ;
[0075] Among them, , is the camera facing direction; , is the right direction of the camera; , is the upward direction of the camera; is the norm, , .
[0076] S3.2, Combine the optimized intermediate latent vector with the k camera parameters corresponding to the back-of-the-head view (including camera intrinsics and extrinsics), and use the fine-tuned 3D generative adversarial network to decode and obtain the generated image under the back-of-the-head view:
[0077] .
[0078] S3.3, Combine the generated image , the camera extrinsics and intrinsics corresponding to the generated image, in the format of the original multi-view dataset to obtain the back-of-the-head view dataset.
[0079] Specifically, step S4 specifically includes the following steps:
[0080] S4.1, Bind the initialized Gaussian point cloud to the parametric human head model FLAME as the Gaussian character 3D human head model, and use Gaussian rendering combined with the camera parameters of the input view to obtain the rendered image .
[0081] S4.2, For the problem of the 3D space offset between the back-of-the-head view dataset generated by the 3D generative adversarial network and the parametric human head model, construct a coordinate transformation matrix, including the rotation matrix , the scaling matrix and the translation matrix ; when the input is the back-of-the-head dataset, transform the coordinates of all mesh vertices of the Gaussian character 3D human head model to the same position as the back-of-the-head dataset before rendering:
[0082] ;
[0083] Among them, , , , , is the number of vertices, is the after rendering.
[0084] S4.3. Use the back-of-the-head perspective dataset and the multi-perspective dataset as the real images , and calculate the structural similarity loss between the real image and the rendered image :
[0085] ;
[0086] Among them, is the structural similarity index, is the weight coefficient of the structural similarity loss.
[0087] S4.4. In each iteration, use the structural similarity loss between the real image and the rendered image and the geometric constraints between the Gaussian points to adjust the parameters of the parameterized human head model and the Gaussian point cloud. The back-of-the-head perspective dataset additionally optimizes the coordinate transformation matrix to make the rendered image gradually approach the real image .
[0088] The present invention is implemented under the PyTorch framework, uses the Adam algorithm to optimize the model, the maximum number of iterations of the 3D generative adversarial network is 1000, and the maximum number of iterations of the Gaussian character 3D human head model is 600000.
[0089] After introducing the back-of-the-head completion dataset, the reconstruction integrity of the Gaussian character 3D human head model without real back-of-the-head data is significantly improved, and details such as hair and ears in the back-of-the-head area that do not appear in the dataset can be generated. The comparison between the back-of-the-head part of the Gaussian character 3D human head model trained only with the multi-perspective dataset and the back-of-the-head part of the Gaussian character 3D human head model trained jointly with the back-of-the-head perspective dataset and the multi-perspective dataset is as Figure 2 and Figure 3 shown.
[0090] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.
Claims
1. A multi-view 3D head model reconstruction method based on back of head completion, characterized in that: The specific steps include: S1, obtain a multi-view dataset of the target person, and use a pose estimation model to estimate the camera pose parameters of the front face orientation of the image in the multi-view dataset; S2, taking multi-view images and corresponding camera pose parameters as supervision information, iteratively optimizes the 3D generative adversarial network; S3, based on the optimized 3D generative adversarial network, renders the image of the back of the head, and uses the image and the corresponding camera parameters as the back of the head perspective dataset; S4, uses the back of the head view dataset and the original multi-view dataset as supervision information to train a high-quality 3D human head model; Step S3 specifically includes the following steps: S3.1, determine the camera internal parameters And render the k camera positions and orientations of the back of the head, and according to the camera position and camera orientation Calculate the corresponding camera extrinsic matrix : ; in, , is the focal length, are the principal point coordinates of the image; ; in, , is the camera’s direction; , is the right direction of the camera; , is the upward direction of the camera; is the norm, , ; S3.2, the optimized intermediate latent vector k camera parameters corresponding to the back of the head perspective Combined with fine-tuned 3D generative adversarial network Decode and get the generated image from the back of the head : ; S3.3, will generate an image , generate the camera extrinsic parameters and camera intrinsic parameters corresponding to the image, and combine them together in the format of the original multi-view dataset to obtain the back of the head perspective dataset.
2. The method for reconstructing a multi-view 3D human head model based on back of head completion according to claim 1, characterized in that: Step S1 specifically includes the following steps: S1.1, obtain the multi-view dataset of the target person, perform face detection on the image, and select and retain the valid face area Images, cropped using the face alignment model 3DDFA Multi-view images , so that the multi-view image The coordinate origin is aligned with the center of the face; S1.2, based on multi-view images The face orientation information in the image is estimated using the face alignment model 3DDFA. Camera parameters corresponding to each face : ; in, It is the face alignment model 3DDFA.
3. The method for reconstructing a multi-view 3D head model based on back of head completion according to claim 1, characterized in that: Step S2 specifically includes the following steps: S2.1, randomly sample Gaussian noise from a latent space containing arbitrary Gaussian noise , and the Gaussian noise Camera parameters under input viewing angle Combined, by mapping the network Map randomly sampled Gaussian noise to an intermediate latent vector , using a 3D generative adversarial network Decoding the intermediate latent vector , and get the generated image under the input perspective : ; S2.2, Calculation of multi-view images And generate images The mean square error loss and perceived loss : ; ; in, is the input image index, VGG16 neural network for extracting deep features of images; S2.3, based on mean square error loss and perceived loss , and the regularization term jointly optimizes the intermediate latent vector , updating the intermediate latent vector in each iteration , so that the generated image With multi-view images gradually approaching; S2.4, in the intermediate latent vector After the optimization is completed, the intermediate latent vector is fixed And optimize the 3D generative adversarial network Parameters.
4. The method for reconstructing a multi-view 3D head model based on back of head completion according to claim 1, characterized in that: Step S4 specifically includes the following steps: S4.1, bind the initialized Gaussian point cloud and the parameterized head model FLAME as the Gaussian character 3D head model, and use Gaussian rendering combined with the camera parameters of the input perspective to obtain the rendered image of the corresponding perspective ; S4.2, construct coordinate transformation matrix, including rotation matrix , scaling matrix and translation matrix ; When the input is the back of the head dataset, all mesh vertices of the Gauss character 3D head model are The coordinates are transformed to the same position as the back of the head dataset before rendering: ; in, , , , , is the number of vertices, For the rendered ; S4.3, the back of the head view dataset and the multi-view dataset are used as real images , calculate the structural similarity loss between the real image and the rendered image : ; in, is the structural similarity index, is the weight coefficient of structural similarity loss; S4.4, in each round of iteration, the structural similarity loss between the real image and the rendered image and the geometric constraints between the Gaussian points are used to adjust the parameters of the parameterized head model and the Gaussian point cloud. The back of the head perspective dataset additionally optimizes the coordinate transformation matrix so that the rendered image With real image Gradually approaching.
Citation Information
Patent Citations
Three-dimensional head model generation method and device fused with real face, electronic equipment and storage medium
CN114419255A
Micro-expression recognition method based on local facial region reconstruction and memory contrast learning
CN116311483A