Three-dimensional head model reconstruction method based on single-view face image
By combining the three-dimensional diffusion model and the three-dimensional generative adversarial network, using single-view images to estimate images from other angles of the human head and optimize potential vectors, the problem of difficulty in generating high-quality three-dimensional human head models under single-view input is solved, and efficient and low-cost three-dimensional human head model reconstruction is achieved.
Patent Information
- Application Number
- CN202510074550.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-17
AI Technical Summary
It is difficult for the prior art to use only a single-view image as input to generate a high-quality three-dimensional human head model that maintains spatial consistency at any angle.
Using a method based on the three-dimensional diffusion model and the three-dimensional generative adversarial network, the combination of the pose estimation model, the three-dimensional diffusion model and the three-dimensional generative adversarial network is used to estimate the images from other angles of the human head, and the potential vectors are iteratively optimized to generate a high-quality three-dimensional human head model.
It realizes the reconstruction of a high-quality, high-fidelity three-dimensional human head model with geometric texture consistency under the premise of single-view input images, reducing reconstruction costs and improving efficiency.
Smart Images

Figure CN120047614A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional human head model reconstruction, and particularly relates to a method for reconstructing a three-dimensional human head model based on a single-view face image. Background Art
[0002] The research on three-dimensional reconstruction of digital humans aims to reconstruct a virtual three-dimensional human body model from image or video data taken from different angles, so as to simulate the appearance and details of a real human body. Reconstructing the human head is an essential part of the three-dimensional human body reconstruction task and often determines the final quality of the entire reconstruction task.
[0003] In order to generate a three-dimensional human head model with a fine appearance, traditional modeling methods need to learn parameterized geometry and texture from a large amount of three-dimensional scan data, and then use voxels such as point clouds and triangular meshes for representation. Therefore, generating high-fidelity three-dimensional digital portraits often requires expensive shooting equipment and a harsh shooting environment, which not only increases the difficulty for ordinary people to participate in virtual reality, but also violates the requirements of fast, low-cost, and high-precision three-dimensional portrait reconstruction. In contrast, the three-dimensional digital human reconstruction technology based on single-view image input does not rely on expensive equipment and a harsh environment, significantly reduces the cost of shape capture, improves the speed of human body modeling, and has broad application value in downstream fields such as biomedicine, film and television production, and virtual reality. The image reconstruction task with single-view input refers to the modeling work of the entire three-dimensional human head model that can be achieved only by relying on a single photo containing a frontal face. However, a single frontal face input image often cannot simultaneously contain the features and information of the face, side, and back of the head, resulting in difficulty for traditional methods to reconstruct a complete head model, such as structures like hairstyles and the back of the head. Therefore, how to generate a high-quality three-dimensional human head model that maintains spatial consistency at any angle using only a single-view image as input is the key research direction at present.
[0004] Therefore, there is a need for a method that can generate a high-quality three-dimensional human head model that maintains spatial consistency at any angle using only a single-view image as input. Summary of the Invention
[0005] The main purpose of the present invention is to provide a method for reconstructing a three-dimensional human head model based on a single-view face image, so as to solve the problem in the prior art that a high-quality three-dimensional human head model that maintains spatial consistency at any angle cannot be generated using only a single-view image as input.
[0006] To achieve the above purpose, the present invention provides a method for reconstructing a three-dimensional human head model based on a single-view face image, which specifically includes the following steps:
[0007] S1, obtain a single-view input frontal face image of the person to be reconstructed, and use a pose estimation model to estimate the camera pose parameters of the frontal face orientation in the input image.
[0008] S2. Estimate the images of the human head at other angles using a three-dimensional diffusion model according to the frontal face image and the corresponding camera pose parameters of the frontal face image.
[0009] S3. Use a three-dimensional generative adversarial network to decode the latent vector representing the human head model, and use the frontal face image, the images at other angles, and the corresponding camera parameters as supervision to iteratively optimize the latent vector.
[0010] S4. Render the three-dimensional human head images at any angle based on the optimized latent vector, and extract the corresponding three-dimensional human head geometric model.
[0011] Furthermore, step S1 specifically includes the following steps:
[0012] S1.1. Obtain the RGB image of the front of the person to be reconstructed, and use the 3DDFA model to crop the input image I gt and align the origin of the image coordinates with the center position of the human face.
[0013] S1.2. According to the orientation of the face in image I gt , continue to use the 3DDFA model E to estimate the homogeneous matrix c of the frontal face relative to the coordinate system front :
[0014] c front = E(I gt ).
[0015] S1.3. Add the rotation matrix R to the homogeneous matrix to obtain the camera pose matrix of the human head at any angle, and select the camera parameters c back of the back of the head opposite to the front as the camera pose matrix of the human head at any angle:
[0016]
[0017] c back = R × c front .
[0018] Furthermore, step S2 specifically includes the following steps:
[0019] S2.1. Under the condition of given k camera viewpoints {π 1 , π 2 ,..., π k} and given image y, form the joint distribution p(z) from the RGB images:
[0020] p(z) = p c (x 1:k | y);
[0021] Among them, pc Denote the joint distribution of the RGB image x observed by the 3D diffusion model conditioned on the image y. 1:k
[0022] S2.2. Fit p by learning the 3D diffusion model f c to synthesize the corresponding RGB images (x 1 , x 2 ,..., x k ) under k camera views {π 1:k}:
[0023] (x 1:k ) = f(y, π 1:k );
[0024] S2.3. Represent x 1:k in the form of a Markov chain in the diffusion model:
[0025]
[0026] where denotes the random Gaussian noise at the k camera views in the T-th time step, P(·) is the 3D diffusion model distribution, t represents the time step, denotes the RGB image at the k camera views in the t-th time step, and ∏ t (·) represents the product over t time steps.
[0027] Furthermore, step S3 specifically includes the following steps:
[0028] S3.1. Randomly sample Gaussian noise from the latent space containing arbitrary Gaussian noise and concatenate it with the camera parameter c 1:k at the input view to obtain
[0029] S3.2. Map to the intermediate latent vector w through the mapping network M.
[0030] S3.3. Decode the intermediate latent vector w using the 3D generative adversarial network G to obtain the generated image
[0031] at the input view. S3.4. Calculate the mean squared error L between the real image mse and the generated image p :
[0032]
[0033]
[0034] Among them, 1:k is the input image index, and F(·) is the VGG16 neural network used to extract the deep features of the image.
[0035] S3.5, calculate the total loss L:
[0036] L = L p + L mse 。
[0037] The present invention has the following beneficial effects:
[0038] The present invention proposes a method for reconstructing a single-view input three-dimensional human head model based on a three-dimensional diffusion model and a three-dimensional generative adversarial network. The three-dimensional diffusion model is used to provide spatial information from other perspectives, and the three-dimensional generative adversarial network is used to decode the latent vector. Finally, the mean square error and perceptual loss are used to iteratively optimize the latent vector, so that the latent vector can better represent the spatial features of the three-dimensional human head, thereby reconstructing a high-quality, high-fidelity three-dimensional human head model with geometric texture consistency under the premise of a single-view input image. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0040] Figure 1 Shows a flowchart of a method for reconstructing a three-dimensional human head model based on a single-view face image of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0042] As Figure 1 shown, a method for reconstructing a three-dimensional human head model based on a single-view face image specifically includes the following steps:
[0043] S1, obtain a single-view input frontal face image of the person to be reconstructed, and use a pose estimation model to estimate the camera pose parameters of the frontal face orientation in the input image.
[0044] S2. Estimate the images of the human head at other angles using a three-dimensional diffusion model based on the frontal face image and the corresponding camera pose parameters of the frontal face image.
[0045] S3. Decode the latent vector representing the human head model using a three-dimensional generative adversarial network, and use the frontal face image, the images at other angles, and the corresponding camera parameters as supervision to iteratively optimize the latent vector.
[0046] S4. Render the three-dimensional human head images at any angle based on the optimized latent vector, and extract the corresponding three-dimensional human head geometric model.
[0047] Specifically, step S1 specifically includes the following steps:
[0048] S1.1. Obtain the RGB image of the front of the person to be reconstructed, and use the 3DDFA model to crop the input image I gt and align the origin of the image coordinates with the center position of the human face. Among them, the three-dimensional dense face alignment model (3DDFA) is a method for estimating the pose parameters of the frontal face towards the camera according to the input image.
[0049] S1.2. According to the orientation of the face in the image I gt , continue to use the 3DDFA model E to estimate the homogeneous matrix c of the frontal face relative to the coordinate system front : The matrix size is 4×6, and the homogeneous matrix c front contains external information such as the spatial information of the camera and internal information such as the focal length;
[0050] c front = E(I gt ).
[0051] S1.3. Add the rotation matrix R to the basis of the homogeneous matrix to obtain the camera pose matrix of the human head at any angle. In order to reduce the interference of redundant images on human head modeling, after verification, we select the camera parameters c back of the back of the head part opposite to the front.
[0052]
[0053] c back = R × c front .
[0054] Specifically, step S2 specifically includes the following steps:
[0055] S2.1. Under the condition of given k camera viewpoints {π 1 , π 2 ,..., π k} and given image y, compose the joint distribution p(z) from the RGB images:
[0056] p(z) = p c (x 1:k |y);
[0057] where p c represents the joint distribution of the RGB image x observed by the three-dimensional diffusion model conditioned on the image y 1:k of.
[0058] S2.2, Fit p by learning the three-dimensional diffusion model f c to synthesize the corresponding RGB images (x 1 , π 2 ,..., π k ) at k camera views {π 1:k}:
[0059] (x 1:k ) = f(y, π 1:k );
[0060] S2.3, Express x 1:k in the form of a Markov chain in the diffusion model:
[0061]
[0062] where represents the random Gaussian noise at the k camera views at the T-th time step, P(·) is the three-dimensional diffusion model distribution, t represents the time step, represents the RGB image at the k camera views at the t-th time step, ∏ t (·) represents the product of t time steps. The RGB image and the normal map at the target view can be sampled from the Markov chain.
[0063] Specifically, step S3 specifically includes the following steps:
[0064] S3.1, Randomly sample Gaussian noise from the latent space containing arbitrary Gaussian noise Concatenate with the camera parameter c 1:k at the input view to obtain
[0065] S3.2, Map to the intermediate latent vector w through the mapping network M.
[0066] S3.3, Decode the intermediate latent vector w using the three-dimensional generative adversarial network G to obtain the generated image at the input view
[0067] S3.4, Calculate the real image and the generated image Mean squared error L mse and perceptual loss L p :
[0068]
[0069]
[0070] where 1:k is the input image index, F(·) is the VGG16 neural network for extracting deep features of the image. The VGG16 network is a deep convolutional neural network architecture constructed with continuous small convolutional kernels and pooling layers, and the network depth reaches 16 layers.
[0071] S3.5. Calculate the total loss L:
[0072] L = L p + L mse .
[0073] Supervise the iterative optimization process of the intermediate latent vector w through the total loss L, so that the generated image decoded by w gradually approaches
[0074] Specifically, the specific process of step 4 is as follows:
[0075] Based on the intermediate latent variable w optimized in step 3, a voxel-level reconstruction algorithm (Marching cubes) can be used to extract the geometric model of the three-dimensional human head;
[0076] Based on the intermediate latent variable w optimized in step 3, generate three-dimensional human head images from arbitrary viewpoints.
[0077] The present invention is implemented under the PyTorch framework, and the Adam algorithm is used to optimize the model, and the maximum number of iterations is 1000.
[0078] To verify the feasibility and superiority of the present invention, the following comparative experiments were carried out. The experiments were all carried out under the FFHQ dataset and the K-Hairs dataset.
[0079] Select the current state-of-the-art PanoHead method for the three-dimensional human head generation task under single-view input, and compare the reconstruction results with the recognition results of the present invention. The comparison results are shown in Table 1. The PanoHead method only uses single-view images as input and does not use a three-dimensional diffusion model to expand the spatial information contained in the input.
[0080] The present invention selects three evaluation metrics, namely Mean Squared Error (MSE), Perceptual Loss (Perception), and Fréchet Inception Distance (FID), to evaluate the trained model. MSE and Perception are used to evaluate the difference between the generated frontal face images and the real frontal face images. The smaller the value, the higher the accuracy of the model. FID is divided into two parts: FID-front and FID-back. FID-front evaluates the mean and covariance of the frontal face generated images and the input real images in the deep feature space, and then calculates the Fréchet distance between the two Gaussian distributions of features. Secondly, since the FFHQ dataset does not contain real images of the back of the head, and the K-Hairs dataset only contains images of the back of the head, it is impossible to compare and evaluate the generated results of the back of the head with the real data at the pixel level. To address this issue, the K-Hairs test set, PanoHead, and the generated results of the present invention are divided into three groups according to the yaw angle: 90° ≤ yaw < 180°, yaw = 180°, and 180° < yaw ≤ 270°. The FID of each group is calculated according to the corresponding angle, and finally, the FID values of the three groups are averaged to obtain FID-back, which is used as the final index to evaluate the quality of the back of the head images. The smaller the value of FID, the higher the accuracy of the model. Among the above evaluation metrics, MSE, Perception, and FID-front are calculated on the FFHQ dataset, and FID-back is calculated on the K-Hairs dataset.
[0081] Table 1 Comparison of the generation results of the method of the present invention and PanoHead on the FFHQ and K-Hairs datasets
[0082]
[0083] As can be seen from Table 1, for the MSE, Perception, and FID-front on the FFHQ dataset using the method proposed in the present invention, the values reach 1.8, 6.1, and 13.5 respectively, which are lower than 2.1, 6.9, and 14.0 of PanoHead. For FID-back on the K-Hairs dataset, the method proposed in the present invention also achieves the best result of 52.1, which is lower than 64.9 of PanoHead, effectively improving the reconstruction quality of the 3D human head model.
[0084] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.
Claims
1. A method for reconstructing a three-dimensional head model based on a single-view face image, characterized in that: The specific steps include: S1, obtaining a single-view input frontal face image of the person to be reconstructed, and using a pose estimation model to estimate the camera pose parameters of the frontal face in the input image; S2, based on the frontal face image and the camera pose parameters corresponding to the frontal face image, the three-dimensional diffusion model is used to estimate the image of the head at other angles; S3 uses a 3D generative adversarial network to decode the latent vector representing the head model, and uses the frontal face image, images from other angles, and the corresponding camera parameters as supervision to iteratively optimize the latent vector; S4, based on the optimized latent vector, renders a three-dimensional human head image at any angle and extracts the corresponding three-dimensional human head geometric model.
2. The method for reconstructing a 3D head model based on a single-view face image according to claim 1, characterized in that: Step S1 specifically includes the following steps: S1.1, obtain the RGB image of the front of the person to be reconstructed, and use the 3DDFA model to transform the input image I gt Perform cropping and align the image coordinate origin with the center of the face; S1.2, according to image I gt The orientation of the mid-face, continue to use the 3DDFA model E to estimate the homogeneous matrix c of the front face relative to the coordinate system front : c front =E(I gt ); S1.3, add the rotation matrix R to the homogeneous matrix to obtain the camera pose matrix at any angle of the head, and select the camera parameter c of the back of the head opposite to the front back The camera pose matrix for any angle of the head: c back =R×c front 。 3. The method for reconstructing a 3D head model based on a single-view face image according to claim 1, characterized in that: Step S2 specifically includes the following steps: S2.1, given k camera perspectives {π1, π2, ..., π k } and given image y, the joint distribution p(z) composed of RGB images is: p(z)=p c (x 1:k |y); Among them, p c represents the RGB image x observed by the 3D diffusion model conditioned on image y 1:k The joint distribution of S2.2, fitting p by learning the three-dimensional diffusion model f c , thereby synthesizing the k camera perspectives {π1, π2, ..., π k } under the corresponding RGB image (x 1:k ): (x 1:k )=f(y,π 1:k ); S2.3, x 1:k Expressed as a Markov chain in the diffusion model: in, represents the random Gaussian noise from the k camera perspectives in the Tth time step, P(·) is the three-dimensional diffusion model distribution, t represents the time step, represents the RGB image from the k camera perspectives in the t-th time step, Π t (·) represents the cumulative multiplication of t time steps.
4. The method for reconstructing a 3D head model based on a single-view face image according to claim 1, characterized in that: Step S3 specifically includes the following steps: S3.1, randomly sample Gaussian noise from a latent space containing arbitrary Gaussian noise Will and the camera parameters c under the input viewing angle 1:k Splice and get S3.2, by mapping the network M Mapped to an intermediate latent vector w; S3.3, use the 3D generative adversarial network G to decode the intermediate potential vector w and obtain the generated image under the input perspective S3.4, Calculate the real image And generate images The mean square error L mse With the perceptual loss L p : Where 1:k is the input image index, F(·) is the VGG16 neural network used to extract deep features of the image; S3.5, calculate the total loss L: L=L p +L mse .
Citation Information
Patent Citations
Three-dimensional head model reconstruction method
CN102663820A
A method for 3D reconstruction and texture generation from single-view face based on multi-task learning
CN109255831A
Three-dimensional reconstruction method and device, electronic equipment and storage medium
CN114445562A
Self-supervised single-view 3D reconstruction via semantic consistency
US20210287430A1