Three-dimensional head portrait model generation method based on single photo

Through the methods of face feature point detection, depth estimation and Gaussian point cloud rendering, the time-consuming and cost-effective problems in the existing technology are solved, and high-precision and automated three-dimensional avatar generation are realized, which can accurately restore the detailed characteristics of the photo.

CN120339503APending Publication Date: 2025-07-18SUZHOU MUJINZHI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510320597.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing face modeling techniques are time-consuming, cost-effective and difficult to accurately restore personalized features in photos, such as wrinkles, beards or hairstyles, especially in the side face area, which leads to insufficient similarity deviation between the model and the photo and structural consistency.

Method used

Using face feature point detection, SAM algorithm background segmentation, Depth Anything model depth estimation, SV3D video diffusion and Gaussian splashing three-dimensional reconstruction methods, 360° rotation video is generated and point cloud data is extracted. The video frame rate is optimized through multi-view consistency loss and optical flow interpolation algorithm, and Gaussian point cloud rendering is finally performed to generate a high-precision three-dimensional avatar model.

Benefits of technology

It realizes the automatic generation of high-precision three-dimensional avatars from a single photo, and the model is highly matched with the input photos, retaining rich detailed features, reducing labor costs and improving model authenticity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339503A_ABST
    Figure CN120339503A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of photo three-dimensional processing, and discloses a three-dimensional head portrait model generation method based on a single photo, which comprises the following specific steps: photo input processing: firstly, extracting a face region from the single photo by using a face feature point detection algorithm; background segmentation is carried out by using an SAM algorithm, and a face region in the picture is extracted; finally, the depth information of the human face is speculated through a Depth Anything model, and the depth information is used for guiding generation of Gaussian point cloud; according to the three-dimensional head model generation method based on the single photo, the high-precision three-dimensional head model can be generated from the single photo. According to the method, the 360-degree rotating video is firstly generated, and then three-dimensional reconstruction is carried out by utilizing Gaussian splashing, so that the finally generated 3D head model is highly matched with the input picture, and the similarity of the model is relatively high. Besides, the method can retain abundant detail features, and the Gaussian splashing point cloud can accurately restore personalized features of the human face, such as wrinkles, skin colors, hairlines and the like, so that the model is more real and natural.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional photo processing, and specifically to a method for generating a three-dimensional avatar model based on a single photo. Background Art

[0002] With the rise of the metaverse concept, three-dimensional digital humans are widely used in industries such as virtual reality, e-commerce live streaming, film and television games. The core technology of three-dimensional digital humans is the construction of three-dimensional face models. Currently, face modeling mainly relies on the following two methods:

[0003] Manual modeling by artists: It mainly relies on three-dimensional modeling software (such as Max, Maya, ZBrush) for manual carving to construct a three-dimensional head model that meets the requirements. The main advantage of this method is that artists can finely adjust the model to generate a high-quality face model.

[0004] 3DMM modeling (3DMorphable Model): This method is based on statistical learning. By constructing a parametric 3D face model and using a single image to fit the 3DMM parameters, the corresponding 3D head model can be generated. Its main advantage is that this method only requires a small amount of input data to generate a complete 3D structure, and at the same time, the computational cost is small, and the model construction can be completed quickly.

[0005] However, there are deficiencies in manual modeling by artists. First of all, the modeling process is very time-consuming, and it usually takes several days or even a month to complete a high-quality 3D avatar. Secondly, manual modeling requires experienced modeling engineers, so the labor cost is relatively high. In addition, since the model is manually carved by artists, it is difficult to accurately reproduce the face features in the input photo, resulting in a certain similarity deviation between the final model and the original photo.

[0006] 3DMM also has some limitations. Since it mainly relies on linear deformation for fitting, the model lacks in detail performance and is difficult to accurately restore personalized features in the photo, such as wrinkles, beards or hairstyles. In addition, since 3DMM only fits a limited number of parameters, the modeling accuracy in the side face area is relatively low, which may lead to a large error during perspective conversion, thus affecting the consistency of the model structure. Summary of the Invention

[0007] To solve the technical problems raised in the above background art, the present invention is realized through the following technical solutions: A method for generating a three-dimensional avatar model based on a single photo, including the following specific steps:

[0008] Photo Input Processing: First, use a facial landmark detection algorithm to extract the face region from a single photo; then use the SAM algorithm for background segmentation to extract the face region in the picture; finally, infer the depth information of the face through the Depth Anything model to guide the generation of Gaussian point clouds;

[0009] Generate 360° Rotating Video: Use the SV3D video diffusion model to generate multi-view views, ensure the geometric consistency of multi-view images in three-dimensional space according to the 3D consistency loss function, and finally use the real-time optical flow evaluation interpolation algorithm to improve the video frame rate to complete the generation of 360° rotating video from a single photo;

[0010] Gaussian Projection 3D Reconstruction: Use Colmap to extract point cloud data from video frames, and use the Gaussian projection algorithm to reconstruct a three-dimensional head model according to the depth information of the photo.

[0011] Furthermore, the facial landmark detection algorithm uses a multi-stage cascaded deep learning method for facial landmark detection, specifically:

[0012] Use the RetinaFace detector to locate the face region in the photo;

[0013] Use a multi-task convolutional neural network to locate several facial feature points, including: feature points of the eye contour and pupils, feature points of the eyebrows, feature points of the nose contour, feature points of the outer and inner contours of the lips, and feature points of the outer contour of the face;

[0014] Normalize the feature point coordinates to the range [0,1] and construct a semantic grouping of feature points.

[0015] Furthermore, the extraction of the face region is specifically as follows:

[0016] Load the pre-trained model and configure the inference parameters;

[0017] Use the detected face center point as a key prompt point and use the SAM algorithm to automatically generate a segmentation mask;

[0018] Extract the foreground portrait according to the mask, that is, the face region.

[0019] Furthermore, use the Depth Anything model to infer the depth information of a single photo and construct a three-dimensional structure prior, specifically:

[0020] Preprocess the image of the face region:

[0021]

[0022] where I is the set of input image pixels, I normis the set of standardized image pixels, where μ and σ are the mean and standard deviation of the dataset;

[0023] Adjust the depth scale according to the facial feature points and normalize the depth values to the standard range [0, 1];

[0024] Depth point cloud generation: Construct the camera projection matrix and calculate the point cloud coordinates based on the pixel coordinates of the depth map.

[0025] Furthermore, the 3D consistency loss function ensures the geometric consistency of multi-view images in 3D space, specifically:

[0026] First, extract the feature maps of two views, and then project the feature map of one view F1 to another view F2 through depth information and view transformation;

[0027] Based on the depth-based feature warping calculation, transform the image of one view F1 to another view F2 to ensure the 3D structure is consistent. First, back-project from one view F1 to the 3D space, apply the view transformation in the 3D space, project back to another view F2, and use bilinear interpolation to calculate the feature weights;

[0028] Multi-scale consistency constraint calculation to ensure that both details and the overall structure are consistent and prevent the model from deviating at different resolutions. Calculate the consistency loss at multiple scales i:

[0029] L ms = ∑λ i *L feat (F1 i , F2 i );

[0030] Where: L ms is the deviation accumulated at multiple scales, L feat is the error under different views, λ i is the weight at different scales, F1 i , F2 i are the feature maps at different resolutions;

[0031] Geometric regularization term loss calculation, smooth the depth map, reduce noise, and keep the edges clear at the same time. Calculate the loss according to the depth gradient:

[0032]

[0033] Where, L smooth is the smoothing loss, λ is the degree of edge smoothing, is the depth gradient, is the gradient of the input image, that is, the edge information of the image;

[0034] Cyclic consistency constraint calculation ensures that after the transformation from perspective F1 → perspective F2 → perspective F1, it can still return to the original state to avoid cumulative errors:

[0035] L cycle = ||W(W(F1, d1, T 1→2 ), d2, T 2→1 ) - F1|| 2 ;

[0036] where L cycle is the cyclic consistency loss, W(F, d, T) is the feature transformation function based on depth d and transformation matrix T, d is the depth map, T is the transformation matrix, d1 is the depth map of perspective F1, d2 is the depth map of perspective F2, T 1→2 is the transformation matrix from perspective F1 to perspective F2, and T 2→1 is the transformation matrix from perspective F2 to perspective F1;

[0037] The complete 3D consistency loss calculation formula is:

[0038] L 3D = αL ms + βL smooth + γL cycle ;

[0039] where L 3D is the final 3D loss, which is the weighted sum of multiple sub-loss functions, and α, β, γ are the weights of different loss terms.

[0040] Furthermore, the use of the real-time optical flow evaluation and interpolation algorithm to improve the video frame rate is specifically as follows:

[0041] Content-based motion interpolation generates intermediate frames based on the motion information of the previous and next frames to make the motion smoother. Use PWC-Net to estimate the forward optical flow and backward optical flow, and predict the position of the intermediate frame according to the optical flow;

[0042] Utilize optical flow consistency verification to check whether the optical flow is accurate to avoid distorted images caused by incorrect interpolation. Generate a reliable region mask according to the consistency error of the forward and backward optical flows, and perform interpolation only in the regions where the error is allowed:

[0043]

[0044] M(x, y) = C(x, y) < τ;

[0045] where (x, y) represents the position of the pixel point in the image; F 1→2 (x, y) is the optical flow field from perspective 1 to perspective 2; F 2→1 (x, y) is the optical flow field from perspective 2 to perspective 1, is the horizontal displacement, representing the optical flow component in the x direction; is the vertical displacement, representing the optical flow component in the y direction; C(x, y) is the cycle consistency error, M(x, y) is the consistency mask, τ = 0.01 * max(H, W), where H and W are the image height and width;

[0046] Optimize the optical flow using face region perception to improve the interpolation quality of the face region, avoid face deformation, and generate a face weight map according to the distance D(x, y) from the pixel to the nearest face feature point:

[0047] W face (x, y) = exp(-D(x, y) 2 / 2σ0 2 );

[0048] where W face (x, y) is the face weighted weight, and σ0 controls the speed of weight attenuation;

[0049] Use a recursive refinement strategy to gradually increase the number of interpolation frames to make the video smoother. First, insert intermediate frames between key frames to generate a 2x frame rate video; then recursively apply the same algorithm to each pair of adjacent frames until the target frame rate is reached;

[0050] Use temporal consistency filtering to reduce flicker and make the interpolated video more stable, and smooth the time series:

[0051]

[0052] where, is the smoothed frame, I t-1 , I t+1 is the frame at adjacent time, I t is the original frame;

[0053] Process the head and tail connection of the 360° rotated video to avoid frame breakage, and perform transitional blending on the head and tail frames:

[0054] I new (i) = (1 - α i ) * I n-l+i + α i * I i ;

[0055] where i represents the index of the current interpolation, and the range is 0 ≤ i < l; l represents the total number of interpolation steps, which determines the smoothness of the transition; I new (i) is the new frame after smooth interpolation, I n-l+i is the historical frame, I i is the current frame, is the linear blending ratio for smooth transition.

[0056] Further, the Gaussian projection 3D reconstruction is specifically as follows:

[0057] Use Colmap to extract Colmap point cloud from the generated 360° rotation video;

[0058] Point cloud alignment: Use the ICP algorithm for optimized alignment to minimize the error between the two groups of point clouds and align the two groups of point clouds in the same 3D coordinate system;

[0059] Point cloud weighted fusion: Calculate the credibility of each point, use the Colmap point cloud in the local area, and use the depth estimation point cloud in the structural area;

[0060] Generate Gaussian point cloud: Set Gaussian distribution parameters for each point cloud, including center coordinates, color information, scale parameters, and rotation parameters;

[0061] Optimize the Gaussian point cloud: Use the regularization loss to maintain the uniform distribution of the point cloud, and use the 3D consistency loss to constrain the consistency of the point cloud under different perspectives, thereby reducing the artifacts in Gaussian reconstruction;

[0062] Gaussian point cloud rendering: Use Gaussian projection for 3D reconstruction.

[0063] Compared with the prior art, the present invention has the following beneficial effects:

[0064] The method for generating a 3D avatar model based on a single photo can generate a high-precision 3D head model from a single photo. Compared with traditional methods, first, this method has the characteristic of high automation and completely eliminates the need for manual modeling. The AI can directly generate a complete 3D avatar from a single photo. Second, by first generating a 360° rotation video and then using Gaussian projection for 3D reconstruction, the finally generated 3D head model highly matches the input photo, and the similarity of the model is relatively high. In addition, this method can retain rich detail features, and the Gaussian projection point cloud can accurately restore the personalized features of the human face, such as wrinkles, skin color, hairline, etc., making the model more realistic and natural. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is the flowchart of the method of the present invention;

[0066] Figure 2 is a single face photo selected in the embodiment of the present invention;

[0067] Figure 3 is a schematic diagram of feature points for face area recognition in the embodiment of the present invention;

[0068] Figure 4 is in the multi-view view generated by using SV3D in the embodiment of the present invention Figure 1 ;

[0069] Figure 5 In the embodiment of the present invention, for generating multi-view views using SV3D Figure 2 ;

[0070] Figure 6 In the embodiment of the present invention, for generating multi-view views using SV3D Figure 3 ;

[0071] Figure 7 Schematic diagram of generating point cloud using depth map and Colmap in the embodiment of the present invention;

[0072] Figure 8 Schematic diagram of the 3D avatar generated in the embodiment of the present invention. Detailed implementation manners

[0073] Next, with reference to the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0074] An embodiment of the method for generating a three-dimensional avatar model based on a single photo is as follows:

[0075] Please refer to Figure 1 , the method for generating a three-dimensional avatar model based on a single photo includes the following specific steps:

[0076] I. Photo input processing: First, use a face feature point detection algorithm to extract the face region from a single photo; then use the SAM algorithm for background segmentation to extract the face region in the picture; finally, infer the depth information of the face through the Depth Anything model (depth estimation) to guide the generation of the Gaussian point cloud (generate depth point cloud).

[0077] In the single photo input processing stage, this stage mainly completes face region extraction, background segmentation, and depth information estimation, specifically including:

[0078] 1.1 Face feature point detection algorithm

[0079] Use the RetinaFace detector to locate the face region in the photo;

[0080] Locate 106 facial feature points using a Multi-Task Cascaded Convolutional Network (MTCNN), including: eye contours and pupils (24 points), eyebrows (16 points), nose contours (18 points), outer and inner contours of the lips (18 points), outer contours of the face (24 points), and additional key feature points (6 points);

[0081] Normalize the feature point coordinates to the range [0, 1] and construct semantic groupings of the feature points.

[0082] 1.2 Extract the face region

[0083] Load a pre-trained model and configure the inference parameters;

[0084] Use the detected face center point as a key cue point and automatically generate a segmentation mask using the SAM algorithm;

[0085] Extract the foreground portrait, i.e., the face region, according to the mask.

[0086] 1.3 Use the Depth Anything model to infer the depth information of a single photo and construct a 3D structure prior, specifically:

[0087] Preprocess the image of the face region:

[0088]

[0089] where I is the set of input image pixels, I norm is the set of standardized image pixels, and μ and σ are the mean and standard deviation of the dataset;

[0090] Adjust the depth scale according to the facial feature points (obtain the depth map) and normalize the depth values to the standard range [0, 1];

[0091] Depth point cloud generation: Construct a camera projection matrix and calculate the point cloud coordinates based on the pixel coordinates of the depth map.

[0092] II. Generate a 360° rotation video: Use the SV3D (Single View to 3D) video diffusion model to generate multi-view views, ensure the geometric consistency of the multi-view images in 3D space according to the 3D consistency loss function, and finally use the real-time optical flow evaluation interpolation algorithm to improve the video frame rate to complete the generation of a 360° rotation video from a single photo. The algorithm optimization for generating a 360° rotation video using SV3D can ensure the consistency of the multi-view images to reduce errors in the 3D reconstruction process.

[0093] In the 360° rotation video generation stage, this stage mainly completes the generation of a 360° rotation video from a single photo, specifically:

[0094] 2.1 Design of 3D consistency loss function. According to the 3D consistency loss function, the geometric consistency of multi-view images in three-dimensional space is ensured, specifically as follows:

[0095] Calculation of multi-view feature consistency loss to ensure that the feature information remains consistent when observing the same object from different angles. First, extract the feature maps of two views, and then project the feature map of one view F1 to another view F2 through depth information and view transformation;

[0096] Calculation of depth-based feature warping. Convert the image of one view F1 to another view F2 to ensure the consistency of the 3D structure. First, back-project from one view F1 to the 3D space, apply view transformation in the 3D space, project back to another view F2, and use bilinear interpolation to calculate the feature weights;

[0097] Calculation of multi-scale consistency constraint to ensure that both details and the overall structure can be kept consistent and prevent the model from deviating at different resolutions. Calculate the consistency loss at multiple scales i:

[0098] L ms =∑λ i *L feat (F1 i ,F2 i );

[0099] Among them: Scale i refers to different resolutions, that is, different image sizes. For example, the original image size is 1024×1024, and multiple scales are 512×512, 256×256, etc. L ms is the deviation accumulated at multiple scales, L feat is the error under different views, λ i is the weight of different scales, F1 i ,F2 i are feature maps at different resolutions;

[0100] Calculation of geometric regularization term loss. Smooth the depth map, reduce noise, and at the same time keep the edges clear. Calculate the loss according to the depth gradient:

[0101]

[0102] Among them, L smooth is the smoothing loss, λ is the degree of edge smoothing, is the depth gradient, is the gradient of the input image, that is, the edge information of the image;

[0103] Calculation of cyclic consistency constraint to ensure that after the transformation from view F1→view F2→view F1, it can still return to the original state and avoid cumulative errors:

[0104] Lcycle = ||W(W(F1, d1, T 1→2 ), d2, T 2→1 ) - F1|| 2 ;

[0105] Among them, L cycle is the cycle consistency loss, W(F, d, T) is the feature transformation function based on the depth d and the transformation matrix T, d is the depth map, T is the transformation matrix, d1 is the depth map of the viewing angle F1, d2 is the depth map of the viewing angle F2, T 1→2 is the transformation matrix from the viewing angle F1 to the viewing angle F2, T 2→1 is the transformation matrix from the viewing angle F2 to the viewing angle F1;

[0106] The complete 3D consistency loss calculation formula is:

[0107] L 3D = αL ms + βL smooth + γL cycle ;

[0108] Among them, L 3D is the final 3D loss, which is the weighted sum of multiple sub-loss functions, and α, β, and γ are the weights of different loss terms.

[0109] 2.2 Video Frame Rate Optimization and Frame Interpolation Algorithm

[0110] The video frame rate directly generated by the SV3D model is relatively low (8 - 16fps). By designing a real-time optical flow evaluation frame interpolation algorithm, the frame rate can be increased to over 60fps. Specifically:

[0111] Content-based motion interpolation, generating intermediate frames according to the motion information of the front and back frames to make the motion smoother, using PWC-Net to estimate the forward optical flow and the backward optical flow, and predicting the position of the intermediate frame according to the optical flow;

[0112] Using optical flow consistency verification to check whether the optical flow is accurate and avoid picture distortion caused by incorrect frame interpolation. Generating a reliable region mask according to the consistency error of the front and back optical flows, and performing interpolation only in the regions where the error is allowed:

[0113]

[0114] M(x, y) = C(x, y) < τ;

[0115] Among them, (x, y) represents the position of a certain pixel point in the image; F 1→2 (x, y) is the optical flow field from viewing angle 1 to viewing angle 2; F 2→1 (x, y) is the optical flow field from viewing angle 2 to viewing angle 1, is the horizontal displacement, representing the optical flow component in the x direction; is the vertical displacement, representing the optical flow component in the y direction; C(x, y) is the cyclic consistency error, M(x, y) is the consistency mask, τ = 0.01 * max(H, W), where H and W are the image height and width;

[0116] Utilize the optical flow optimization with facial region perception to enhance the interpolation quality of the face region, avoid facial deformation, and generate a facial weight map according to the distance D(x, y) from the pixel to the nearest facial feature point:

[0117] W face (x, y) = exp(-D(x, y) 2 / 2σ0 2 );

[0118] where W face (x, y) is the facial weighted weight, and σ0 controls the speed of weight attenuation;

[0119] Utilize the recursive refinement strategy to gradually increase the number of interpolation frames to make the video smoother. First, insert intermediate frames between key frames to generate a 2x frame rate video; then recursively apply the same algorithm to each pair of adjacent frames until the target frame rate is reached;

[0120] Utilize temporal consistency filtering to reduce flicker and make the interpolated video more stable, and perform smoothing processing on the time series:

[0121]

[0122] where, is the smoothed frame, I t-1 , I t+1 is the frame at adjacent time, I t is the original frame;

[0123] Handle the head and tail connection of the 360° rotated video to avoid frame breakage, and perform transitional blending on the head and tail frames:

[0124] I new (i) = (1 - α i ) * I n-l+i + α i * I i ;

[0125] where i represents the index of the current interpolation, with the range 0 ≤ i < l; l represents the total number of interpolation steps, which determines the smoothness of the transition; I new (i) is the new frame after smooth interpolation, I n-l+i is the historical frame, I i is the current frame, is the linear blending ratio for smooth transition.

[0126] III. Gaussian Projection 3D Reconstruction: Use Colmap to extract point cloud data (i.e., extract 3D point cloud) from video frames, and according to the depth information of the photos, use the Gaussian projection algorithm to reconstruct the three-dimensional head model. By adopting an algorithm that fuses the depth information with the point cloud generated by the video (point cloud fusion and optimization), the quality of the point cloud can be improved, the three-dimensional avatar structure can be guaranteed, and the reconstructed avatar can be made more realistic. Specifically:

[0127] Use Colmap to extract Colmap point cloud from the generated 360° rotating video;

[0128] Point cloud alignment (aligning the depth point cloud with the Colmap point cloud): Use the ICP algorithm for optimized alignment to minimize the error between the two groups of point clouds and align the two groups of point clouds in the same 3D coordinate system;

[0129] Point cloud weighted fusion: Calculate the credibility of each point, use the Colmap point cloud in the local area, and use the depth estimation point cloud in the structural area;

[0130] Generate Gaussian point cloud: Set Gaussian distribution parameters for each point cloud, including center coordinates, color information, scale parameters, and rotation parameters;

[0131] Optimize the Gaussian point cloud: Adopt a regularization loss to maintain the uniform distribution of the point cloud, and adopt a 3D consistency loss to constrain the consistency of the point cloud under different perspectives, thereby reducing the artifacts in Gaussian reconstruction;

[0132] Gaussian point cloud rendering: Adopt Gaussian projection for 3D reconstruction.

[0133] Select a photo (as shown in Figure 2 ), and use the above method to generate a three-dimensional avatar model. This photo is a frontal portrait photo with a resolution of 720×1110 pixels, as shown in Figure 2 . The facial expression of the person in this photo is natural, the lighting is uniform, and the background is a solid color.

[0134] Use the RetinaFace detector to locate the face area in the photo, and use the multi-task convolutional neural network (MTCNN) to locate 106 facial feature points, specifically including: eye contours and pupils (24 points), eyebrows (16 points), nose contours (18 points), outer and inner contours of the lips (18 points), outer contours of the face (24 points), and additional key feature points (6 points), as shown in Figure 3 .

[0135] As shown in Figure 4 , Figure 5 and Figure 6 , use the SV3D video diffusion model to generate multi-view views.

[0136] As Figure 7 shown, a depth map and Colmap are used to generate a point cloud.

[0137] As Figure 8 shown, a 3D avatar is generated using 3D reconstruction rendering with Gaussian projection technology.

[0138] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention.

Claims

1. A method for generating a 3D avatar model based on a single photo, characterized in that, It includes the following specific steps: Photo input processing: First, use a face landmark detection algorithm to extract the face region from a single photo; then use the SAM algorithm for background segmentation to extract the face region in the picture; finally, infer the depth information of the face through the Depth Anything model to guide the generation of Gaussian point clouds; Generate a 360° rotating video: Use the SV3D video diffusion model to generate multi-view views, ensure the geometric consistency of multi-view images in three-dimensional space according to the 3D consistency loss function, and finally use the real-time optical flow evaluation and interpolation algorithm to improve the video frame rate to complete the generation of a 360° rotating video from a single photo; Gaussian splash three-dimensional reconstruction: Use Colmap to extract point cloud data from video frames, and use the Gaussian splash algorithm to reconstruct a three-dimensional head model according to the depth information of the photo.

2. The method for generating a three-dimensional avatar model based on a single photo according to claim 1, wherein: The face landmark detection algorithm uses a multi-stage cascaded deep learning method for face landmark detection, specifically: Use the RetinaFace detector to locate the face region in the photo; Use a multi-task convolutional neural network to locate several facial landmarks, including: landmarks of the eye contour and pupils, landmarks of the eyebrows, landmarks of the nose contour, landmarks of the outer and inner contours of the lips, and landmarks of the outer contour of the face; Normalize the landmark coordinates to the range [0,1] and construct landmark semantic groups.

3. The method for generating a three-dimensional avatar model based on a single photo according to claim 2, wherein: The extraction of the face region is specifically as follows: Load the pre-trained model and configure the inference parameters; Use the detected face center point as the key prompt point and use the SAM algorithm to automatically generate a segmentation mask; Extract the foreground portrait according to the mask, that is, the face region.

4. The method for generating a 3D avatar model based on a single photo according to claim 3, characterized in that: Use the Depth Anything model to infer the depth information of a single photo and construct a three-dimensional structure prior, specifically: Perform standardized preprocessing on the image of the face region: where I is the set of input image pixels, and I norm is the set of image pixels after normalization, and μ and σ are the mean and standard deviation of the data set; Adjust the depth scale according to the face landmarks and normalize the depth values to the standard range [0,1]; Depth point cloud generation: Construct a camera projection matrix and calculate the point cloud coordinates according to the pixel coordinates of the depth map.

5. The method for generating a 3D avatar model based on a single photo according to claim 1, wherein: The ensuring of the geometric consistency of multi-view images in three-dimensional space according to the 3D consistency loss function is specifically as follows: First, extract the feature maps of two views, and then project the feature map of one view F1 to another view F2 through depth information and view transformation; Based on depth-based feature warping calculation, transform the image of one view F1 to another view F2 to ensure 3D structure consistency. First, back-project from one view F1 to 3D space, apply view transformation in 3D space, project back to another view F2, and use bilinear interpolation to calculate the feature weights; Multi-scale consistency constraint calculation to ensure that both details and the overall structure can be kept consistent and prevent the model from deviating at different resolutions. Calculate the consistency loss at multiple scales i: L ms = ∑λ i * L feat (F1 i , F2 i ); Where: L ms is the deviation accumulated at multiple scales, L feat is the error under different perspectives, λ i is the weight of different scales, F1 i , F2 i are feature maps at different resolutions; Geometric regularization term loss calculation to smooth the depth map, reduce noise, and keep the edges clear at the same time. Calculate the loss according to the depth gradient: Among them, L smooth Smoothing loss, λ represents the degree of edge smoothing, is the depth gradient, is the gradient of the input image, that is, the edge information of the image; Cycle consistency constraint calculation to ensure that after the transformation from view F1 → view F2 → view F1, it can still return to the original state and avoid cumulative errors: L cycle = ||W(W(F1, d1, T 1→2 ), d2, T 2→1 ) - F1|| 2 ; Among them, L cycle is the cyclic consistency loss, W(F, d, T) is the feature transformation function based on the depth d and the transformation matrix T, d is the depth map, T is the transformation matrix, d1 is the depth map of the view F1, d2 is the depth map of the view F2, T 1→2 is the transformation matrix from the view F1 to the view F2, T 2→1 is the transformation matrix from the view F2 to the view F1; The complete 3D consistency loss calculation formula is: L 3D = αL ms + βL smooth + γL cycle ; Among them, L 3D is the final 3D loss, which is the weighted sum of multiple sub-loss functions, and α, β, and γ are the weights of different loss terms.

6. The method for generating a 3D avatar model based on a single photo according to claim 5, wherein: The method of improving the video frame rate by using a real-time optical flow evaluation frame interpolation algorithm is as follows: Content-based motion interpolation is used to generate intermediate frames based on the motion information of the previous and subsequent frames, making the motion smoother. The PWC-Net is used to estimate the forward optical flow and backward optical flow, and the position of the intermediate frame is predicted according to the optical flow; Optical flow consistency verification is used to check whether the optical flow is accurate and avoid image distortion caused by incorrect frame interpolation. A reliable region mask is generated according to the consistency error of the forward and backward optical flows, and interpolation is only performed in the region where the error is allowed: M(x,y) = C(x,y) < τ; Among them, (x, y) represents the position of a pixel point in the image; F 1→2 (x, y) is the optical flow field from view 1 to view 2; F 2→1 (x, y) is the optical flow field from view 2 to view 1, is the horizontal displacement, representing the optical flow component in the x direction; is the vertical displacement, representing the optical flow component in the y direction; C(x, y) is the cyclic consistency error, M(x, y) is the consistency mask, τ = 0.01 * max(H, W), where H and W are the image height and width; Optical flow optimization based on facial region perception is used to improve the frame interpolation quality of the human face region and avoid facial deformation. A facial weight map is generated according to the distance D(x,y) from the pixel to the nearest facial feature point; W face (x,y) = exp(-D(x,y) 2 / 2σ0 2 ); where W face (x, y) is the facial weighting weight, and σ0 controls the speed of weight decay; A recursive refinement strategy is used to gradually increase the number of frame interpolation times to make the video smoother. First, intermediate frames are inserted between key frames to generate a video with a 2-fold frame rate; then the same algorithm is recursively applied to each pair of adjacent frames until the target frame rate is reached; Temporal consistency filtering is used to reduce flicker and make the interpolated video more stable, and the time series is smoothed; Among them, is the smoothed frame, I t-1 , I t+1 is the frame at adjacent time, I t is the original frame; The connection between the head and tail of the 360° rotating video is processed to avoid image breakage, and transitional blending is performed on the head and tail frames; I new (i) = (1 - α i ) * I n-l+i + α i * I i ; Among them, i represents the index of the current interpolation, where 0 ≤ i < l; l represents the total number of interpolation steps, which determines the smoothness of the transition; I new (i) is the new frame after smooth interpolation, I n-l+i is the historical frame, I i is the current frame, is the linear mixing ratio for smooth transition.

7. The method for generating a 3D avatar model based on a single photo according to claim 1, wherein: The specific method of Gaussian projection three-dimensional reconstruction is as follows: Colmap points are extracted from the generated 360° rotating video using Colmap; Point cloud alignment: The ICP algorithm is used for optimized alignment to minimize the error between the two sets of point clouds and align the two sets of point clouds in the same 3D coordinate system; Point cloud weighted fusion: The credibility of each point is calculated, and the Colmap point cloud is used in the local area, and the depth estimation point cloud is used in the structural area; Generate Gaussian point clouds: Gaussian distribution parameters are set for each point cloud, including center coordinates, color information, scale parameters, and rotation parameters; Optimize Gaussian point clouds: Regularization loss is used to maintain the uniform distribution of point clouds, and 3D consistency loss is used to constrain the consistency of point clouds from different perspectives, thereby reducing the artifacts of Gaussian reconstruction; Gaussian point cloud rendering: Gaussian projection is used for three-dimensional reconstruction.

Citation Information

Cited By

  • Extensible multi-view 3D Gaussian splashing method based on credibility driving

    CN120953464A

  • Credibility-driven scalable multi-view 3D gaussian splatting method

    CN120953464B