Head posture data acquisition method based on dynamic Gaussian head portrait

Through the data acquisition method based on dynamic Gaussian head avatar, the existing head pose data sets have been solved in terms of diversity, extreme pose coverage and labeling accuracy, and a high-quality and diverse head pose data sets have been generated, which significantly improves the generalization ability of the model and its performance in complex scenarios.

CN120125931APending Publication Date: 2025-06-10HANGZHOU DIANZI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411652243.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing head pose datasets have limitations in terms of scale, resolution, labeling accuracy and diversity, especially inadequate data diversity, insufficient extreme pose coverage, labeling accuracy and consistency problems, and insufficient sense of reality.

Method used

The head posture data acquisition method based on dynamic Gaussian head avatar is used, and the head posture data set covering -90 degrees to 90 degrees is generated through NeRSemble data set, BackgroundMattingV2, Multiview-3DMM-Fitting, dynamic Gaussian head posture model and Deep Single-Image Portrait Relighting and other technologies, and the diversity processing of background and lighting is achieved.

Benefits of technology

The generated data set has high data diversity, extensive head posture coverage and high sense of reality, and high labeling accuracy, which can effectively improve the generalization ability of the head posture detection model and its performance in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125931A_ABST
    Figure CN120125931A_ABST
Patent Text Reader

Abstract

The invention discloses a head posture data acquisition method based on a dynamic Gaussian head portrait. The method comprises the following steps: carrying out head modeling training by using a Gaussian head portrait model improved on the basis of Gaussian splashing; after a head model is obtained after training is finished, photos of different expressions are taken, and expression diversity is achieved; rendering a head posture head portrait picture with a head posture range from-90 degrees to 90 degrees by using a Gaussian head portrait, and performing normal distribution on a head posture angle; the background pictures are put into a Background MattingV2 model to replace the background pictures, background diversity is achieved, and at least three different types of background pictures of a pure color background, a texture background and an actual life scene background are adopted; the method comprises the following steps of: putting a light source into a light source, putting the light source into a Deep Single-Image Portion Righting for re-lighting, changing illumination, and realizing illumination diversity; the obtained synthesized head posture data set is subjected to dockerface preprocessing to calibrate a face bounding box, a head posture detection model Hopenet is put into training, a multi-task convolutional neural network is used for preprocessing to calibrate the face bounding box, and a head posture detection model FSA-Net is put into training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a method for collecting head pose data based on a dynamic Gaussian head avatar. Background Art

[0002] Head Pose Estimation (HPE) has long been a research hotspot in the field of computer vision due to its wide application in various computer vision technologies. And a high-precision head pose estimation model has become a key element for the development and implementation of many next-generation consumer technologies, including augmented reality and virtual reality (AR / VR) entertainment systems, human-computer interaction technologies for enhancing human attention and behavior analysis, immersive audio systems, and driver monitoring systems (DMS), etc.

[0003] Currently, the most accurate head pose estimation models are trained by supervised learning methods based on deep neural networks. Since the accuracy of model estimation mainly depends on the accuracy of training data, obtaining real head pose data with a wide range of changes in yaw, pitch, and roll angles has become a highly challenging task. Through collection, research, and observation, publicly available head pose datasets have great limitations in terms of scale, resolution, annotation accuracy, and diversity.

[0004] Existing technical solutions include:

[0005] 300W-3D constructs the model by fitting a 3D deformable model (3DMM) using a multi-features framework. Different from the original algorithm, this method uses the landmark marching method to adjust the 3D landmarks. The constraint of 68 landmarks is always adopted throughout the fitting process. After fitting, each sample will be inspected, and a small number of samples with poor quality will be manually adjusted. Further, 300W-LP uses Matlab to perform head pose rotation augmentation on 300W-3D.

[0006] The AFLW2000-3D dataset contains the first 2000 samples in the AFLW dataset and their corresponding real 3D facial models. Compared with 300W-3D, constructing AFLW2000-3D is more challenging because AFLW ignores occluded landmarks (including occluded and self-occluded ones), and its landmarks do not contain expression information (no lip landmarks). Therefore, specific methods need to be adopted to handle these non-standard landmarks during the construction process. The main steps include:

[0007] 1. Train the pose-independent SDM model on 300W-LP to detect 68 landmarks, and train 7 models to handle 7 different yaw angle intervals.

[0008] 2. Roughly align each AFLW sample according to the annotated yaw angle to obtain 68 landmarks.

[0009] 3. Use the multi-feature framework (MFF) to fit the 3DMM. 68 landmarks are used for initialization, and 21 visible landmarks are used for fitting constraints.

[0010] 4. Manually check the results and filter out 389 failed samples.

[0011] 5. Manually annotate additional landmarks for the 389 failed samples, such as eyebrows, eyes, nose, mouth, and facial contours.

[0012] 6. After adding these additional annotated landmarks, re-fit the 389 samples with MFF.

[0013] Head pose datasets in the prior art: The C3I-SynFace head pose dataset, a large-scale synthetic face dataset generated using the iClone7 Character Creator "Realistic Human 100" toolkit, which contains corresponding ground truth annotations for head poses and facial depths, with different nationalities, genders, races, ages, and clothing. The data is generated from 15 female and 15 male synthetic 3D human models in FBX format extracted from the iClone software. Five facial expressions - neutral, angry, sad, happy, and afraid - are added to the face models to increase more variations. With the help of these models, an open-source data generation pipeline in Python is proposed, which imports these models into the 3D computer graphics tool Blender and renders facial images as well as ground truth annotations for head poses and facial depths in the original format. The dataset contains more than 100,000 ground truth samples and their annotations. With the help of virtual human models, this framework can generate a large number of synthetic face datasets (such as head pose or face depth datasets) and has a high degree of control over face and environmental changes (such as pose, lighting, and background). Such large datasets can be used for the improvement and targeted training of deep neural networks.

[0014] The prior art has at least the following disadvantages:

[0015] There are great limitations in terms of scale, resolution, annotation accuracy, and diversity. The specific manifestations are as follows:

[0016] 1. Insufficient data diversity

[0017] 300W-LP dataset: The samples in the dataset are not rich enough in terms of race, age, gender, etc. Especially the number of ethnic minority samples is small, resulting in limited generalization ability of the model across different races and age groups.

[0018] AFLW2000 dataset: The main samples come from a specific race and background, lacking the diversity of different races and genders, which affects the performance of the model in global applications.

[0019] BIWI dataset: The racial and gender diversity is insufficient, mostly Western faces, resulting in possible instability of the model's performance across different races.

[0020] UPNA dataset: The sample distribution is limited in terms of gender, age, and race, making it difficult to train a model that can adapt to multiple characteristics.

[0021] Pandora dataset: The samples of this dataset are mainly concentrated in specific races and environments, lacking various lighting, age, and race data, which limits the generalization of the model.

[0022] 2. Insufficient coverage of extreme poses

[0023] 300W-LP dataset: Most of the poses in the dataset are obtained by generation, and there are few images at extreme angles, resulting in poor performance under extreme side views, top views, or bottom views.

[0024] AFLW2000 dataset: Although it contains samples at multiple angles, the amount of data at extreme angles is still insufficient, making the model less adaptable to large-angle deflections and complex poses.

[0025] AFLW dataset: It lacks samples at extreme angles, such as side views, top views, and bottom views, which affects the performance of the model under large-angle rotation poses.

[0026] Pandora dataset: The provided pose range is limited, especially lacking sufficient data at extreme poses, resulting in unsatisfactory performance of the model in complex scenarios.

[0027] SynHead dataset: Although the synthetic data contains multiple angles, the number of samples at extreme poses is insufficient, limiting the model's adaptability to extreme poses.

[0028] 3. Annotation accuracy and consistency issues

[0029] 300W-LP dataset: The annotations of this dataset are generated through 3D transformation, which will cause certain errors, especially the annotation accuracy is poor when rotating at large angles.

[0030] AFLW2000 Dataset: Compared with other datasets (such as 300W), the number and definition of key points in AFLW2000 are different, and this inconsistency will cause adaptation difficulties during joint training. Most existing head pose datasets have annotation generation errors. For example, in the AFLW2000 dataset, randomly select some pictures from the dataset at 10-degree intervals. By observation, there are visible head pose ground truth annotation errors to the naked eye. After detection using existing pre-trained head pose models, the errors are indeed large. For example, the ground truth annotation is about 51 degrees but actually about 90 degrees; the ground truth annotation is about 80 degrees but actually about 30 degrees; the ground truth annotation is about 65 degrees but actually about 40 degrees, and so on.

[0031] AFLW Dataset: The number of key point annotations is small, which cannot provide fine-grained facial feature points and limits the application of tasks with high requirements for details.

[0032] BIWI Dataset: Although it contains pose annotations, the detailed annotations for key points are few, which affects the performance in precise detection applications.

[0033] 4. Lack of realism

[0034] AFLW Dataset: Despite having different poses, it lacks samples with complex backgrounds.

[0035] CAS-PEAL Dataset: Although it provides rich expression samples, there are few diverse backgrounds and it lacks the ability to handle complex scenes.

[0036] SynHead Dataset: The data is completely synthetically generated. Although the poses are rich, it lacks the detailed features in real images, which may lead to poor cross-dataset generalization of the model.

[0037] 5. Small dataset size

[0038] AFLW2000 Dataset: It contains 2,000 pictures, and the sample size is relatively small, which may lead to underfitting when training deep learning models.

[0039] UPNA Dataset: The dataset size is very small, and the lack of data volume limits the training effect of the model and it is difficult to support large-scale deep learning tasks.

[0040] Pandora Dataset: The number of samples is small. Especially when compared with other large datasets, the scale limitation results in insufficient information during model training.

[0041] BIWI Dataset: There are about 15,000 images, the scale is not large, and the same individual contains multiple pose images, so the actual diversity is low, which affects the training effect of the model.

[0042] 6. Insufficient illumination variation

[0043] BIWI dataset: Most of the images are taken under relatively fixed lighting conditions, lacking different lighting scenarios, which may lead to poor adaptability of the model in the actual environment.

[0044] UPNA dataset: Lacks data under various lighting conditions, and the model has limited performance in dealing with different lighting environments.

[0045] Pandora dataset: Most of the samples are collected under similar lighting conditions, with insufficient lighting changes, resulting in unsatisfactory performance of the model under various lighting conditions. Summary of the Invention

[0046] In view of this, in order to solve the pain point that it is time-consuming and laborious to obtain test data in the application of estimating the concentration of mixed gas components, the present invention provides a method for collecting head pose data based on a dynamic Gaussian head avatar, including the following steps:

[0047] S1, Based on the NeRSemble dataset, there are at least 267 identities in total. Take at least 30 identities and extract the videos among them;

[0048] S2, Use BackgroundMattingV2 to remove the original background;

[0049] S3, Use Multiview-3DMM-Fitting to obtain expression parameters and landmarks;

[0050] S4, Use a Gaussian head avatar model improved based on Gaussian splash for head modeling training;

[0051] S5, After the training is completed and the head model is obtained, take at least three photos with different expressions for each identity to achieve expression diversity;

[0052] S6, Use the Gaussian head avatar to render head pose avatar pictures with the head pose range from -90 degrees to 90 degrees, and make a normal distribution for the head pose angles;

[0053] S7, Put it into the BackgroundMattingV2 model to replace the background picture, achieve background diversity, and at least adopt three different types of background pictures: solid color background, texture background and actual life scene background;

[0054] S8, Put it into Deep Single-Image Portrait Relighting for relighting, change the lighting, and use at least 7 kinds of lighting to achieve lighting diversity;

[0055] S9. Preprocess the obtained synthetic head pose dataset through dockerface to calibrate the face bounding box, and input it into the head pose detection model Hopenet for training. After preprocessing the face bounding box through a multi-task convolutional neural network, input it into the head pose detection model FSA-Net for training;

[0056] S10. Conduct tests on the test head pose dataset.

[0057] Preferably, in S4, the Gaussian head avatar model represents and models the human head using a controllable dynamic 3D Gaussian distribution based on the Gaussian splashing method.

[0058] Preferably, S4 specifically includes the following steps:

[0059] S41. Dynamic Gaussian head representation;

[0060] S42. Geometric deformation;

[0061] S43. Dynamic color, rotation degree, scale, and opacity representation;

[0062] S44. Training loss function;

[0063] S45. Geometric-guided initialization.

[0064] Preferably, S41 specifically includes: The static 3D Gaussian points contain N points, which are defined by their position X, multi-channel color C, rotation Q in quaternion form, scale S, and opacity A;

[0065] The Gaussian points are rasterized and rendered through the camera parameters μ to obtain a multi-channel image I, and this process is expressed as:

[0066] I = R(X, C, Q, S, A; μ)

[0067] Construct a standard neutral Gaussian model with position X 0 ∈R N×3 、feature vector F 0 ∈R N×128 、rotation Q 0 ∈R N×4 、scale S 0 ∈R N×3 and opacity A 0 ∈R N×1 whose attributes can all be fully optimized;

[0068] The MLP-based dynamic generator Φ is used to generate expression-related dynamic changes. The input is the neutral model attributes, expression coefficient θ, and head pose β, and the output is the final dynamic model attributes {X, C, Q, S, A}. The formulaic representation of the entire dynamic Gaussian head avatar model is:

[0069] {X, C, Q, S, A} = Φ(X 0 , F 0 , Q 0 , S 0 , A 0 ; θ, β).

[0070] Preferably, the S42 specifically includes: using two different MLPs to respectively control the deformation of the expression and the head pose, and the displacement calculation formula for each Gaussian point is:

[0071]

[0072] Among them, two MLPs are used, for defining the displacement caused by the expression, for defining the displacement caused by the pose, def represents the meaning of displacement, λ exp (X 0 ) and λ pose (X 0 ) represent the influence of the expression and the head pose on each Gaussian point.

[0073] Preferably, the S43 specifically includes:

[0074] The calculation formula for the dynamic color C' is:

[0075]

[0076] The color is directly predicted by two color MLPs, for defining the dynamic color that changes with the expression, for defining the dynamic color that changes with the pose, col represents the meaning of color;

[0077] The changes in rotation, scale, and opacity are expressed as:

[0078]

[0079] Predicted by two MLPs, for defining the rotation, scaling, and opacity attributes that change with the expression, for defining the rotation, scaling, and opacity attributes that change with the pose, att represents the meaning of rotation, scaling, and opacity;

[0080] Finally, the above attributes are transformed into the world coordinate system:

[0081] {X, Q} = T({X′, Q′}, β)

[0082] {C, S, A} = {C′, S′, A′}.

[0083] Preferably, S44 specifically includes:

[0084] During training, a 3D Gaussian model is generated under the expression condition, then it is rendered into a 32-channel image, and finally a high-resolution head image is generated through a super-resolution network; the training loss function includes L1 loss, VGG perceptual loss, and RGB channel loss, and the total loss function is expressed as:

[0085] L = ||I hr - I gt || 1 + λ vgg VGG(I hr , I gt ) + λ lr |I lr - I gt || 1

[0086] where I hr is the generated high-resolution image, I gt is the real high-resolution image, ||I hr - I gt || 1 : the L1 loss between the generated high-resolution image I hr and the real high-resolution image I gt , λ vgg is the weight hyperparameter that controls the VGG perceptual loss term, VGG(I hr , I gt ) is the perceptual loss term, and the difference between the generated image I hr and the real image I gt in the feature space is calculated using the VGG network, λ lr is the weight hyperparameter that controls the low-resolution L1 loss term, I lr is the generated low-resolution image, ||I lr - I gt || 1 is the L1 loss between the generated low-resolution image Ilr and the real image I gt .

[0087] Preferably, S45 specifically includes:

[0088] First, an MLP representing the signed distance field is defined:

[0089] s, η = f sdf (x)

[0090] where s represents the value of the signed distance field, η represents the feature vector of each point, and x is the position of the point;

[0091] RGB loss L RGB is expressed as:

[0092] L RGB = ||I r,g,b - I gt || 1

[0093] where I r,g,b represents the RGB channels of the generated image, and I gt is the real RGB image, and ||·|| 1 represents the L1 norm;

[0094] Contour loss L sil is expressed as:

[0095] L sil = IOU(M, M gt )

[0096] where M is the mask of the generated image, and M gt is the mask of the real image, and IOU(·,·) is used to calculate the overlap between two masks.

[0097] Preferably, the S6 specifically includes: The head pose generation process starts from the elevation angle and azimuth angle to create a basic head rotation matrix R, which rotates the head from the initial position to the desired viewing angle, and this rotation is achieved through the classical rotation matrix formula:

[0098]

[0099]

[0100] The initial head rotation matrix R is calculated as the product of these two rotation matrices;

[0101] After constructing the basic rotation matrix, the roll angle is added to control the rotation of the head around its own axis, and the final head rotation matrix is obtained by multiplying with the roll rotation matrix;

[0102] This rotation matrix is combined with the translation vector T to form the extrinsic matrix of the camera to describe the position and orientation of the camera in the world coordinate system;

[0103] To convert the head pose from the world coordinate system to the camera coordinate system, a view transformation matrix is calculated, and the view transformation matrix combines the rotation matrix R final and the translation vector T to describe the transformation from the world coordinate system to the camera coordinate system:

[0104]

[0105] To project the 3D head pose onto the 2D image plane, a projection transformation matrix is required. The projection matrix, based on the camera's field of view and other projection parameters, converts 3D coordinates into 2D coordinates. The combination of the view transformation matrix and the projection matrix produces a complete projection transformation matrix, which is used to describe the complete transformation process from the 3D head pose to the 2D image plane:

[0106] FullProj = ProjMatrix × Extrinsic

[0107] This process finally generates a complete projection transformation matrix.

[0108] Preferably, in step S9, it includes: during the process of generating the synthetic head pose dataset, recording three key angle parameters for each pose: pitch angle, yaw angle, and roll angle. These pose parameters are directly extracted during the rendering process and used as ground truth annotations;

[0109] To verify the accuracy of the ground truth data, three head pose estimation models are adopted: Hopenet, 6DRepNet, and FSA-Net.

[0110] The present invention has at least the following beneficial effects:

[0111] 1. Strong data diversity and dataset expandability: Currently, based on the NeRSemble dataset, thirty identities are modeled, and the number of pictures has reached more than 90,000. Real head portraits of people with different genders, ages, and races are used to produce the dataset. The BackgroundMattingV2 is used to achieve different backgrounds for the dataset, and the Deep Single-Image Portrait Relighting is used.

[0112] By performing different lighting treatments to increase the diversity of the dataset, head portraits of specific people can be modeled using videos or images, and an exclusive head pose dataset can be made for that person. This dataset is then input into the head pose detection model for training to obtain an exclusive head pose detector.

[0113] 2. The head pose coverage is wide and controllable, and true value annotation is achieved without supervision: Use Dynamic GaussianAvatar to rotate the head avatar pose angle. With its precise, rapid, and high-resolution rendering capabilities, head avatar pictures at any angle from -90 to 90 can be obtained, and the angle can be precisely controlled to generate and automatically annotate the true value. It has the following advantages: 1) Natural pose continuity and realism: Compared with traditional 3D modeling and rendering methods, avatar dynamic Gaussian rendering can more accurately simulate the dynamic changes of the human head in real scenarios. By modeling the head pose with Gaussian distribution, this technology can generate a highly natural sequence of pose changes, avoiding the problem of discontinuous pose switching in traditional methods and greatly enhancing the realism of the data; 2) Automatically generate diverse data: The dynamic Gaussian rendering technology can automatically generate diverse head pose data by adjusting the Gaussian model parameters, covering scenarios under different angles, distances, and lighting conditions. This automated generation method greatly improves the data generation efficiency compared with manual annotation and traditional 3D rendering, and reduces the subjective errors caused by human operations; 3) Efficient data generation process: Compared with the method based on generative adversarial network (GAN), dynamic Gaussian rendering does not rely on a complex network training process and can directly control the generation of a dataset that meets the requirements through model parameters. Due to the high efficiency and controllability of this technology, the data generation speed is fast, and the hardware requirements are low, significantly reducing the cost and time of dataset creation; 4) Enhanced pose diversity and robustness: By simulating head movements in different situations, dynamic Gaussian rendering can generate a more diverse and robust head pose dataset. This provides more comprehensive training data for subsequent head pose detection models, thereby improving the performance of the model in practical applications, especially the accuracy and stability when dealing with complex pose changes.

[0114] 3. High-precision ground truth annotation of head pose in the dataset and ultra-high-fidelity synthetic images: Dynamic GaussiansAvatar is used for head modeling and rendering, which is a head avatar modeling method based on dynamic Gaussian distribution. The recent four-dimensional Gaussian work has been proposed for reconstructing dynamic scenes, but they all belong to static Gaussian splatter modeling and perform poorly when used for more animated head avatar modeling. How to effectively control the deformation of the three-dimensional Gaussian model and model the dynamic appearance through expression coefficients is the key issue in animatable head character modeling. Dynamic Gaussian Avatar proposes an efficient and well-designed geometric guidance initialization strategy. Instead of starting from a random Gaussian or FLAME model, it first optimizes the implicit signed distance function (SDF) field, as well as the color field and deformation MLP, which are used to model the basic geometry, color, and deformation of the head avatar under expression conditions respectively. The SDF field is converted into a mesh through DMTet (Deep Marching Tetrahedra), and the color and deformation of the vertices are predicted by the MLP. Then the mesh is rendered and optimized under the supervision of multi-view RGB images. Finally, the mesh with per-vertex features from the SDF field is used to initialize the 3D Gaussian function located on the basic head surface, while the color and deformation MLP are transferred to the next stage to ensure stable convergence training. Ultra-high-fidelity synthetic images with a resolution of 2K can be obtained after rendering.

[0115] 4. The dataset is produced using free and open-source projects, while the existing head pose dataset C3I-SynFace is created using the paid commercial 3D asset creation tool iClone and the open-source 3D computer graphics software Blender for rendering.

[0116] 5. Strong sense of realism: The dataset is modeled using real human head avatars instead of virtual human figures (such as the C3I-SynFace head pose dataset, a synthetic face dataset generated using the iClone 7 Character Creator "Realistic Human100" toolkit), thereby improving the cross-dataset generalization of the model. Description of the Drawings

[0117] Figure 1 It is a flowchart of the steps of the head pose data acquisition method based on dynamic Gaussian head avatars according to an embodiment of the present invention;

[0118] Figure 2 It is a schematic diagram of the S1 dataset of the head pose data acquisition method based on dynamic Gaussian head avatars according to an embodiment of the present invention;

[0119] Figure 3It is a schematic diagram of S5 expression diversity of the head posture data collection method based on the dynamic Gaussian head avatar according to an embodiment of the present invention;

[0120] Figure 4 It is a schematic diagram of rendering different head postures in S6 of the head posture data collection method based on a dynamic Gaussian head avatar according to an embodiment of the present invention;

[0121] Figure 5 Schematic diagram of S7 background diversity of the head posture data collection method based on the dynamic Gaussian head avatar according to an embodiment of the present invention;

[0122] Figure 6 Schematic diagram of S8 illumination diversity of the head posture data collection method based on a dynamic Gaussian head avatar according to an embodiment of the present invention. DETAILED DESCRIPTION

[0123] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0124] See also Figure 1 , is a flowchart of the method steps of the present invention, comprising the following steps:

[0125] S1, based on the NeRSemble dataset, there are at least 267 identities in total, and at least 30 identities are selected to extract videos from them; see Figure 2 , taken from the nersemble dataset (11 screenshots are given here as examples).

[0126] S2, use BackgroundMattingV2 to remove the original background;

[0127] S3, expression parameters and landmarks were obtained using Multiview-3DMM-Fitting;

[0128] S4, uses the improved Gaussian head portrait model based on Gaussian splashing for head modeling training;

[0129] S5, after training and obtaining the head model, each identity takes at least three photos with different expressions to achieve expression diversity; see Figure 3 .

[0130] S6, uses Gaussian head avatars to render head pose avatar images with head poses ranging from -90 degrees to 90 degrees, and makes a normal distribution of head pose angles; see Figure 4 , rendering schematics of different head poses.

[0131] S7. Put it into the BackgroundMattingV2 model to replace the background image and achieve background diversity. At least three different types of background images, namely solid-color background, textured background, and real-life scene background, are adopted; see Figure 5 。

[0132] S8. Put it into Deep Single-Image Portrait Relighting for relighting, change the lighting, and use at least 7 kinds of lighting to achieve lighting diversity; see Figure 6 。

[0133] S9. Calibrate the face bounding box of the obtained synthetic head pose dataset through dockerface preprocessing, input it into the head pose detection model Hopenet for training, calibrate the face bounding box through multi-task convolutional neural network preprocessing, and input it into the head pose detection model FSA-Net for training;

[0134] S10. Test on the test head pose dataset.

[0135] In S4, the Gaussian Splatting technology is widely recognized due to its low computational cost and efficient modeling and rendering capabilities, and has become a leading technology in complex scene reconstruction and editing. A new 3D head avatar modeling method - Dynamic Gaussian Head Avatar - uses a controllable dynamic 3D Gaussian distribution to represent and model the human head. This method integrates a fully learned deformation field based on multi-layer perceptrons (MLPs) and combines a geometry-guided initialization strategy, enabling ultra-high-fidelity rendering at 2K resolution under lightweight and sparse view settings. Experimental results show that this method is significantly superior to existing state-of-the-art technologies in capturing dynamic details and accurately transmitting expressions..

[0136] S4 specifically includes the following steps:

[0137] S41. Dynamic Gaussian head representation;

[0138] S42. Geometric deformation;

[0139] S43. Dynamic color, rotation, scale, and opacity representation;

[0140] S44. Training loss function;

[0141] S45. Geometry-guided initialization.

[0142] S41 specifically includes: The static 3D Gaussian points contain N points, which are defined by their position X, multi-channel color C, rotation Q in quaternion form, scale S, and opacity A;

[0143] The Gaussian points are rasterized and rendered through the camera parameter μ to obtain a multi-channel image I, and this process is expressed as:

[0144] I = R(X, C, Q, S, A; μ)

[0145] Construct a standard neutral Gaussian model with the position X 0 ∈ R N×3 , the feature vector F 0 ∈ R N×128 , the rotation Q 0 ∈ R N×4 , the scale S 0 ∈ R N×3 and the opacity A 0 ∈ R N×1 of attributes, all of which can be fully optimized;

[0146] The MLP-based dynamic generator Φ is used to generate the dynamic changes related to expressions. The input is the neutral model attributes, the expression coefficient θ, and the head pose β, and the output is the final dynamic model attributes {X, C, Q, S, A}. The formulation of the entire dynamic Gaussian head avatar model is:

[0147] {X, C, Q, S, A} = Φ(X 0 , F 0 , Q 0 , S 0 , A 0 ; θ, β).

[0148] S42 specifically includes: Using two different MLPs to control the deformations of expressions and head poses respectively. The displacement calculation formula for each Gaussian point is:

[0149]

[0150] Among them, two MLPs are used, used to define the displacement caused by expressions, used to define the displacement caused by poses. def represents the meaning of displacement, and λ exp (X 0 ) and λ pose (X 0 ) represent the influences of expressions and head poses on each Gaussian point.

[0151] S43 specifically includes:

[0152] The calculation formula for the dynamic color C’ is:

[0153]

[0154] The color is directly predicted by two color MLPs, which is used to define the dynamic color that changes with the expression, which is used to define the dynamic color that changes with the pose, where col represents the color meaning;

[0155] The changes in rotation, scale, and opacity are expressed as:

[0156]

[0157] Predicted by two MLPs, which is used to define the rotation, scaling, and opacity attributes that change with the expression, which is used to define the rotation, scaling, and opacity attributes that change with the pose, where att represents the meaning of rotation, scaling, and opacity;

[0158] Finally, the above attributes are transformed into the world coordinate system:

[0159] {X,Q} = T({X′,Q′},β)

[0160] {C,S,A} = {C′,S′,A′}.

[0161] S44 specifically includes:

[0162] During training, a 3D Gaussian model is generated under the expression condition, then it is rendered into a 32-channel image, and finally a high-resolution head image is generated through a super-resolution network; the training loss function includes L1 loss, VGG perceptual loss, and RGB channel loss, and the total loss function is expressed as:

[0163] L = ||I hr - I gt || 1 + λ vgg VGG(I hr , I gt ) + λ lr ||I lr - I gt || 1

[0164] where I hr is the generated high-resolution image, I gt is the real high-resolution image, ||I hr - I gt || 1 : the L1 loss between the generated high-resolution image I hr and the real high-resolution image I gt , λ vgg is the weight hyperparameter that controls the VGG perceptual loss term, VGG(I hr , Igt ) is the perceptual loss term, and the VGG network is used to calculate the generated image I hr and the real image I gt in the feature space. λ lr is the weight hyperparameter that controls the low-resolution L1 loss term. I lr is the generated low-resolution image. ||I lr -I gt || 1 is the L1 loss between the generated low-resolution image I lr and the real image I gt .

[0165] S45 specifically includes:

[0166] First, an MLP representing the signed distance field is defined:

[0167] s, η = f sdf (x)

[0168] where s represents the value of the signed distance field, η represents the feature vector of each point, and x is the position of the point;

[0169] The RGB loss L RGB is expressed as:

[0170] L RGB = ||I r,g,b -I gt || 1

[0171] where I r,g,b represents the RGB channels of the generated image, I gt is the real RGB image, and ||·|| 1 represents the L1 norm;

[0172] The contour loss L sil is expressed as:

[0173] L sil = IOU(M, M gt )

[0174] where M is the mask of the generated image, M gt is the mask of the real image, and IOU(·, ·) is used to calculate the overlap between the two masks.

[0175] S6 specifically includes: The head pose generation process starts from the elevation angle and azimuth angle to create the basic head rotation matrix R, which rotates the head from the initial position to the desired viewing angle. This rotation is achieved through the classical rotation matrix formula:

[0176]

[0177]

[0178] The initial head rotation matrix R is calculated as the product of these two rotation matrices;

[0179] After constructing the basic rotation matrix, the roll angle is added to control the rotation of the head around its own axis, and the final head rotation matrix is obtained by multiplying with the roll rotation matrix;

[0180] This rotation matrix combines with the translation vector T to form the extrinsic matrix of the camera, which describes the position and orientation of the camera in the world coordinate system;

[0181] To transform the head pose from the world coordinate system to the camera coordinate system, a view transformation matrix is calculated, and the view transformation matrix combines the rotation matrix R final and the translation vector T to describe the transformation from the world coordinate system to the camera coordinate system:

[0182]

[0183] To project the 3D head pose onto the 2D image plane, a projection transformation matrix is required. The projection matrix converts 3D coordinates to 2D coordinates based on the camera's field of view and other projection parameters. The combination of the view transformation matrix and the projection matrix produces a complete projection transformation matrix, which is used to describe the complete transformation process from the 3D head pose to the 2D image plane:

[0184] FullProj = ProjMatrix × Extrinsic

[0185] This process finally generates a complete projection transformation matrix.

[0186] S9 includes: During the process of generating the synthetic head pose dataset, three key angular parameters of each pose are recorded: pitch angle, yaw angle, and roll angle. These pose parameters are directly extracted during the rendering process and used as ground truth annotations;

[0187] To verify the accuracy of the ground truth data, we adopted three existing head pose estimation models: Hopenet, 6DRepNet, and FSA-Net. The mean absolute error of the test on the generated dataset. It is worth noting that as the number of identities in the dataset increases, the MAE gradually decreases, indicating that the generated dataset has high accuracy and reliability, providing a solid foundation for the training and evaluation of subsequent head pose estimation models.

[0188] Innovative dataset generation application of the present invention: For the first time, the Dynamic Gaussian Avatar method is applied to the generation of head pose datasets. Through its efficient rendering and pose control capabilities, diverse and high-fidelity head pose data is generated.

[0189] Diversity and scalability of the dataset: Using real-person head modeling, covering different genders, ages, and ethnicities, combined with various background and lighting treatments, to ensure the diversity and generalization ability of the dataset. It is possible to create personalized head pose datasets based on the videos or images of specific individuals for training exclusive detection models.

[0190] Wide controllable head pose coverage and automatic ground truth annotation: With the help of Dynamic Gaussian Avatar, head pose generation at any angle from -90° to 90° is achieved, and automatic ground truth annotation is generated without manual intervention, ensuring the natural continuity and high realism of the pose data.

[0191] Low-cost dataset production process: Using open-source projects and existing technologies (such as BackgroundMattingV2 and Deep Single-Image Portrait Relighting), avoiding commercial 3D tools, thus reducing the cost of dataset production.

[0192] Modeling based on real people to improve generalization ability across datasets: Using real-person head portraits instead of synthetic face datasets to enhance the cross-dataset adaptability of the model and its reliability in practical applications.

[0193] Method for generating a head pose dataset using the technical solution of the present invention with Dynamic Gaussian Avatar: A method for generating a head pose dataset at a wide range of angles using the precise rendering and pose control of Dynamic Gaussian Avatar.

[0194] System for automatically generating diverse head pose data: A system that automatically generates diverse head pose data covering different angles, lighting, and backgrounds by adjusting Gaussian model parameters, thereby reducing manual annotation errors and improving generation efficiency.

[0195] Method for constructing a head pose dataset based on real-person head modeling: A method for generating a diverse head pose dataset by modeling real-person heads and combining background and lighting processing technologies (such as BackgroundMattingV2 and Deep Single-Image Portrait Relighting).

[0196] Method for generating exclusive dataset for personalized head pose detection model: A method for creating an exclusive head pose dataset based on specific person images or videos and using it to train a personalized head pose detection model.

[0197] Process for generating low-cost head pose dataset using open-source technology: A process for completing dataset production using free and open-source tools, avoiding commercial 3D asset tools, thereby significantly reducing production costs.

[0198] Application of dynamic Gaussian rendering technology in pose diversity and robustness: A method for simulating various head movement scenarios through DynamicGaussian Avatar rendering technology to generate a head pose dataset with high pose diversity and robustness, enhancing the performance of the detection model in complex scenarios.

[0199] In a specific embodiment, for S6, the head pose in a real video is annotated using a video annotation tool, the pose parameters are extracted, and the annotated data is used as the ground truth. Its advantages: The generated data is real video data, with good realism and generalization ability. Its disadvantages: The annotation process of video data is cumbersome, and it is difficult to precisely control the range and angle of the pose. In addition, it is impossible to control large angles, different lighting, and backgrounds, and the diversity of the dataset is limited.

[0200] In a specific embodiment, for S1, S4, S5, S6, the FLAME (Face Learned with Expressions and Pose) or a similar 3D face model is used to generate head poses, and different expressions and poses are controlled by optimizing the 3D mesh and parameters. Its advantages: The FLAME model can better represent the changes in head poses and expressions, and existing research has used this model to generate head pose data. Its disadvantages: The FLAME model is not as good as dynamic Gaussian rendering in terms of rendering accuracy and high-resolution performance, and is limited in terms of control details and continuity. In addition, the dataset generated using FLAME may lack the effect of real person modeling, and its generalization ability may be inferior to that of real modeling data.

[0201] In a specific embodiment, for S4, S5, S6, S7, S8, traditional 3D modeling and rendering software (such as Maya, 3dsMax, or Blender) is used to manually create head models with different poses and expressions, and each pose is manually rendered and annotated. Its advantages: 3D modeling software has strong flexibility and can precisely control poses, expressions, and lighting conditions. Its disadvantages: The process of generating the dataset highly depends on manual operations, with low efficiency and high costs. The rendering speed of traditional 3D software is slow, and it is difficult to automatically generate large-scale data, resulting in much higher production costs and time costs than automated methods.

[0202] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A head posture data collection method based on a dynamic Gaussian head avatar, characterized in that: The following steps are involved: S1, based on the NeRSemble dataset, there are at least 267 identities in total, and at least 30 identities are selected to extract videos; S2, use BackgroundMattingV2 to remove the original background; S3, expression parameters and landmarks were obtained using Multiview-3DMM-Fitting; S4, uses the improved Gaussian head portrait model based on Gaussian splashing for head modeling training; S5, after the head model is obtained after training, each identity takes at least three photos with different expressions to achieve expression diversity; S6, uses Gaussian head avatars to render head pose avatar images with head poses ranging from -90 degrees to 90 degrees, and makes a normal distribution of head pose angles; S7, insert the BackgroundMattingV2 model to replace the background image to achieve background diversity, and use at least three different types of background images: pure color background, texture background, and real life scene background; S8, put in Deep Single-Image Portrait Relighting to relight, change the lighting, use at least 7 types of lighting to achieve lighting diversity; S9, the obtained synthetic head posture data set is preprocessed by dockerface to calibrate the face bounding box, and then put into the head posture detection model Hopenet training, and then preprocessed by multi-task convolutional neural network to calibrate the face bounding box, and then put into the head posture detection model FSA-Net for training; S10, testing on the test head pose dataset.

2. The head posture data collection method based on a dynamic Gaussian head portrait according to claim 1, characterized in that: The Gaussian head portrait model in S4 uses a controllable dynamic 3D Gaussian distribution to represent and model the human head based on the Gaussian splash method.

3. The head posture data collection method based on a dynamic Gaussian head portrait according to claim 1, characterized in that: The S4 specifically comprises the following steps: S41, dynamic Gaussian head representation; S42, geometric deformation; S43, dynamic color, rotation, scale, and opacity representation; S44, training loss function; S45, geometry-guided initialization.

4. The head posture data collection method based on a dynamic Gaussian head portrait according to claim 3, characterized in that: The S41 specifically includes: the static 3D Gaussian points include N points, which are defined by their position X, multi-channel color C, rotation Q in quaternion form, scale S and opacity A; The Gaussian points are rasterized and rendered through the camera parameters μ to obtain a multi-channel image I. This process is expressed as: I=R(X,C,Q,S,A;μ) Construct a standard neutral Gaussian model with position X0∈R N×3 , feature vector F0∈R N×128 , rotate Q0∈R N×4 , scale S0∈R N×3 and opacity A0∈R N×1 properties, all of which are fully optimizable; The MLP-based dynamic generator Φ is used to generate dynamic changes related to expression. The input is the neutral model attribute, expression coefficient θ and head posture β, and the output is the final dynamic model attribute {X, C, Q, S, A}. The formulation of the entire dynamic Gaussian head avatar model is: {X,C,Q,S,A}=Φ(X0,F0,Q0,S0,A0;θ,β).

5. The head posture data collection method based on dynamic Gaussian head portrait according to claim 4, characterized in that: The S42 specifically includes: using two different MLPs to respectively control the deformation of the expression and the head posture, and the displacement calculation formula of each Gaussian point is: Among them, two MLPs are used, Used to define the displacement caused by the expression, ∈Φ is used to define the displacement caused by the posture, def represents the meaning of displacement, λ exp (X0) and λ pose (X0) represents the influence of expression and head posture on each Gaussian point.

6. The head posture data collection method based on dynamic Gaussian head portrait according to claim 5, characterized in that: The S43 specifically includes: The calculation formula of dynamic color C' is: Color is predicted directly by two color MLPs, Used to define dynamic colors that change with expressions. Used to define dynamic colors that change with posture, col represents the meaning of color; Changes in rotation, scale, and opacity are expressed as: Through two MLP predictions, Used to define rotation, scale, and opacity properties that change with expressions. Used to define the rotation, scale and opacity attributes that change with the posture. att represents the meaning of rotation, scale and opacity. Finally, the above properties are converted to world coordinates: {X,Q}=T({X′,Q′},β) {C,S,A}={C′,S′,A′}.

7. The head posture data collection method based on dynamic Gaussian head portrait according to claim 6, characterized in that: The S44 specifically includes: During training, a 3D Gaussian model is generated under the expression condition, then rendered as a 32-channel image, and finally a high-resolution head image is generated through a super-resolution network; the training loss function includes L1 loss, VGG perception loss and RGB channel loss, and the total loss function is expressed as: L=||I hr -AND gt ||1+λ vgg VGG(I hr ,AND gt )+λ lr ||And lr -AND gt ||1; Among them, I hr To generate high-resolution images, I gt is a real high-resolution image, |||I hr -I gt ∣∣1: Generate high-resolution image I hr and the real high-resolution image I gt The L1 loss between vgg To control the weight hyperparameters of the VGG perceptual loss term, VGG(I hr ,I gt ) is the perceptual loss term, and the VGG network is used to calculate the generated image I hr and the real image I gt The difference in feature space, λ lr To control the weight hyperparameter of the low-resolution L1 loss term, I lr is the generated low-resolution image, ||I lr -I gt ||1 is the generated low-resolution image I lr and the real image I gt The L1 loss.

8. The head posture data collection method based on dynamic Gaussian head avatar according to claim 7, characterized in that: The S45 specifically includes: First, an MLP representing the signed distance field is defined: s,η=f sdf (x) Where s represents the value of the signature distance field, η represents the feature vector of each point, and x is the position of the point; RGB Loss L RGB It is expressed as: L RGB =||I r,g,b -I gt ||1 Among them, I r,g,b Represents the RGB channels of the generated image, I gt is a real RGB image, ||·||1 represents the L1 norm; Contour loss L sil It is expressed as: L sil =IOU(M,M gt ) Among them, M is the mask of the generated image, M gt is the mask of the real image, and IOU(·,·) is used to calculate the overlap between the two masks.

9. The head posture data collection method based on dynamic Gaussian head portrait according to claim 1, characterized in that: S6 specifically includes: the head posture generation process starts from the elevation angle and the azimuth angle to create a basic head rotation matrix R, and rotates the head from the initial position to the desired viewing angle. This rotation is achieved by the classic rotation matrix formula: The initial head rotation matrix R is calculated as the product of these two rotation matrices; After constructing the basic rotation matrix, the roll angle is added to control the rotation of the head around its own axis. The final head rotation matrix is ​​obtained by multiplying it with the roll rotation matrix. This rotation matrix is ​​combined with the translation vector T to form the camera's external parameter matrix to describe the camera's position and orientation in the world coordinate system; To transform the head pose from the world coordinate system to the camera coordinate system, a view transformation matrix is ​​calculated, which combines the rotation matrix R final and a translation vector T to describe the transformation from the world coordinate system to the camera coordinate system: In order to project the 3D head pose onto the 2D image plane, a projection transformation matrix is ​​required. The projection matrix converts 3D coordinates into 2D coordinates based on the camera's field of view and other projection parameters. The combination of the view transformation matrix and the projection matrix produces a complete projection transformation matrix, which is used to describe the complete conversion process from 3D head pose to 2D image plane: FullProj=ProjMatrix×Extrinsic This process ultimately generates a complete projection transformation matrix.

10. The head posture data collection method based on dynamic Gaussian head avatar according to claim 1, characterized in that: The S9 includes: in the process of generating the synthetic head posture data set, recording three key angle parameters of each posture: pitch angle, yaw angle and roll angle, these posture parameters are directly extracted in the rendering process and used as true value annotation; In order to verify the accuracy of the true value data, three head pose estimation models are used: Hopenet, 6DRepNet and FSA-Net.

Citation Information

Patent Citations

  • Sight line area estimation method and system based on head posture and space attention

    CN113361441A

  • Three-dimensional reconstruction method and system for indoor scene

    CN116797742A

  • Virtual anchor generation method and system based on neural radiation field and hidden attributes

    CN117171392A

  • Nerve relighting method applied to mixed reality flight simulator

    CN118429583A