Real-time online construction method for dynamic three-dimensional gaussian human head model

WO2026174413A1PCT designated stage Publication Date: 2026-08-27ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/077746
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2026-08-27

Smart Images

  • Figure CN2025077746_27082026_PF_FP_ABST
    Figure CN2025077746_27082026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present invention is a real-time online construction method for a dynamic three-dimensional Gaussian human head model. By means of the method, batches of three-dimensional Gaussian rendering is performed, so that the training throughput is greatly improved, and a model achieves an almost real-time convergence rate; in addition, a local-global combined data sampling algorithm applicable to online modeling is designed, wherein newly inputted training samples are rapidly fitted by means of local sampling, and global sampling is used to alleviate the problem of model forgetting, thereby achieving a reconstruction quality similar to that of offline modeling. Compared with conventional methods, the present invention implements real-time online construction of a three-dimensional Gaussian face model without any preprocessing or post-processing. In addition, online modeling can acquire data and also provide real-time feedback to users, thereby providing a better user experience. After reconstruction, the model can be used for synthesizing high-fidelity human head animations with new expressions; and the present invention can be widely applied in film and animation production, online games and remote conferences.
Need to check novelty before this filing date? Find Prior Art

Description

A method for real-time online construction of a dynamic 3D Gaussian model of the human head Technical Field

[0001] This invention relates to the fields of real-time online 3D reconstruction, facial motion capture, and real-time animation technology, and particularly to a real-time online construction method for a dynamic 3D Gaussian model of the human head. Background Technology

[0002] In recent years, 3D Gaussian Splatting (3DGS) (Kerbl B, Kopanas G, Leimkühler T, et al. 3D Gaussian Splatting for Real-Time Radiance Field Rendering[J]. ACM Trans. Graph., 2023, 42(4): 139: 1-139: 14.) has made significant progress in scene reconstruction and inspired a series of related studies. MonoGaussianAvatar (Chen Y, Wang L, Li Q, et al. Monogaussianavatar: Monocular gaussian point-based head avatar[C] / / ACM SIGGRAPH 2024Conference Papers. 2024: 1-9.) uses 3D Gaussian to replace the point cloud in PointAvatar to improve rendering quality. GaussianHead (Wang J, Xie JC, Li X, et al. Gaussianhead: Impressive head avatars with learnable gaussian diffusion[J]. arXiv preprint arXiv:2312.01632,2023.) utilizes a three-plane storage to store appearance-related attributes and achieves animation control through an MLP-based deformation field. Gaussian Head Avatar (Xu Y, Chen B, Li Z, et al. Gaussian head avatar: Ultra high-fidelity head avatar via dynamic gaussians[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition.2024:1931-1941.) achieves high-fidelity head avatar rendering through a super-resolution network. Based on the interpretability of 3DGS, some studies have bound Gaussians to template meshes to achieve more convenient head expression control.Among them, GaussianAvatars (Qian S, Kirschstein T, Schoneveld L, et al. Gaussianavatars: Photorealistic head avatars with rigged 3d gaussians[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition.2024:20299-20309.) achieves better image alignment by jointly optimizing the mesh. FlashAvatar (Xiang J, Gao X, Guo Y, et al. FlashAvatar: High-fidelity Head Avatar with Efficient Gaussian Embedding[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition.2024:1802-1812.) uses MLP to add dynamic spatial offsets to Gaussians. Another type of method uses a set of mixed bases to model the head avatar. For example, HeadGaS (Dhamo H, Nie Y, Moreau A, et al. Headgas: Real-time animatable head avatars via 3d gaussian splatting[C] / / European Conference on Computer Vision. Springer, Cham, 2025:459-476.) blends learnable latent features and decodes the color and opacity of Gaussians using an MLP. GEM compresses the reconstructed Gaussian head image into a feature base. GaussianBlendshapes (Ma S, Weng Y, Shao T, et al. 3d gaussian blendshapes for head avatar animation[C] / / ACM SIGGRAPH 2024 Conference Papers.2024:1-10.) directly blends all Gaussian properties and ensures semantic consistency with the mesh blend shape, thus achieving high reconstruction quality and rendering performance. Although these methods can reconstruct highly realistic face models, they all require offline reconstruction after acquiring video data, limiting their practical application scenarios.

[0003] Weise et al. (Weise T, Bouaziz S, Li H, et al. Realtime performance-based facial animation[J]. ACM transactions on graphics(TOG),2011,30(4):1-10.) first achieved real-time facial expression capture by fitting a parameterized hybrid shape model based on RGB-D data from a consumer-grade depth sensor. Subsequent research focused on shape correction (Li H, Yu J, Ye Y, et al. Realtime facial animation with on-the-fly correctives[J]. ACM Trans. Graph., 2013, 32(4):42:1-42:10.), dynamic facial expression space (Bouaziz S, Wang Y, Pauly M. Online modeling for realtime facial animation[J]. ACM Transactions on Graphics(ToG), 2013, 32(4):1-10.), and non-rigid mesh deformation (Chen YL, Wu HT, Shi F, et al. Accurate and robust 3d facial capture using a single rgbd camera[C] / / Proceedings of the IEEE International Conference on Computer Vision. 2013:3615-3622.). Although these works demonstrate impressive results, they rely on depth data, which is typically unavailable in most video footage. Cao et al. (Cao C, Weng Y, Lin S, et al. 3D shape regression for real-time facial animation[J]. ACM Transactions on Graphics(TOG),2013,32(4):1-10.) proposed a real-time facial animation system that fits a hybrid shape model from two-dimensional video frames.DDE (Cao C, Hou Q, Zhou K. Displaced dynamic expression regression for real-time facial tracking and animation[J]. ACM Transactions on graphics(TOG),2014,33(4):1-10.) uses a universal regressor, eliminating the need for calibration for each user. Although these methods can track and reconstruct facial models in real time, they can only use three-dimensional meshes, lacking realism. Summary of the Invention

[0004] This invention addresses the shortcomings of existing technologies by proposing a real-time online construction method for a dynamic 3D Gaussian model of the human head. This method is suitable for general users and requires no special preprocessing for specific users. Any new user can directly use it in their daily life environment using a home computer and a webcam. This invention acquires a facial video stream through a monocular camera. After removing the background from the video frames and tracking the head motion parameters, this method can optimize the 3D Gaussian model of the human head online and achieve real-time convergence. After modeling, this method can drive the reconstructed 3D Gaussian model of the human head in real time using given head motion parameters to synthesize high-fidelity facial animations. This invention can be widely applied in film and animation production, video chat, online games, and other applications, possessing broad applicability.

[0005] This invention is achieved through the following technical solution:

[0006] A method for real-time online construction of a dynamic 3D Gaussian model of the human head includes the following steps:

[0007] (1) Initialization: Calibrate the intrinsic parameters of the monocular camera, initialize the dynamic three-dimensional Gaussian model of the human head; create a local sampling cache and a global sampling cache to store online training data samples, and keep the cache size constant during system operation;

[0008] (2) Online video stream input image acquisition and processing: Use a monocular camera to capture face video, calculate the head foreground mask and head motion parameters for the input image; store them in the local sampling buffer and global sampling buffer created in step (1);

[0009] (3) Initialize 3D Gaussian color: For the first few frames of input image, the pixel color is projected onto the 3D Gaussian according to the camera parameters to accelerate the initialization of the color attributes of the 3D Gaussian and accelerate the convergence speed in the early stage of training.

[0010] (4) Online data sample sampling: Based on the number of samples already in the local sampling cache and the global sampling cache during system operation, the local data batch and the global data batch are adaptively calculated; further, the local data batch and the global data batch are sampled from the local and global sampling caches respectively, and finally the two are combined to obtain the training data batch, which is used for the current iteration optimization step;

[0011] (5) Batch parallel optimization of dynamic 3D Gaussian model: Based on the training data batch obtained in step (3), firstly, the human head dynamic 3D Gaussian model is driven in batches according to the human head motion parameters to obtain 3D Gaussian under different motion parameters; secondly, the 3D Gaussian is rendered in batches; then the rendering results are compared with the real images in the training data batch to calculate the loss; finally, the optimization parameters of the human head dynamic 3D Gaussian model are updated iteratively through image loss.

[0012] (6) Human head animation generation: Based on the human head motion parameters and camera perspective given by the user, the modeled human head dynamic 3D model is driven and the head animation under new perspective and expression is generated in real time.

[0013] Specifically, the training samples to be stored in the local sampling cache and the global sampling cache in step (1) include human head motion parameters, real images and human head foreground masks; and the capacity of the local sampling cache and the global sampling cache remains unchanged after creation and will not increase indefinitely as the input data increases. During system operation, the local sampling cache is maintained in the form of a first-in-first-out queue, and the global sampling cache is maintained in the form of a reservoir sampling method.

[0014] Furthermore, step (2) specifically includes the following sub-steps:

[0015] (2.1) Maintain the global sampling buffer by sampling from a reservoir. When the local sampling buffer is full, the last data sample D in the local sampling buffer will be stored. j Take it out, with The probability of D j Store it in the global sampling buffer, where j is the sample number that was retrieved. The current capacity of the global sampling cache; when D j Stored At that time, randomly discard one The sample in the reservoir sampling ensures that each piece of data in the data stream has an equal probability of remaining in the cache;

[0016] (2.2) Maintain the local sampling buffer using a first-in-first-out queue. When the local sampling buffer is full, the last sample is retrieved and stored in D. i Otherwise, store directly in D.i .

[0017] Furthermore, in step (2), the convergence speed in the early stage of training is accelerated by using the batch parallel optimization method of the dynamic three-dimensional Gaussian model of the human head to accelerate convergence; specifically, the batch rendering of the three-dimensional Gaussian model is used to improve the training throughput, so that the model can achieve a real-time convergence rate and can be applied to online modeling.

[0018] Further, step (3) specifically involves: first, projecting the three-dimensional Gaussian kernel onto a two-dimensional image using camera parameters; then, performing convolution calculations between the two-dimensional Gaussian kernel and the pixel colors of the input image; and finally, assigning the calculation result to the color attribute of the Gaussian kernel, the expression of which is as follows:

[0019] Where W and H are the width and height of the image, respectively, and I... xy w represents the RGB color value of the (x,y) pixel in the image. xy The weights of Gaussian splatting are applied to the (x,y) pixels of the image.

[0020] Furthermore, step (4) includes the following sub-steps:

[0021] (4.1) Based on the ratio of the sum of the sampling probability densities of the current local sampling cache and the global sampling cache, the training data batch is adaptively divided to ensure that the probability of sampling from the global sampling cache is less than or equal to the probability of sampling from the local cache.

[0022] (4.2) Randomly sample local training data batches from the local sampling buffer;

[0023] (4.3) Sample global data batches from the global sampling buffer with importance based on the length of time the samples have been added to the buffer. The mathematical formula for calculating the sampling density is as follows: W(k) = exp(w l ·k / |M l |);

[0024] That is, when the input video frame is the i-th frame, Sample D of the j-th frame j The probability of being sampled, where For the present The number of samples in w l The sampling coefficients w are manually set. l The larger the value, the higher the sampling weight assigned to the new sample.

[0025] Specifically, in step (4), online data sample sampling is achieved by using a local-global combined data sample sampling algorithm suitable for online modeling. First, local sampling is used to quickly fit the new input training samples, and then global sampling is used to alleviate the model forgetting problem, thereby achieving a reconstruction quality similar to that of offline modeling.

[0026] Furthermore, the batch rendering of the three-dimensional Gaussian in step (5) specifically includes the following sub-steps:

[0027] (5.1) Batch preprocessing of 3D Gaussian: Based on the camera parameters, each 3D Gaussian is projected onto a 2D plane and the number of Gaussian instances is recorded. Each group of Gaussians in the 3D Gaussian batch is calculated on a different CUDA Stream.

[0028] (5.2) GPU-CPU synchronization: After all CUDA Stream calculations in step (5.1) are completed, perform a GPU-CPU synchronization to transfer the number of 3D Gaussian instances to the CPU and create an array of the corresponding size on the GPU to store the screen space depth of the Gaussian instance, the corresponding pixel block number and the corresponding 3D Gaussian number.

[0029] (5.3) 3D Gaussian Batch Rasterization: Based on the projected 3D Gaussian parameters obtained in step (5.1) and the 3D Gaussian instances obtained in step (5.2), the contribution value of each Gaussian instance to the color of the current pixel is calculated in the order of depth and the transparency is mixed to obtain the color of the current pixel; each group of Gaussians in the 3D Gaussian batch is calculated on a different CUDA Stream, and finally the batch image after rendering is obtained.

[0030] The beneficial effects of this invention are as follows:

[0031] This method can model 3D Gaussian faces online without any preprocessing or post-processing, simplifying the data acquisition process. Furthermore, online modeling provides real-time feedback to users during data acquisition, making it particularly suitable for ordinary users. The reconstructed model can be used for high-fidelity facial animation synthesis. Compared to previous methods, this method significantly shortens the face modeling process, providing a better user experience; the entire process of modeling a person's face can be completed in approximately 3 minutes. Attached Figure Description

[0032] Figure 1 is a flowchart of the operation after the initialization steps of the operating system of the present invention;

[0033] Figure 2 is a system operation diagram after the rapid initialization step of the three-dimensional Gaussian color in this invention;

[0034] Figure 3 is a system operation diagram of the modeling process of this invention;

[0035] Figure 4 is a system operation diagram of the generated human head animation after the modeling of the present invention is completed;

[0036] Figure 5 is a screenshot of the system running the generated human head animation after the modeling of the present invention is completed. Detailed Implementation

[0037] The core of this invention lies in real-time online modeling of human face models. A monocular video stream is input, and a dynamic three-dimensional Gaussian model of the human head is optimized in real time. After reconstruction, it can be used to generate human head images and animations from new perspectives and with different expressions.

[0038] The dynamic 3D Gaussian model of the human head used in this invention is a Gaussian blend shape. The Gaussian blend shape consists of a neutral Gaussian model B0 and 50 base expression Gaussian models {B1, B2, ..., B...} corresponding to basic facial expressions. 50 The above models are all composed of a set of three-dimensional Gaussians. Each Gaussian has the following basic properties: position x∈R 3 Opacity α∈R, rotation q (quaternion representation), size s∈R 3 Spherical harmonic coefficients SH∈R 16×3 ; The Gaussian model B0 of neutral Gaussian model and the Gaussian model B of each base expression k There is a one-to-one correspondence between the Gaussians in the equations. A Gaussian model of a human face with any expression can be calculated using the following formula:

[0039] Where ΔB k =B k -B0,{ψ k} represents the expression coefficient for each base expression. B ψ Images can be rendered using Gaussian splatting.

[0040] Based on the aforementioned Gaussian mixture shape, this invention discloses a real-time online construction method for a dynamic three-dimensional Gaussian model of the human head (including but not limited to Gaussian mixture shape), specifically including the following steps: system initialization, online video stream input image acquisition and processing, rapid initialization of three-dimensional Gaussian color, local-global combined online data sample sampling, and batch parallel optimization of the dynamic three-dimensional Gaussian model.

[0041] The following will be explained in detail with reference to Figure 1 (Figure 1 is a screenshot of the system after the initialization steps of the present invention; the left side of Figure 2 shows the control panel and the input video frame; the right side of Figure 2 shows the rendering result of the dynamic three-dimensional Gaussian model of the human head after rapid color initialization).

[0042] 1. System Initialization

[0043] 1.1 Initialize the camera and sampling buffer

[0044] Before modeling begins, the camera intrinsics need to be calibrated, and then a local sampling buffer needs to be created. With global sampling cache These caches are used to store training data samples during the online modeling process. In the specific implementation, these two caches can store 150 and 1000 data samples respectively, and their size remains constant during system operation.

[0045] 1.2 Initialize Gaussian Mixture Shape

[0046] For the neutral Gaussian model B0, Poisson disk sampling is used to uniformly sample points on the neutral facial expression mesh M0 of the human head, which are then used as the initial positions of the Gaussians. The Gaussians are initialized to be isotropic, specifically: the size s is initialized to the average distance to the centers of the three nearest-neighbor Gaussians, and the rotation q is initialized to a unit quaternion. Based on 50 basic facial expression Gaussian models {B... k}, extract the deformation gradient from the FLAME base expression mesh, that is, from the neutral expression mesh M0 of the human head to each base expression mesh M. k The deformation is calculated, and the deformation gradient is applied to the neutral Gaussian model B0 to obtain the base expression Gaussian model {B}. k}, as initialization. The number of 3D Gaussians initialized is 60k.

[0047] 2. Online video stream input image acquisition and processing

[0048] 2.1 Online Video Stream Image Acquisition

[0049] When a user enters the camera's field of view, the system is activated. The ambient lighting is required to remain relatively constant during filming, and the user must make different head poses and facial expressions. Each time the camera inputs a frame, training data samples are first obtained through the method in step 2.2, and then the cache is updated according to steps 2.3 and 2.4. and

[0050] 2.2 Input Image Processing

[0051] For each frame of input image I iThis invention uses Robust Video Matting (Lin S, Yang L, Saleemi I, et al. Robust high-resolution video matting with temporal guidance[C] / / Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision.2022:238-247.) to calculate the head foreground mask to remove the image background. At the same time, it uses DDE (Cao C, Hou Q, Zhou K. Displaced dynamic expression regression for real-time facial tracking and animation[J]. ACM Transactions on graphics(TOG),2014,33(4):1-10.) to track the motion parameters of the human head. Then, the motion parameters, the image, and the foreground mask are combined to form a training data sample D. i .

[0052] 2.3 Update the local sampling cache

[0053] The local sampling buffer is maintained using a first-in, first-out (FIFO) queue: the training data samples D of the current frame are... i Add local sampling buffer like When full, then The last data sample D in j Remove it, freeing up storage space for one sample, and then put D... i Put in like If the storage is not full, directly save D. i join in That's it; step 1.4 is not required.

[0054] 2.4 Update global sampling cache

[0055] The global sampling cache is maintained using a reservoir sampling method: it will be from... Sample D was taken out j Add to global sampling buffer like If the buffer is full, a random integer k is selected from [0,j). If k is less than the size of the global sampling buffer... Then Discard the k-th data in D and set D jPlace it in an empty space, otherwise place D. j Discard. This is equivalent to discarding. The probability of D j Storing samples in a global sampling buffer ensures that each sample in the video stream has an equal probability of remaining in the buffer. middle.

[0056] 3. Fast initialization of 3D Gaussian color

[0057] As shown in Figure 2 (the left side of Figure 2 shows the control panel and the input video frames, and the right side shows the rendering result of the dynamic 3D Gaussian model of the human head after rapid color initialization), for the first few input frames, the pixel colors in the images can be mapped to the unoptimized 3D Gaussian model based on the camera parameters, thereby quickly initializing the Gaussian color (i.e., 0th order spherical harmonic) and accelerating the convergence speed. The specific implementation steps are as follows: Project the 3D Gaussian model onto the 2D model using the same calculation method as in step 4.1, and perform convolution calculation with the input image using the 2D Gaussian kernel to obtain the estimated color c of each Gaussian model. init

[0058] Where W and H are the width and height of the image, respectively, and I... xy w represents the RGB color value of the (x,y) pixel in the image. xy This represents the weights of the Gaussian splatter on the (x,y) pixels of the image. This color initialization operation is only executed when the weights of the Gaussian splatter first exceed the threshold δ = 0.1; that is, Gaussians with poor visibility will not be color initialized. This operation is also executed only once, providing a relatively accurate initial value for subsequent gradient optimization and avoiding interference with the optimization process.

[0059] 4. Online data sampling combining local and global methods

[0060] As shown in Figure 3 (the left side of Figure 3 shows the control panel and the input video frame, and the right side shows the rendering result of the dynamic 3D Gaussian model of the human head during the modeling process), in traditional offline modeling schemes, since the entire video data has already been preprocessed, the training process only requires randomly selecting a sample from the dataset, driving and obtaining the Gaussian model of that frame, rendering it into an image, calculating the loss between it and the real image, and finally backpropagating the gradient to update the model parameters, and so on iteratively. However, in online modeling, since the input data is streamed, random sampling cannot obtain high-quality results. New data samples have the fewest opportunities for optimization, so more sampling weights need to be allocated, while optimized data cannot be discarded and still needs to participate in the optimization to prevent the model from forgetting. Therefore, this invention proposes a local-global combined online data sample sampling algorithm, which iteratively optimizes the face model based on the existing data samples in the cache while processing the input data, achieving a reconstruction quality similar to offline modeling.

[0061] 4.1 Local-Global Sampling

[0062] For each iteration step in the online optimization process, this invention uses local sampling caches. and global sampling cache Medium sampling local sample batch and global sample batch Merge Then, batch parallel optimization is performed (see step 5). For global sample batches... From global sampling cache Random sampling is sufficient. However, for local sample batches... To assign more sampling weights to new samples, the probability distribution is as follows: Medium sampling W(k) = exp(w l ·k / |M l |)

[0063] This means that when the input video frame is the i-th frame, Sample D of the j-th frame j The probability of being sampled, where For the present The number of samples in w l The sampling coefficients w are manually set. l The larger the value of w, the higher the sampling weight assigned to the new sample. In the specific implementation, w l =1.0.

[0064] 4.2 Divide the data into batches

[0065] During system operation, the data batch size during optimization and and The number of samples will increase dynamically, from and The probability density ratio of the mid-sample will also change dynamically, so adaptive calculation is required. and The division ratio η i (Right now occupy (the proportion) to ensure that from The probability of sampling any sample from the middle is less than that from the middle. The probability of sampling any sample from the middle. The specific calculation is as follows:

[0066] Where η max To determine the maximum value of the specified division ratio, in the specific implementation, η max =0.7. After obtaining the division ratio η i Then, the specific data batch partitioning process is as follows: Calculation A number of random numbers between 0 and 1, of which the number is less than η i The number of is The quantity, and vice versa The quantity. After determining... After the initial sampling, the batch to be optimized in the current iteration step can be obtained according to step 4.1.

[0067] 5. Batch Parallel Optimization of 3D Gaussian Models

[0068] This invention optimizes a batch of data samples in each iteration step, thereby significantly improving training throughput and enabling the model convergence speed to be applied to online real-time modeling. For each iteration step, the batch of data samples to be optimized is obtained from step 4. Firstly, according to A Gaussian model driven by parallel computation of face pose and expression parameters in medium batches is then rendered in batches; subsequently, the rendering results are compared with... The loss is calculated using real images; finally, the optimization parameters of the 3D Gaussian face model are iteratively updated using image loss. In batch optimization, the core of this invention lies in batch parallel rendering of 3D Gaussians, the specific steps of which are as follows:

[0069] 5.1 Batch Preprocessing of 3D Gaussian Models

[0070] Each model in a batch 3D Gaussian model consists of a set of 3D Gaussians. The following steps describe the calculation process for each set of 3D Gaussians, which are executed on different CUDA Streams.

[0071] First, parallel computation is performed for each 3D Gaussian: the 3D Gaussian is projected onto a 2D model using camera parameters, with the following mathematical formula: Σ′=JWΣW T J T x′=PWx

[0072] Where Σ′ and Σ are the 3D covariance matrices before and after Gaussian projection, respectively; x′ and x are the position coordinates in the 3D world space before Gaussian projection, the coordinates on the 2D image after projection, and the depth in screen space, respectively. W is the camera's View matrix, P is the camera projection matrix, and J is the Jacobian matrix estimated by the affine transformation of P. 3D The Gaussian rotation (q) and scale (s) can be calculated using the 3D Gaussian. Further, based on the 2D covariance matrix Σ′, the lengths of the minor and major axes of the 2D Gaussian can be calculated through eigenvalue decomposition (mathematically, the Gaussian distribution has values ​​throughout the entire space, but in practice, the influence of the Gaussian on values ​​beyond three standard deviations can generally be ignored). The length of the major axis allows calculation of the enclosing circle of the Gaussian in screen space, thus determining how many pixel tiles the Gaussian contacts. Each Gaussian in contact with a pixel tile is recorded as an "instance". Finally, the Gaussian spherical harmonics is converted to RGB color according to the camera's viewing direction. Therefore, this stage of calculation yields the following data: Gaussian screen space depth, Gaussian radius (and major axis length), 2D coordinates of the Gaussian on the image, 2D covariance matrix of the Gaussian, RGB color of the Gaussian, and the number of generated Gaussian instances.

[0073] Then, a prefix sum is calculated on the array storing the number of all Gaussian instances. The last data in the prefix sum array is the sum of the number of Gaussian instances generated by all Gaussian instances. This sum is transferred from GPU storage to memory for use by the CPU in subsequent steps.

[0074] 5.2 GPU-CPU Synchronization

[0075] After step 5.1 completes all batch computations and data transfers on different CUDA Streams, allocate GPU storage of the corresponding size based on the sum of the number of Gaussian instances in each group.

[0076] 5.3 Batch Rasterization of 3D Gaussian Models

[0077] Each model in a batch 3D Gaussian model consists of a set of 3D Gaussians. The following steps describe the calculation process for each set of 3D Gaussians, which are executed on different CUDA Streams.

[0078] First, parallel computation is performed for each 3D Gaussian: based on the Gaussian radius and its 2D coordinates on the image, each pixel block in contact with the Gaussian, i.e., each Gaussian instance, is recalculated. Simultaneously, a key-value pair is calculated for each Gaussian instance: the key is 64 bits long, with the high 32 bits representing the pixel block number in contact and the low 32 bits representing the Gaussian's depth in screen space; the value is the Gaussian's index. Then, the key-value pair is stored in the allocated GPU memory.

[0079] Then, based on the key value, the key-value pairs are sorted using radix sort. In the sorted array, Gaussian instances of the same pixel block are stored in a contiguous block, and the Gaussians within the same pixel block are arranged from lightest to darkest screen depth. After sorting, parallel computation is needed for each 3D Gaussian instance to determine the start and end positions of the same pixel block within the array.

[0080] Finally, parallel computation is performed for each image pixel, with threads belonging to the same thread group (Block): First, the attributes of the Gaussian instances within a pixel block are loaded into the thread group's shared memory. After loading, a synchronization is performed within the thread group. Then, each Gaussian instance in the shared memory is traversed in depth-order, and the color splattered by each Gaussian instance is calculated. The mathematical formula is as follows:

[0081] Where x p Let be the 2D coordinates of the pixel, and 'c' be the RGB color of the Gaussian pixel. After obtaining the Gaussian splash color, it is blended with the pixel's own color using transparency. If the cumulative opacity of the pixels reaches 1.0, the loop terminates prematurely. The entire image is rendered after calculating the color of each pixel.

[0082] Human head animation generation. As shown in Figure 4, the left side of Figure 4 shows the control panel and the input video frames, while the right side shows the animated image synthesized using the same user's head motion parameters to drive a dynamic 3D Gaussian model of the human head. That is, the completed 3D dynamic Gaussian model of the human head (Gaussian mixture shape) can be used to synthesize new images and animations from a given viewpoint and expression. The user only needs to provide the human head motion parameters ψ and camera parameters to generate the Gaussian model B for each frame. ψ Furthermore, Gaussian splashing technology is used to create highly realistic images and animations. Parameters can be manually edited by the user or obtained from any human head video using a face tracker.

[0083] Example 1

[0084] The inventors implemented an example of this invention on a desktop computer equipped with an Intel Core i7-13700KF (5.40GHz) CPU and an NVIDIA RTX 4090 GPU, and a webcam providing 1280×960 resolution at 30 frames per second. During operation, approximately 50 milliseconds are required for each frame to calculate the human head foreground mask and track the human head motion parameters. The throughput of the 3D Gaussian model training reached approximately 320 samples per second, enabling the system to model the human head in real time.

[0085] The inventors invited various new users to test the system based on this invention. The results showed that the system can model a dynamic 3D Gaussian head model online in real time for any new user without any pre-processing or post-processing. As shown in Figure 5, the left side of Figure 5 shows the control panel and the input video frames, while the right side shows the animated images rendered and synthesized from the dynamic 3D Gaussian head model driven by different users' head motion parameters. That is, after modeling is completed, the model can be driven in real time to obtain human facial images and animations under different head postures and expressions. (Furthermore, all human portraits used in the accompanying drawings have been used with the consent of the individuals involved.)

[0086] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0087] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A method for real-time online construction of a dynamic three-dimensional Gaussian model of the human head, characterized in that, Includes the following steps: (1) Initialization: Calibrate the intrinsic parameters of the monocular camera, initialize the dynamic three-dimensional Gaussian model of the human head; create a local sampling cache and a global sampling cache to store online training data samples, and keep the cache size constant during system operation; (2) Online video stream input image acquisition and processing: Use a monocular camera to capture face video, calculate the head foreground mask and head motion parameters for the input image; store them in the local sampling buffer and global sampling buffer created in step (1); (3) Initialize 3D Gaussian color: For the first few frames of input image, the pixel color is projected onto the 3D Gaussian according to the camera parameters to accelerate the initialization of the color attributes of the 3D Gaussian and accelerate the convergence speed in the early stage of training. (4) Online data sample sampling: Based on the number of samples already in the local sampling cache and the global sampling cache during system operation, the local data batch and the global data batch are adaptively calculated; further, the local data batch and the global data batch are sampled from the local and global sampling caches respectively, and finally the two are combined to obtain the training data batch, which is used for the current iteration optimization step; (5) Batch parallel optimization of dynamic 3D Gaussian model: Based on the training data batch obtained in step (3), firstly, the human head dynamic 3D Gaussian model is driven in batches according to the human head motion parameters to obtain 3D Gaussian under different motion parameters; secondly, the 3D Gaussian is rendered in batches; then the rendering results are compared with the real images in the training data batch to calculate the loss; finally, the optimization parameters of the human head dynamic 3D Gaussian model are updated iteratively through image loss. (6) Human head animation generation: Based on the human head motion parameters and camera perspective given by the user, the modeled human head dynamic 3D model is driven and the head animation under new perspective and expression is generated in real time.

2. The real-time online construction method for a dynamic three-dimensional Gaussian model of the human head according to claim 1, characterized in that, In step (1), the training samples to be stored in the local sampling cache and the global sampling cache include human head motion parameters, real images, and human head foreground masks; and the capacity of the local sampling cache and the global sampling cache remains unchanged after creation and will not increase indefinitely as the input data increases. During system operation, the local sampling cache is maintained in the form of a first-in-first-out queue, and the global sampling cache is maintained in the form of a reservoir sampling method.

3. The real-time online construction method for a dynamic three-dimensional Gaussian model of the human head according to claim 1, characterized in that, Step (2) specifically includes the following sub-steps: (2.1) Maintain the global sampling buffer by sampling from a reservoir. When the local sampling buffer is full, the last data sample D in the local sampling buffer will be stored. j Take it out, with The probability of D j Store it in the global sampling buffer, where j is the sample number that was retrieved. The current capacity of the global sampling cache; when D j Stored At that time, randomly discard one The sample in the reservoir sampling ensures that each piece of data in the data stream has an equal probability of remaining in the cache; (2.2) Maintain the local sampling buffer using a first-in-first-out queue. When the local sampling buffer is full, the last sample is retrieved and stored in D. i Otherwise, store directly in D. i .

4. The real-time online construction method for a dynamic three-dimensional Gaussian model of the human head according to claim 1, characterized in that, In step (2), the convergence speed in the early stage of training is accelerated by using the batch parallel optimization method of the dynamic three-dimensional Gaussian model of the human head to accelerate convergence; specifically, the batch rendering of the three-dimensional Gaussian model is used to improve the training throughput, so that the model can achieve a real-time convergence rate and can be applied to online modeling.

5. The real-time online construction method for a dynamic three-dimensional Gaussian model of the human head according to claim 1, characterized in that, The specific steps (3) are as follows: First, the three-dimensional Gaussian is projected onto the two-dimensional plane through camera parameters. Then, the two-dimensional Gaussian kernel is convolved with the pixel color of the input image. Finally, the calculation result is assigned to the color attribute of the Gaussian, and its expression is as follows: Where W and H are the width and height of the image, respectively, and I... xy w represents the RGB color value of the (x,y) pixel in the image. xy The weights of Gaussian splatting are applied to the (x,y) pixels of the image.

6. The real-time online construction method for a dynamic three-dimensional Gaussian model of the human head according to claim 1, characterized in that, Step (4) includes the following sub-steps: (4.1) Based on the ratio of the sum of the sampling probability densities of the current local sampling cache and the global sampling cache, the training data batch is adaptively divided to ensure that the probability of sampling from the global sampling cache is less than or equal to the probability of sampling from the local cache. (4.2) Randomly sample local training data batches from the local sampling buffer; (4.3) Sample global data batches from the global sampling buffer with importance based on the length of time the samples have been added to the buffer. The mathematical formula for calculating the sampling density is as follows: W(k)=exp(w l ·k / |M l |); That is, when the input video frame is the i-th frame, Sample D of the j-th frame j The probability of being sampled, where For the present The number of samples in w l The sampling coefficients w are manually set. l The larger the value, the higher the sampling weight assigned to the new sample.

7. The real-time online construction method for a dynamic three-dimensional Gaussian model of the human head according to claim 1, characterized in that, In step (4), online data sample sampling is achieved by using a local-global combined data sample sampling algorithm suitable for online modeling. First, local sampling is used to quickly fit the new input training samples, and then global sampling is used to alleviate the model forgetting problem, thereby achieving a reconstruction quality similar to that of offline modeling.

8. The real-time online construction method for a dynamic three-dimensional Gaussian model of the human head according to claim 1, characterized in that, The batch rendering of the three-dimensional Gaussian in step (5) specifically includes the following sub-steps: (5.1) Batch preprocessing of 3D Gaussian: Based on the camera parameters, each 3D Gaussian is projected onto a 2D plane and the number of Gaussian instances is recorded. Each group of Gaussians in the 3D Gaussian batch is calculated on a different CUDA Stream. (5.2) GPU-CPU synchronization: After all CUDA Stream calculations in step (5.1) are completed, perform a GPU-CPU synchronization to transfer the number of 3D Gaussian instances to the CPU and create an array of the corresponding size on the GPU to store the screen space depth of the Gaussian instance, the corresponding pixel block number and the corresponding 3D Gaussian number. (5.3) 3D Gaussian Batch Rasterization: Based on the projected 3D Gaussian parameters obtained in step (5.1) and the 3D Gaussian instances obtained in step (5.2), the contribution value of each Gaussian instance to the color of the current pixel is calculated in the order of depth and the transparency is mixed to obtain the color of the current pixel; each group of Gaussians in the 3D Gaussian batch is calculated on a different CUDA Stream, and finally the batch image after rendering is obtained.