High-precision three-dimensional face real-time reconstruction method and system based on monocular video

CN122737331APending Publication Date: 2026-09-11HANGZHOU GONGSHU DISTRICT HOLOGRAPHIC INTELLIGENT TECHNOLOGY RESEARCH INSTITUTE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611201320.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-10
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0006]针对现有基于稠密先验的优化重建方法普遍依赖一阶优化器、每帧需数百次迭代才能收敛、难以部署为实时流程的问题,本发明提供一种基于稠密纹理坐标先验与高斯牛顿优化的基于单目视频的高精度三维人脸实时重建方法及系统

Benefits of technology

[0025] This invention transforms the dense texture coordinate prior into a vertex-by-vertex least squares objective and fits it with a Gauss-Newton solver that descends in block coordinates. The convergence speed of this invention is much faster than that of the traditional first-order optimizer, and the solution efficiency is improved by orders of magnitude.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122737331A_ABST
    Figure CN122737331A_ABST
Patent Text Reader

Abstract

This invention discloses a high-precision real-time 3D face reconstruction method and system based on monocular video. The method processes monocular face images or video streams, predicts correspondences, relative depths, and uncertainties in the texture coordinate domain of a parameterized face model using a prior prediction network. Vertex-by-vertex fitting targets are obtained by sampling at the texture coordinates of each vertex, constructing the 3D face parameter fitting as a confidence-weighted nonlinear least squares problem. Identity and dynamic parameters are optimized using a block coordinate descent Gauss-Newton solver with closed-form Jacobians. This invention also designs a keyframe buffer and identity incremental optimization strategy, progressively refining shared identity parameters while tracking frame-by-frame online. This invention offers high solution efficiency and good real-time performance, uniformly supporting three modes: single-image fitting, offline sequence reconstruction, and online tracking. It can be widely applied in scenarios such as face motion capture, 3D avatar driving, remote conferencing, and virtual reality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision and 3D reconstruction technology, and in particular to a high-precision real-time 3D face reconstruction method and system based on monocular video. Background Technology

[0002] High-fidelity monocular 3D face reconstruction is a core foundational technology for applications such as remote conferencing, facial animation, and 3D virtual avatars. Parametric face models (such as 3D deformation models and parametric face models) constrain the underdetermined image-to-geometric mapping problem to a low-dimensional space spanned by identity, expression, and pose parameters, providing strong prior constraints for monocular reconstruction. Existing monocular 3D face reconstruction methods are mainly divided into two categories: feedforward regression methods and optimization-based fitting methods.

[0003] The first category is feedforward regression methods, which directly regress the parameters of a parameterized face model from a single image, eliminating the need for test-time optimization and thus exhibiting lower inference latency. Among these, detailed representational face models are a widely used baseline for single-image reconstruction based on parameterized face models; emotion-driven face capture further enhances expression reconstruction through emotion-aware supervision; and facial component tokens introduce facial component tokens and combine 2D and 3D supervision to improve reconstruction accuracy. Recent methods improve feedforward regression through more expressive rendering modeling, such as introducing analysis-synthesis supervision based on neural renderers through neural synthesis analysis, binding 3D Gaussian to the predicted head mesh using RGB reconstruction loss for 2D Gaussian self-supervised head geometry prediction, and focusing on metrically accurate neutral identity reconstruction through metric-accurate methods. Although feedforward methods are highly efficient, their reconstruction accuracy is often still inferior to optimization-based fitting methods in challenging scenarios such as large head poses and strong facial expressions.

[0004] The second category is optimization-based fitting methods, which explicitly fit a parameterized face model to the input image during testing. These methods often offer more flexible adaptation capabilities and higher reconstruction accuracy. Early real-time face capture systems utilized depth-based red-green-blue (IRB) inputs to obtain robust geometric evidence for tracking. Subsequent work extended face tracking to monocular IRB inputs. Among these, displacement-based dynamic expression regression uses sparse keypoints for real-time expression tracking, face reconstruction relies on photometric consistency constraints, and dynamic rigid prior-stabilized tracking methods combine optical flow with learned dynamic rigid priors to improve tracking stability. More recent fitting frameworks, such as metric photometric trackers and adaptive appearance prior head alignment, leverage more expressive parameterized face models and appearance modeling to improve reconstruction quality, but all are designed as offline optimization processes. Furthermore, the sparse keypoints, photometric terms, or optical flow relied upon by these methods provide limited dense geometric evidence for high-fidelity reconstruction.

[0005] To overcome the aforementioned limitations, another type of work utilizes deep networks to predict intermediate dense priors for parametric model fitting, while retaining the flexibility of optimization-based reconstruction. Dense keypoint 3D face reconstruction methods and continuous keypoint detection methods predict dense keypoints for parametric face model fitting. More recent methods predict richer dense priors, such as texture coordinate optical flow face tracking predicting the correspondence between texture coordinates and images for 3D face tracking, screen space prior single-image reconstruction predicting screen space texture coordinates and normal priors for single-image fitting, and neural parametric head model regression exploring refined fitting when testing more expressive neural parametric head models. While these methods improve reconstruction quality, their fitting stage generally relies on first-order optimizers, such as adaptive moment estimation, requiring hundreds of iterations per frame to converge, primarily targeting offline reconstruction scenarios. Therefore, high-fidelity fitting based on dense priors is difficult to deploy as a lightweight real-time reconstruction process, which is precisely the core technical problem that this invention aims to solve. Summary of the Invention

[0006] To address the problems of existing dense prior-based optimization and reconstruction methods that generally rely on first-order optimizers, require hundreds of iterations per frame to converge, and are difficult to deploy as real-time processes, this invention provides a high-precision real-time 3D face reconstruction method and system based on monocular video, which is based on dense texture coordinate priors and Gauss-Newton optimization.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A high-precision real-time 3D face reconstruction method based on monocular video includes the following steps:

[0009] S1. Input a monocular face video stream, initialize the keyframe buffer and determine the camera intrinsic parameters; perform single-image fitting on the first frame to obtain the initial identity parameters and dynamic parameters of the parameterized face model; the identity parameters are shared by all frames, while the dynamic parameters, including expression, global rotation, local joints and global translation, are unique to each frame.

[0010] S2. For each input frame, the image is fed into the prediction network to predict the correspondence map, the relative depth map, and their respective log-variance maps in the texture coordinate domain of the parameterized face model. The correspondence map maps each texture coordinate position to image space coordinates, the relative depth map predicts the relative depth value at each position, and the log-variance map represents the uncertainty of the corresponding parameters.

[0011] S3. At the predefined texture coordinates of each parameterized face model vertex, sample each image obtained in step S2 to obtain the two-dimensional corresponding target, relative depth target, and vertex-by-vertex confidence weight determined by the logarithmic variance of the vertex, which are used as the fixed least squares target term for the frame.

[0012] S4. Divide the face parameters into identity parameter blocks and dynamic parameter blocks; drive the parameterized face model with the two types of parameters to obtain the face vertices with pose, calculate the image projection coordinates and relative depth of each vertex, and subtract them from the corresponding target and relative depth target in step S3 to obtain the vertex-by-vertex residual; weight the residuals with confidence weights and add a regularization term to form the least squares fitting energy.

[0013] S5. Fix one parameter block, linearize the residual of the other currently active parameter block, construct the Gauss-Newton normal equation with damping terms using the closed Jacobian and solve it, and update the active parameter block with the obtained increment.

[0014] S6. For each newly arrived frame, use the dynamic parameters of the previous frame as initial values ​​and fix the shared identity. Update the dynamic parameters according to step S5 to complete the tracking of the frame. At the same time, maintain the key frame buffer. When updating the buffer, combine all key frame data to optimize the shared identity. Only use past frames and the current frame throughout the process and output the 3D face model of each frame in real time.

[0015] Preferably, the relative depth mentioned in step S2 is defined relative to the neck joint of the parameterized face model. The relative depth of any vertex is the camera space depth of that vertex minus the camera space depth of the neck joint. The relative depth is used to eliminate the inconsistency in absolute scale caused by random cropping and scaling during training, remove global translation components and focus on the local three-dimensional geometry of the face. The global translation is left to be recovered by the optimizer during the fitting stage.

[0016] Preferably, the prediction network in step S2 consists of an image encoder and two texture coordinate domain decoding branches; the image encoder extracts multi-scale features from the input image; a set of learnable texture coordinate feature maps bound to the texture coordinate layout of the parameterized face model is used as a query, and each decoding branch fuses the query with the image features to generate a texture coordinate domain feature map; the correspondence and relative depth use their own independent decoding branches, and each branch is followed by a dense prediction head to predict the target map and its log-variance map respectively.

[0017] Preferably, the vertex-by-vertex confidence weight in step S3 is obtained by exponential transformation of the logarithmic variance obtained by sampling the vertex. The higher the prediction uncertainty, i.e. the larger the logarithmic variance, the lower the weight of the vertex in the fitting energy.

[0018] Preferably, in step S4, the dynamic parameter block energy is composed of the weighted corresponding residual and the weighted relative depth residual per vertex, plus the expression and pose regularization terms; the identity parameter block energy is composed of the weighted residual accumulated from multiple keyframes per vertex, plus the identity regularization term; the identity regularization term is the L2 norm regularization of the identity parameters with respect to the zero vector.

[0019] Preferably, in step S5, the residual vector of the currently active parameter block is linearized, a Gauss-Newton normal equation with damping terms is constructed, the parameter increment is obtained by Choleski decomposition, and the active parameter block is updated accordingly; when updating the dynamic parameter block and the identity parameter block, the closed Jacobian of the residual with respect to the dynamic parameter and the identity parameter is used respectively, and the closed Jacobian is calculated by a custom graphics processor operator.

[0020] A high-precision 3D face real-time reconstruction system based on monocular video includes an initialization module, a dense prior prediction module, a vertex-by-vertex sampling module, a least-squares modeling module, a Gauss-Newton solving module, and an online tracking module. The initialization module receives the monocular face video stream and outputs the initial identity parameters, initial dynamic parameters, and camera intrinsic parameters of the parameterized face model, completing the keyframe buffer initialization. The dense prior prediction module receives the input frame image and outputs a correspondence map, a relative depth map, and their log-variance map in the texture coordinate domain of the parameterized face model. The vertex-by-vertex sampling module samples each map at the texture coordinates of each vertex of the parameterized face model. The system outputs vertex-by-vertex 2D corresponding targets, relative depth targets, and confidence weights. The least squares modeling module divides face parameters into identity parameter blocks and dynamic parameter blocks, calculates vertex-by-vertex residuals and weights them, and outputs the least squares fitting energy in combination with regularization terms. The Gauss-Newton solution module uses a block coordinate descent method, fixes one parameter block and linearizes the residuals of another activated parameter block, constructs a damped Gauss-Newton normal equation using closed Jacobian, and outputs parameter increments. The online tracking module updates the current frame's dynamic parameters using the previous frame's dynamic parameters as initial values, maintains a keyframe buffer, and jointly optimizes shared identities during buffer updates, outputting a 3D face model for each frame.

[0021] Preferably, the dense prior prediction module includes an image encoder and two texture coordinate domain decoding branches; the image encoder extracts multi-scale features of the input image; a learnable texture coordinate feature map bound to the texture coordinate layout of the parameterized face model is used as a query, and each decoding branch fuses the query and image features to generate a texture coordinate domain feature map; the correspondence and relative depth are output as target map and log-variance map through independent decoding branches and dense prediction head, respectively.

[0022] Preferably, the Gauss-Newton solver module linearizes the residual vector of the current active parameter block, constructs the normal Gauss-Newton equation with damping terms, and uses Choleski decomposition to obtain the parameter increment and update the active parameter block; when updating the dynamic parameter block and the identity parameter block, the closed Jacobian of the residual with respect to the dynamic parameter and the identity parameter is used respectively, and the closed Jacobian is calculated in parallel by the graphics processor.

[0023] Preferably, the online tracking module maintains a keyframe buffer of fixed capacity. Every few frames, it calculates the head rotation of the current frame and calculates the minimum geodesic distance between the current frame and the head rotation of existing keyframes in the buffer on the three-dimensional rotation group as the novelty. When the buffer is not full and the novelty exceeds the threshold, the current frame is inserted. When the buffer is full, it is inserted only if replacing a keyframe can improve the pose coverage. Each buffer update adds an identity optimization budget, which is distributed to several subsequent frames. Each frame performs at most one Gauss-Newton iteration to incrementally improve the shared identity parameters.

[0024] Compared with the prior art, the beneficial effects of the present invention are:

[0025] This invention transforms the dense texture coordinate prior into a vertex-by-vertex least squares objective and fits it with a Gauss-Newton solver that descends in block coordinates. The convergence speed of this invention is much faster than that of the traditional first-order optimizer, and the solution efficiency is improved by orders of magnitude.

[0026] This invention can process video streams in real time on a single consumer-grade graphics processor while maintaining near-offline reconstruction accuracy; the offline sequence reconstruction speed is significantly improved compared to optimized baseline methods.

[0027] Based on the publicly available single-view face reconstruction benchmark, this invention outperforms existing best methods in both neutral and pose-inclusive reconstruction, reaching the current advanced level in the field.

[0028] This invention provides a single solver that uniformly supports three reconstruction modes: single-graph fitting, offline sequence reconstruction, and online tracking. The framework is simple, easy to deploy, and can be flexibly adapted to the needs of different application scenarios.

[0029] This invention utilizes a keyframe buffering and identity optimization budget sharing mechanism to progressively refine shared identity parameters as video input progresses during online tracking, achieving incremental accuracy improvements without accessing future frames and meeting strict real-time causality constraints. Attached Figure Description

[0030] Figure 1 This is an overall flowchart of the method of the present invention;

[0031] Figure 2 This is a diagram showing the overall architecture of the system of the present invention;

[0032] Figure 3 Intermediate and final result images for reconstructing the first user's 3D face model using the method of this invention;

[0033] Figure 4 Intermediate and final result images for reconstructing a second user's 3D face model using the method of this invention;

[0034] Figure 5Intermediate and final result images for reconstructing a third user's 3D face model using the method of this invention;

[0035] Figure 6 Intermediate and final results of reconstructing the 3D face model of the fourth user using the method of this invention. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0037] Example 1

[0038] like Figure 1 As shown, this embodiment provides a high-precision real-time 3D face reconstruction method based on monocular video, using dense texture coordinate prior and Gauss-Newton optimization. The system input is a single monocular face image or a monocular face video stream, preferably a face video stream acquired in real time by a monocular camera. The system output is a 3D face model corresponding to each frame, including parameterized face model identity parameters, expression parameters, head pose parameters, local joint parameters, global translation parameters, and a 3D face mesh generated from these parameters.

[0039] The identity parameters of the parameterized face model are denoted as: Used to control the shape of a neutral human face; the dynamic parameters of the current frame are denoted as... ,in For facial expression parameters, For global rotation, For local joint parameters, This is a global translation. Given identity parameters. With dynamic parameters The parametric face model outputs pose-indicated face vertices and joints through linear blending shape and linear blending skin, denoted as:

[0040]

[0041] In the formula, For the first The face vertex coordinate matrix of the frame; For the first The joint coordinate matrix of the frame; This is the forward computation function for the parameterized face model.

[0042] The runtime status variables in this embodiment include: current shared identity parameters. Previous frame dynamic parameters Current frame dynamic parameters Camera intrinsic parameter matrix Fixed-capacity keyframe buffer And the optimization budget used for distributing identity optimization. The keyframe buffer stores the vertex-by-vertex fitted target, dynamic parameters, and head rotation of several historical frames, which are used to gradually refine the shared identity parameters during online processing.

[0043] Step S1, Initialization:

[0044] Input a monocular face video stream, initialize the keyframe buffer, and determine the camera intrinsics. Perform single-image fitting on the first frame to obtain the initial identity parameters and dynamic parameters of the parameterized face model. The identity parameters are shared by all frames, while the dynamic parameters, including expression, global rotation, local joints, and global translation, are unique to each frame.

[0045] Video stream access and face detection:

[0046] After system startup, it receives a monocular face video stream and performs face detection and alignment on the first frame. The face detector outputs a face bounding box and five facial key points, which the system uses to crop and scale the face region to the network input resolution. The cropping transformation matrix is ​​recorded for subsequent coordinate transformation between the original image space and the cropped image space.

[0047] Camera internal parameters confirmed:

[0048] When the camera intrinsic parameters are known, the system directly uses the given camera intrinsic parameters for subsequent projection and fitting. When the camera intrinsic parameters are unknown, the focal length is estimated by performing a one-dimensional golden section search on the camera's vertical field of view.

[0049] First, a search interval for the field of view (LAD) is defined, and two LADs are selected within this interval using the golden ratio. For each LAD, the system constructs a candidate camera intrinsic parameter matrix and performs attitude initialization and single-image fitting sequentially on the current frame under that candidate camera. The sum of squared residuals of the per-vertex correspondence after fitting is used as the energy value of that candidate LAD.

[0050] After comparing the energy of the two detection field-of-view angles, the sub-interval with the smaller energy is retained, and new detection field-of-view angles are added according to the golden ratio. This shrinkage process is repeated until the search interval width is less than a set threshold or the maximum number of iterations is reached. The camera focal length is calculated by converting the field-of-view angle corresponding to the midpoint of the final interval. The conversion formula is:

[0051]

[0052] In the formula, This is the camera focal length, expressed in pixels. The pixel resolution in the vertical direction of the image; Define the camera's vertical field of view.

[0053] The estimated camera intrinsic parameters were then used for fitting the first frame and for tracking all subsequent online video stream frames, remaining constant throughout the tracking process.

[0054] First frame single image fitting:

[0055] A single-graph fitting is performed on the first frame of the video stream to initialize shared identity parameters and first-frame dynamic parameters. The single-graph fitting employs a two-stage strategy:

[0056] The first stage is pose initialization: fix the expression parameters and identity parameters as zero vectors, optimize only global rotation and global translation, and use a large damping coefficient to ensure numerical stability, so as to obtain a stable initial value of head pose.

[0057] The second stage is alternating optimization: starting with the initial pose value, Gaussian-Newton updates are performed alternately between the dynamic parameter block and the identity parameter block. In each iteration, the identity parameters are first fixed, and the energy of the dynamic parameter block is minimized to update the dynamic parameters. Then, the dynamic parameters are fixed, and the energy of the identity parameter block is minimized to update the identity parameters. After several rounds of alternating execution, the identity parameter estimates and dynamic parameter estimates for the first frame are obtained.

[0058] Keyframe buffer initialization:

[0059] After the first frame is fitted, the vertex-by-vertex fitted target, dynamic parameters, and head rotation of the first frame are stored in the keyframe buffer as the first keyframe. The keyframe buffer adopts a fixed capacity design, and the capacity remains unchanged during the operation. Subsequent keyframes are gradually added or replaced according to the pose novelty rule.

[0060] Step S2, Dense Texture Coordinate Prior Prediction:

[0061] For each input frame, the image is fed into the prediction network, which predicts a correspondence map, a relative depth map, and their respective log-variance maps in the texture coordinate domain of the parameterized face model. The correspondence map maps each texture coordinate position to image space coordinates, the relative depth map predicts the relative depth value at each position, and the log-variance map characterizes the uncertainty of the corresponding parameters.

[0062] Predict the overall network structure:

[0063] The dense texture coordinate prior prediction network consists of an image encoder and two texture coordinate domain decoding branches. The image encoder extracts multi-scale features from the input image; a set of learnable texture coordinate feature maps bound to the texture coordinate layout of a parameterized face model is used as a query, and each decoding branch fuses this query with the image features to generate a texture coordinate domain feature map. Correspondence and relative depth use separate decoding branches, each followed by a dense prediction head to predict the target map and its log-variance map, respectively.

[0064] Image encoder:

[0065] The image encoder employs a visual transformer architecture, dividing the input image into a sequence of patches and extracting multi-scale image features through multi-layer self-attention transformation. The encoder outputs four sets of feature maps at different resolutions, each corresponding to a different downsampling rate, for subsequent decoding branches to perform multi-scale feature fusion.

[0066] Cross-attention fusion:

[0067] Each decoding branch uses a learnable texture coordinate query feature map as the query vector and multi-scale image features output by the encoder as the key vector. Through cross-attention operations, the image features are converged onto the texture coordinate grid to generate a texture coordinate aligned feature map. The spatial layout of the texture coordinate query feature map is strictly consistent with the standard texture coordinate unfolding method of the parametric face model, with each spatial location corresponding to a sampling point in the texture coordinate domain.

[0068] Dense prediction head and output:

[0069] The dense prediction head employs multi-level convolution and upsampling structures to progressively refine the fused texture coordinate feature map, ultimately outputting target prediction and log-variance prediction. The correspondence branch outputs the correspondence map and the log-variance map of the two channels; the relative depth branch outputs the relative depth map and the log-variance map of a single channel.

[0070] Relative depth definition:

[0071] The relative depth is defined relative to the neck joint of the parametric face model, using the neck joint as a stable anchor point. For any vertex... Its relative depth is the camera space depth at the vertex minus the camera space depth at the neck joint, that is:

[0072]

[0073] In the formula, For the first Frame vertex The relative depth; For the first Frame vertex Depth value in camera coordinate system; For the first The depth value of the frame neck joint in the camera coordinate system.

[0074] Using relative depth instead of absolute depth can eliminate the inconsistency in absolute scale caused by random cropping and scaling during training, remove global translation components and focus on the local 3D geometry of the face, and leave the global translation components to be recovered by the optimizer during the fitting stage.

[0075] The physical meaning of logarithmic variance:

[0076] The log-variance plot represents the uncertainty of the network's prediction for each location. A larger log-variance indicates lower confidence in the network's prediction for that location; a smaller log-variance indicates a more reliable prediction. This uncertainty is automatically learned by the network during training and is subsequently used to construct the confidence weights for each vertex.

[0077] Step S3: Vertex-by-vertex fitting of target sampling:

[0078] At the predefined texture coordinates of each parameterized face model vertex, the images obtained in step S2 are sampled to obtain the two-dimensional corresponding target, the relative depth target, and the vertex-by-vertex confidence weight determined by the logarithmic variance, which are used as the fixed least squares target term for that frame.

[0079] Vertex texture coordinate lookup table:

[0080] Each vertex of the parametric face model has predefined texture coordinates, which are fixed during model definition and do not change with identity or expression parameters. The system pre-stores the texture coordinates of all vertices as a lookup table, which is directly read at runtime.

[0081] Bilinear sampling:

[0082] For the current frame, after obtaining four texture coordinate domain maps, the system iterates through all the vertices of the parameterized face model and performs bilinear sampling on the four maps at the texture coordinates of each vertex. Bilinear sampling reads the values ​​of the four pixels surrounding the floating-point position of the texture coordinates and obtains the precise sampled value at that vertex through weighted interpolation.

[0083] After sampling, the two-dimensional corresponding target, relative depth target, logarithmic variance of the correspondence, and logarithmic variance of the relative depth are obtained for each vertex. The above sampled values ​​remain fixed during the Gaussian-Newton optimization process of the current frame and are used as the vertex-by-vertex least squares fitting target for that frame.

[0084] Confidence weight calculation:

[0085] The confidence weights for each vertex are obtained by exponentially transforming the logarithmic variance. The correspondence residual weights and relative depth residual weights each incorporate an exponential transformation factor, ensuring that vertices with higher prediction uncertainty (i.e., larger logarithmic variance) have lower weights in the fitted energy. The formula for calculating the confidence weights is:

[0086]

[0087] In the formula, For the first Frame vertex Confidence weights; For the first Frame vertex The logarithmic variance at point ; It is a natural exponential function.

[0088] Through this adaptive weighting mechanism based on prediction confidence, the weights of vertices that are difficult to predict accurately, such as self-occluded regions and large-pose side face regions, are automatically reduced, thereby improving the overall robustness of the fitting.

[0089] Step S4, Nonlinear Least Squares Modeling:

[0090] The face parameters are divided into identity parameter blocks and dynamic parameter blocks. The parameterized face model is driven by the two types of parameters to obtain the face vertices with pose. The image projection coordinates and relative depth of each vertex are calculated and subtracted from the corresponding target and relative depth target in step S3 to obtain the vertex-by-vertex residual. The residuals are weighted with confidence weights and regularization terms are added to form the least squares fitting energy.

[0091] Parameter block partitioning:

[0092] All face parameters to be optimized are divided into two independent parameter blocks: an identity parameter block and a dynamic parameter block. The identity parameter block contains only identity parameters. Shared across all frames; the dynamic parameter block contains facial expression parameters. Global rotation Local joint parameters With global translation Each frame is unique.

[0093] Vertex projection and residual calculation:

[0094] Given the current identity parameters Current frame dynamic parameters and camera intrinsic parameter matrix First, the pose-integrated face vertices and joints are obtained through forward computation using a parameterized face model. Then, perspective camera projection is performed on each vertex to obtain the current 2D projected coordinates.

[0095]

[0096] In the formula, For the first Frame vertex Two-dimensional coordinates projected onto the image plane; For perspective projection functions; For the first Frame vertex 3D coordinates; This is the camera intrinsic parameter matrix.

[0097] Simultaneously, the current relative depth is calculated based on the vertex depth and neck joint depth. The current projected coordinates and the current relative depth are then subtracted from the sampled target values ​​to obtain the corresponding residual and the relative depth residual, respectively.

[0098] Weighted residuals:

[0099] The confidence-weighted correspondence residuals and relative depth residuals are written as follows:

[0100]

[0101]

[0102] In the formula, For the first Frame vertex The weighted correspondence residual vector; For the first Frame vertex Weighted relative depth residual scalar; These are the global weighting coefficients for the corresponding residuals; The global weighting coefficients for the relative depth residuals; For the first Frame vertex The logarithmic variance of the corresponding relationships; For the first Frame vertex The logarithmic variance at relative depth; This represents element-wise multiplication; For the first Frame vertex The corresponding target coordinates; For the first Frame vertex The relative depth target value.

[0103] Dynamic parameter block energy:

[0104] The dynamic parameter block energy consists of the correspondence residuals of all vertices in the current frame, the relative depth residuals, and the expression and pose regularization terms:

[0105]

[0106] In the formula, The total energy of the dynamic parameter block; Let L be the L2 norm of the vector; The weighting coefficients for the facial expression regularization term; These are the weighting coefficients for the local joint regularization term; For the first The expression parameter vector of the frame; For the first The local joint parameter vector of the frame.

[0107] Both regularization terms are L2 squared regularizations, used to prevent facial or pose parameters from deviating from a reasonable range and improve the stability of the fit.

[0108] Identity parameter block energy:

[0109] The identity parameter block energy is obtained by summing the vertex-by-vertex residuals over multiple keyframes, with the identity parameter regularization term added:

[0110]

[0111] In the formula, The total energy of the identity parameter block; This is the set of indices for keyframes; The weight coefficients for the identity regularization term; This is a vector of identity parameters.

[0112] The identity regularization term is a L2 norm squared regularization of the identity parameters with respect to the zero vector, which constrains the identity parameters to shrink towards the zero vector to avoid overfitting and ensures that the reconstruction result conforms to the statistical average face shape.

[0113] Step S5: Gauss-Newton solution for block coordinate descent:

[0114] Fix one parameter block, linearize the residual of the other currently active parameter block, construct the Gauss-Newton normal equation with damping terms using the closed Jacobian and solve it, and update the active parameter block with the resulting increment.

[0115] Block coordinate descent strategy:

[0116] A block coordinate descent strategy is used to alternately optimize two parameter blocks: for dynamic parameter updates, the shared identity parameter is fixed. Based on the current frame dynamic parameters To activate variables, minimize the energy of the dynamic parameter block; for identity parameter updates, fix the dynamic parameters of the relevant frames, and use the identity parameters... To activate variables, minimize the energy of the identity parameter block. Optimize by activating only one parameter block at a time, while keeping the other parameter block fixed.

[0117] Gauss-Newton linearization:

[0118] For any activation parameter block, let the current parameter be... The corresponding residual vector is Linearizing the residual vector with respect to the current parameters yields a first-order approximation:

[0119]

[0120] In the formula, For parameter increments; Let be the Jacobian matrix of the residual with respect to the parameters.

[0121] By minimizing the sum of squared residuals after linearization, the Gauss-Newton normal equations can be derived. After adding a Levenberg-Marquardt style damping term, the damped Gauss-Newton normal equations are:

[0122]

[0123] In the formula, This is the transpose of the Jacobian matrix; The damping coefficient; It is the identity matrix; This is the current residual vector.

[0124] The damping term limits the step size and improves iteration stability when the current parameter is far from the optimum; when it approaches the optimum, the damping effect weakens, restoring the fast convergence of the standard Gauss-Newton method.

[0125] Choleski Decomposition Solution

[0126] Because the variables in the dynamic parameter block and identity parameter block have low dimensionality, and the residuals come from a large number of vertices, the coefficient matrix of the normal equation is a small-scale symmetric positive definite matrix. Choleski decomposition is used for efficient solution.

[0127] The first step is to perform Cholliski decomposition on the coefficient matrix to obtain the lower triangular matrix. ,satisfy ;

[0128] The second step is to perform a forward-backward substitution solution. To obtain the intermediate vector ;

[0129] The third step is to perform a backward substitution solution. , obtain parameter increment .

[0130] The computational complexity of the Choreski decomposition is O(n). ,in To activate the dimension of the parameter block. Due to the low parameter dimension, the decomposition operation is very fast.

[0131] Parameter update:

[0132] After obtaining the parameter increment, add the increment to the currently active parameter block to complete one Gauss-Newton iteration update:

[0133]

[0134] In the formula, The parameter vector for activating the parameter block; This is the parameter increment obtained from the solution.

[0135] Closed Jacobi:

[0136] The Jacobian matrix uses an analytically derived closed-form Jacobian, rather than being calculated numerically through automatic differentiation. When dynamic parameters are updated, the Jacobian includes the derivatives of vertex projected coordinates and relative depth with respect to expression parameters, global rotation, local joints, and global translation; when identity parameters are updated, the Jacobian includes the derivatives of vertex projected coordinates and relative depth with respect to identity parameters.

[0137] Specifically, the derivative of vertex position with respect to the identity parameter can be directly obtained from the identity blending shape basis matrix of the parameterized face model; the derivative of vertex position with respect to the expression parameter is obtained from the expression blending shape basis matrix; the derivative of vertex position with respect to the rotation parameter is derived from the axial-angular derivative of the rotation matrix; and the derivative of vertex position with respect to the translation parameter is the identity matrix. Further, by using the chain rule and combining it with the Jacobian of perspective projection, the analytical derivatives of the projection coordinates with respect to each parameter can be obtained. The Jacobian of relative depth is obtained by subtracting the Jacobians of the depth components.

[0138] By using the analytical Jacobian, this invention avoids the problem of requiring numerous iterations from general automatic differentiation or first-order optimizers, thus enabling real-time optimization-based 3D face reconstruction based on dense priors. The closed-form Jacobian is implemented in parallel by a custom graphics processor operator.

[0139] Step S6: Frame-by-frame online tracking and incremental identity optimization:

[0140] For each newly arrived frame, the dynamic parameters of the previous frame are used as initial values ​​and the shared identity is fixed. The dynamic parameters are updated according to step S5 to complete the tracking of the frame. At the same time, the key frame buffer is maintained. When the buffer is updated, the shared identity is optimized by combining all key frame data. Only past frames and the current frame are used throughout the process, and the 3D face model of each frame is output in real time.

[0141] Frame-by-frame dynamic tracking:

[0142] For each newly arrived input frame, first perform dense prior prediction in step S2 and vertex-by-vertex sampling in step S3 to obtain the vertex-by-vertex fitting target for the current frame. Then, use the dynamic parameters of the previous frame as the initial values ​​of the dynamic parameters of the current frame, keep the current shared identity parameters unchanged, and perform several steps of Gauss-Newton iteration updates only on the dynamic parameter block to obtain the dynamic parameters of the current frame.

[0143] Since the current frame uses the result of the previous frame as its initial value, and the motion continuity between video frames is good, convergence can be achieved with only a small number of iterations per frame. This step is executed stably every frame and is the main path for online tracking.

[0144] 3D face output:

[0145] After obtaining the identity and dynamic parameters of the current frame, a 3D face mesh with pose is generated through forward computation using a parametric face model and output to downstream applications. The output includes vertex coordinates, triangle topology, identity parameters, expression parameters, and pose parameters, which can be directly used for 3D rendering, animation-driven processing, or data storage.

[0146] Keyframe buffer maintenance:

[0147] The system maintains a fixed-capacity keyframe buffer to collect diverse head pose samples to support incremental optimization of identity parameters. Every few frames, the system evaluates whether the current frame should be added to the keyframe buffer.

[0148] First, the head rotation of the current frame is calculated, which is obtained by combining the global rotation and the neck joint rotation. Then, the minimum geodesic distance between the head rotation of the current frame and the head rotations of all existing keyframes in the buffer on the 3D rotation group is calculated, and the minimum value is taken as the pose novelty of the current frame.

[0149] The formula for calculating the geodesic distance between two rotations on a three-dimensional rotation group is:

[0150]

[0151] In the formula, The geodesic distance between the two rotations; , These are the rotation matrices corresponding to the two rotations; It is the matrix logarithm function; Let be the Frobenius norm of the matrix.

[0152] When the keyframe buffer is not full and the pose novelty is greater than the set threshold, the current frame is inserted into the keyframe buffer; when the keyframe buffer is full, the change in buffer coverage after replacing each existing keyframe is evaluated, and replacement is only performed if replacing a keyframe can significantly improve the overall head pose coverage of the buffer.

[0153] Identity optimization budget allocation:

[0154] When the keyframe buffer is updated, a full identity optimization is not performed immediately; instead, an identity optimization budget is added. The budget is an integer number of steps, representing the total number of subsequent identity optimization iterations that can be performed.

[0155] When processing ordinary frames, if the identity optimization budget is greater than zero, the system loads all key frame data from the key frame buffer after completing the dynamic parameter tracking of the current frame, performs at most one Gauss-Newton update on the shared identity parameters, and decrements the budget by one.

[0156] Through this budget-sharing mechanism, identity refinement is distributed across multiple subsequent frames, avoiding significant computational spikes during keyframe insertion and ensuring stable frame rates. Simultaneously, as the head poses and expressions covered by the keyframes become increasingly rich, the shared identity parameters are gradually improved, achieving incremental enhancements in identity accuracy.

[0157] Causality guarantee:

[0158] The entire online tracking process strictly adheres to causal constraints. All calculations use only data from the current and historical frames, without accessing any future frame information. The keyframe buffer only stores processed historical frames, and identity optimization is based solely on existing keyframes, ensuring the system can be deployed as a strictly real-time online system, meeting the requirements of low-latency applications such as live streaming and video conferencing.

[0159] Compatibility mode extension:

[0160] The same solver is also compatible with single-image fitting and offline sequence reconstruction modes. Single-image fitting mode processes only a single input image, alternately performing Gaussian-Newton updates between dynamic parameter blocks and identity parameter blocks, outputting the corresponding 3D face parameters and mesh. Offline sequence reconstruction mode, after obtaining the complete video sequence, first tracks frame by frame and then refines the identity by selecting keyframes through farthest-point sampling. This can be repeated multiple times to achieve higher accuracy, making it suitable for offline applications where high accuracy is required but real-time performance is not critical.

[0161] Example 2

[0162] like Figure 2 As shown, this embodiment provides a high-precision real-time 3D face reconstruction system based on monocular video, using dense texture coordinate priors and Gauss-Newton optimization. The system employs a modular design, with data transferred between modules via pre-allocated video memory caches to minimize data transfer overhead between the central processing unit and the graphics processing unit. The system comprises six core modules: an initialization module, a dense prior prediction module, a vertex-by-vertex sampling module, a least-squares modeling module, a Gauss-Newton solution module, and an online tracking module. The specific structure and implementation details of each module are described below.

[0163] Initialize module:

[0164] The initialization module is responsible for various preparatory tasks during system startup, including video stream access, camera intrinsic parameter estimation, first frame single-image fitting, and keyframe buffer initialization. The module's input is the first frame image of the monocular face video stream, and its output includes initial shared identity parameters, first frame dynamic parameters, the camera intrinsic parameter matrix, and the initialized keyframe buffer.

[0165] The initialization module integrates a face detector, a cropping and alignment unit, a camera search unit, and a first-frame fitting unit. The face detector is implemented using a lightweight convolutional neural network, performing face detection on the input image and returning the face bounding box and five key points. The cropping and alignment unit crops the face region according to the detection results and scales it to the network input resolution, while recording the cropping transformation matrix for subsequent coordinate space transformation.

[0166] The camera search unit performs a golden ratio search for the field of view (LAV) angle when the camera intrinsic parameters are unknown. This unit maintains the upper and lower bounds of the search interval and the positions of the two detection points, iteratively calls the first frame fitting unit to evaluate the fitting energy of each candidate LAV angle, and narrows the search interval based on the energy comparison results. The number of search iterations is configurable; the default settings are sufficient to keep the LAV angle estimation accuracy within a small range.

[0167] The first frame fitting unit performs a two-stage single-image fitting: the first stage is pose initialization, fixing expression and identity parameters, optimizing only global rotation and global translation, and setting the damping coefficient to a large value to ensure stability; the second stage is alternating optimization, alternately performing Gaussian-Newton updates between the dynamic parameter block and the identity parameter block, with each update performing one iteration per block, and the number of alternation rounds is configurable. After the first frame fitting is completed, the first frame data is written to the keyframe buffer as the first keyframe.

[0168] The initialization module is also responsible for the pre-allocation of video memory resources. At system startup, the initialization module allocates all necessary caches in the graphics processor's video memory at once, based on the number of vertices, the keyframe buffer capacity, and the dimensions of each intermediate tensor. These caches include vertex caches, joint caches, residual caches, Jacobian caches, and normal equation matrix caches. This pre-allocation strategy avoids the overhead and fragmentation caused by dynamic memory allocation and release at runtime, ensuring frame rate stability.

[0169] Dense Prior Prediction Module:

[0170] The dense prior prediction module is responsible for performing forward inference of dense texture coordinates prior for each input frame, and outputs a correspondence map, a relative depth map, and its log-variance map. The input of this module is a cropped and aligned face image, and the output is four texture coordinate domain feature maps.

[0171] The dense prior prediction module consists of an image encoder submodule and two parallel decoding branch submodules. The image encoder submodule adopts a visual transformer architecture, dividing the input image into a sequence of patches and extracting multi-scale features through multi-layer self-attention transformation. The encoder outputs multiple sets of feature maps at different resolutions for subsequent decoding branches to perform multi-scale fusion.

[0172] Each decoding branch submodule consists of a cross-attention fusion layer and a dense prediction head. The cross-attention fusion layer uses a learnable texture coordinate query feature map as the query vector and multi-scale image features output by the encoder as the key vector. Through cross-attention operations, it converges the image features onto a texture coordinate grid to generate a texture coordinate aligned feature map. Each position in the texture coordinate query feature map corresponds to a sampling point in the texture coordinate space of the parameterized face model, and its layout is strictly consistent with the standard texture coordinate unfolding method of the parameterized face model.

[0173] The dense prediction head employs a dense prediction transformer structure, incorporating multiple convolutional layers and upsampling operations. It progressively refines the fused texture coordinate feature map and ultimately outputs the target prediction and log-variance prediction. The correspondence branch outputs the correspondence map and the log-variance map of the two channels; the relative depth branch outputs the relative depth map and the log-variance map of a single channel.

[0174] All computations in the dense prior prediction module are performed on the GPU, and network weights are stored as half-precision floating-point numbers to save GPU memory and accelerate inference. Tensor cores are enabled during inference to further improve throughput.

[0175] Vertex-by-vertex sampling module:

[0176] The vertex-by-vertex sampling module is responsible for converting the texture coordinate domain map output by the dense prior prediction module into a fitting target and confidence weight for each vertex of the parametric face model through bilinear sampling. The input of this module is four texture coordinate domain maps and a vertex texture coordinate lookup table, and the output is the vertex-by-vertex two-dimensional corresponding target, relative depth target, correspondence confidence weight, and relative depth confidence weight.

[0177] The vertex texture coordinate lookup table is a pre-computed and stored constant array containing the texture coordinate values ​​corresponding to each vertex of the parameterized face model. This lookup table is loaded into video memory during system initialization and remains unchanged during runtime.

[0178] The sampling operation is implemented in parallel using the graphics processor, with each thread responsible for the sampling calculation of one vertex. For each vertex, the thread locates its floating-point position in the texture coordinate map based on its texture coordinates, reads the values ​​of the four pixels surrounding that position, and performs bilinear interpolation to obtain the corresponding target, relative depth target, and corresponding log-variance at that vertex. Subsequently, the thread calculates the confidence weight based on the log-variance, i.e., performs an exponential operation and multiplies it by the corresponding global weight coefficient.

[0179] The output of the sampling module is organized into four vertex-level arrays, which are stored in a pre-allocated video memory cache and can be directly read by the subsequent least squares modeling module without the need for transmission between the central processing unit and the graphics processing unit.

[0180] Least squares modeling module:

[0181] The least squares modeling module is responsible for calculating the vertex-by-vertex residuals and the Jacobian matrix of the residuals with respect to the activation parameter blocks based on the current parameter state, providing input for the Gauss-Newton solver module to construct the normal equations. The inputs to this module are the current identity parameters, current dynamic parameters, camera intrinsic parameters, and the vertex-by-vertex fitting target and weights; the outputs are the weighted residual vector and the weighted Jacobian matrix.

[0182] The least squares modeling module comprises three parallel computation stages: the parametric face model forward pass stage, the projection and residual calculation stage, and the Jacobian calculation stage. These three stages are executed sequentially, with each stage exhibiting vertex-level parallelism.

[0183] In the forward pass phase of the parametric face model, the neutral space vertex positions are calculated based on the current identity and expression parameters. Then, linear blending skinning is performed based on the pose parameters to obtain the vertex coordinates and joint positions with pose parameters. Simultaneously, this phase calculates the partial derivative matrix of the vertex positions with respect to the activation parameters in parallel, depending on the type of activation parameter block.

[0184] In the projection and residual calculation stage, perspective projection is performed on each vertex, projecting the 3D vertex coordinates onto the 2D image plane to obtain the current projected coordinates. Simultaneously, the current relative depth of each vertex is calculated. Then, the current projected coordinates and relative depth are subtracted from the target value to obtain the original residual, which is then multiplied by a pre-calculated confidence weight to obtain the weighted residual. The regularization term and its weight are appended to the end of the residual vector in this stage.

[0185] In the Jacobian calculation stage, the Jacobian of the residuals with respect to the activation parameters is derived using the chain rule. The Jacobian of the projected coordinates is obtained by left-multiplying the vertex position Jacobian by the perspective projection Jacobian; the Jacobian of the relative depth is obtained by subtracting the depth component Jacobians. The Jacobian of the weighted residuals is further multiplied by the confidence weight coefficient. The Jacobian corresponding to the regularization term is the identity matrix multiplied by the square root of the regularization weight, and directly appended to the bottom of the Jacobian matrix.

[0186] The least squares modeling module is the core computational part of the Gauss-Newton iteration, and it needs to be executed again in each iteration. The computational cost is relatively small when updating dynamic parameters; however, when updating identity parameters, the computational cost is proportional to the number of keyframes due to the accumulation of multiple keyframes.

[0187] Gauss-Newton solver module:

[0188] The Gauss-Newton solver module is responsible for constructing and solving the normal Gauss-Newton equations with damping terms, obtaining parameter increments, and updating the active parameter block. This module is the core of the numerical solution for the entire system; its inputs are the weighted residual vector and the weighted Jacobian matrix, and its output is the updated parameter vector.

[0189] The Gauss-Newton solver module internally includes a normal equation construction unit, a Cholliski decomposition unit, and a parameter update unit. The normal equation construction unit performs matrix multiplication and matrix-vector multiplication to obtain the left-hand matrix and right-hand vector of the normal equation, and then adds a damping term on the diagonal. Since the number of rows in the Jacobian matrix is ​​much greater than the number of columns, constructing small-scale normal equations using right-hand multiplication is efficient.

[0190] The Cholesky decomposition unit performs Cholesky decomposition on the coefficient matrix of the symmetric positive definite normal equations to obtain a lower triangular matrix, followed by two triangular back-substitutions to obtain the parameter increments. The computational complexity of the Cholesky decomposition is proportional to the cube of the parameter dimension. Since the dimensions of the dynamic parameters and identity parameters are both low, the decomposition operation is very fast.

[0191] The parameter update unit adds the increment obtained from the solution to the current parameter vector, completing one Gauss-Newton iteration. This unit can also be configured with a line search or trust region mechanism to adaptively adjust the damping coefficient according to the energy decrease, but to ensure real-time performance, a simplified strategy with a fixed damping coefficient is used by default.

[0192] The Gauss-Newton solver module supports block coordinate descent scheduling, meaning the upper-level control module determines which parameter block is currently active, and the solver module only performs an update once for the active parameter block. This design allows the same solver to flexibly serve different scenarios such as dynamic parameter tracking and identity parameter optimization, simply by switching the input Jacobian and residuals.

[0193] Online tracking module:

[0194] The online tracing module is the system's scheduling and control center, responsible for coordinating the execution sequence of the entire online pipeline, maintaining its operational status, managing keyframe buffers, and controlling the allocation of identity optimization budgets. This module does not directly perform numerical calculations; instead, it acts as a state machine to drive the other five modules to run in the correct order.

[0195] The internal states of the online tracking module include: current frame number, shared identity parameters, previous frame dynamic parameters, current frame dynamic parameters, keyframe buffer array, identity optimization budget counter, and keyframe insertion interval counter. Upon arrival of each frame, the module schedules execution in a fixed order: first, dense prior prediction is run; then, vertex-by-vertex sampling is run; next, dynamic parameter tracking and updates are performed; then, keyframe insertion conditions are evaluated; and finally, the identity optimization step is executed based on the budget.

[0196] Keyframe buffers are organized as circular queues, with each slot storing the per-vertex target array, dynamic parameter vector, and head rotation matrix for that keyframe. The buffer capacity is fixed; when full, a replacement strategy is used to maximize pose coverage. The identity optimization budget is an integer counter; several steps of budget are added with each keyframe insertion, and one step is deducted with each identity optimization step, stopping when it reaches zero. This budget amortization mechanism ensures that identity optimization does not cause frame rate fluctuations.

[0197] The online tracking module also handles post-processing of the output data, including converting parametric face model vertices to world or camera coordinates, generating triangular meshes, and calculating normal vectors, for use in downstream applications such as virtual avatar driving and motion capture data export. The output interface supports multiple data formats and can be directly connected to 3D rendering engines or animation software.

[0198] System Deployment and Scalability:

[0199] This system can be deployed on various hardware platforms, from high-end desktop graphics processors to mobile graphics processors. On high-end desktop graphics cards, the online tracking mode can achieve high end-to-end frame rates, meeting the needs of professional-grade motion capture and virtual production. On mobile graphics processors, through network quantization and operator optimization, the system can still achieve real-time tracking speeds, meeting the needs of mobile augmented reality, virtual makeup try-on, and other applications.

[0200] Regarding system memory usage, the graphics processor's memory usage mainly comes from network weights and various cache tensors, which is sufficient for most mainstream consumer-grade graphics cards. The central processing unit's memory usage is relatively small, mainly used for video frame buffering and configuration data.

[0201] This system supports a wide range of input resolutions and two standard input frame rates: 30 frames per second and 60 frames per second. The output 3D mesh's vertex and triangle counts conform to the standard topology of parametric face models and can be directly used for subsequent processing such as expression-driven processing and texture mapping.

[0202] In summary, this system achieves high-precision and high-efficiency real-time monocular 3D face reconstruction by combining dense texture coordinate priors with a block coordinate descent Gaussian-Newton solver. It uniformly supports three modes: single image, offline sequence, and online tracking, and has broad application prospects and deployment flexibility.

[0203] Example 3

[0204] like Figures 3 to 6 These are intermediate and final result images of the method for reconstructing several user 3D face models according to the present invention. Figure 3 A to Figure 6 In this context, A represents the input monocular face image. Figure 3 B to Figure 6In the image, B represents the visualization of the correspondence between the face detection bounding box and texture coordinates and the image. Figure 3 C to Figure 6 In the diagram, C represents the visualization result of overlaying the reconstructed 3D face mesh with the input image. Figure 3 D to Figure 6 In the figure, D represents the final result of reconstructing the parametric face model into a 3D face model. As can be seen from the figure, the method of the present invention can stably output 3D face reconstruction results for users with different identities, head postures, and expressions in actual operation.

[0205] The above-described embodiments are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.

Claims

1. A high-precision real-time 3D face reconstruction method based on monocular video, characterized in that, Includes the following steps: S1. Input a monocular face video stream, initialize the keyframe buffer and determine the camera intrinsic parameters; perform single-image fitting on the first frame to obtain the initial identity parameters and dynamic parameters of the parameterized face model; S2. For each input frame, the image is fed into the prediction network to predict the correspondence map, the relative depth map, and their respective log-variance maps in the texture coordinate domain of the parameterized face model. S3. At the predefined texture coordinates of each parameterized face model vertex, sample each image obtained in step S2 to obtain the two-dimensional corresponding target, relative depth target, and vertex-by-vertex confidence weight determined by the logarithmic variance of the vertex, which are used as the fixed least squares target term for the frame. S4. Divide the face parameters into identity parameter blocks and dynamic parameter blocks; drive the parameterized face model with the two types of parameters to obtain the face vertices with pose, calculate the image projection coordinates and relative depth of each vertex, and subtract them from the corresponding target and relative depth target in step S3 to obtain the vertex-by-vertex residual; weight the residuals with confidence weights and add a regularization term to form the least squares fitting energy. S5. Fix one parameter block, linearize the residual of the other currently active parameter block, construct the Gauss-Newton normal equation with damping terms using the closed Jacobian and solve it, and update the active parameter block with the obtained increment. S6. For each newly arrived frame, use the dynamic parameters of the previous frame as initial values ​​and fix the shared identity. Update the dynamic parameters according to step S5 to complete the tracking of the frame. At the same time, maintain the key frame buffer. When updating the buffer, combine all key frame data to optimize the shared identity. Only use past frames and the current frame throughout the process and output the 3D face model of each frame in real time.

2. The high-precision real-time 3D face reconstruction method based on monocular video according to claim 1, characterized in that, The relative depth mentioned in step S2 is defined relative to the neck joint of the parameterized face model. The relative depth of any vertex is the camera space depth of that vertex minus the camera space depth of the neck joint. The relative depth is used to eliminate the inconsistency in absolute scale caused by random cropping and scaling during training, remove global translation components and focus on the local three-dimensional geometry of the face. The global translation is left to be recovered by the optimizer during the fitting stage.

3. The high-precision real-time 3D face reconstruction method based on monocular video according to claim 1, characterized in that, The prediction network described in step S2 consists of an image encoder and two texture coordinate domain decoding branches. The image encoder extracts multi-scale features from the input image. A set of learnable texture coordinate feature maps bound to the texture coordinate layout of the parameterized face model is used as a query. Each decoding branch fuses the query with the image features and generates a texture coordinate domain feature map. Correspondence and relative depth use their own independent decoding branches. Each branch is followed by a dense prediction head to predict the target map and its log-variance map, respectively.

4. The high-precision real-time 3D face reconstruction method based on monocular video according to claim 1, characterized in that, The vertex confidence weights mentioned in step S3 are obtained by exponential transformation of the logarithmic variance obtained from sampling the vertex. The higher the prediction uncertainty, i.e. the larger the logarithmic variance, the lower the weight of the vertex in the fitting energy.

5. The high-precision real-time 3D face reconstruction method based on monocular video according to claim 1, characterized in that, In step S4, the dynamic parameter block energy is composed of the weighted corresponding residuals and weighted relative depth residuals per vertex, plus expression and pose regularization terms; The identity parameter block energy is composed of the vertex-weighted residual accumulated from multiple keyframes, plus an identity regularization term; the identity regularization term is the L2 norm regularization of the identity parameters with respect to the zero vector.

6. The high-precision real-time 3D face reconstruction method based on monocular video according to claim 1, characterized in that, In step S5, the residual vector of the currently active parameter block is linearized, and a Gauss-Newton normal equation with damping terms is constructed. The parameter increment is obtained by Choleski decomposition and the active parameter block is updated accordingly. When updating the dynamic parameter block and the identity parameter block, the closed Jacobian of the residual with respect to the dynamic parameter and the identity parameter is used respectively. The closed Jacobian is calculated by a custom graphics processor operator.

7. A high-precision real-time 3D face reconstruction system based on monocular video, characterized in that, The system includes an initialization module, a dense prior prediction module, a vertex-by-vertex sampling module, a least-squares modeling module, a Gauss-Newton solving module, and an online tracking module. The initialization module receives the monocular face video stream and outputs the initial identity parameters, initial dynamic parameters, and camera intrinsic parameters of the parameterized face model, completing the keyframe buffer initialization. The dense prior prediction module receives the input frame image and outputs a correspondence map, a relative depth map, and a log-variance map of the two in the texture coordinate domain of the parameterized face model. The vertex-by-vertex sampling module samples each map at the texture coordinates of each vertex of the parameterized face model and outputs the vertex-by-vertex two-dimensional corresponding target, relative depth target, and confidence weight. The least squares modeling module divides the face parameters into identity parameter blocks and dynamic parameter blocks, calculates and weights the vertex-by-vertex residuals, and outputs the least squares fitting energy by combining the regularization term. The Gauss-Newton solution module adopts a block coordinate descent method, fixes one parameter block and linearizes the residuals of the other activated parameter block, constructs a damped Gauss-Newton normal equation using a closed Jacobian, and outputs the parameter increment. The online tracking module updates the current frame's dynamic parameters with the previous frame's dynamic parameters as initial values, maintains a keyframe buffer, and jointly optimizes shared identities when updating the buffer, outputting a 3D face model for each frame.

8. The high-precision three-dimensional real-time face reconstruction system based on monocular video according to claim 7, characterized in that, The dense prior prediction module includes an image encoder and two texture coordinate domain decoding branches; the image encoder extracts multi-scale features of the input image; The learnable texture coordinate feature map, which is bound to the texture coordinate layout of the parameterized face model, is used as the query. Each decoding branch fuses the query and image features to generate a texture coordinate domain feature map. The correspondence and relative depth are output as target map and log-variance map through independent decoding branches and dense prediction head, respectively.

9. The high-precision three-dimensional real-time face reconstruction system based on monocular video according to claim 7, characterized in that, The Gauss-Newton solver module linearizes the residual vector of the current active parameter block, constructs the normal Gauss-Newton equation with damping terms, and uses Choleski decomposition to obtain the parameter increment and update the active parameter block. When updating the dynamic parameter block and the identity parameter block, the closed Jacobian of the residual with respect to the dynamic parameter and the identity parameter is used respectively. The closed Jacobian is calculated in parallel by the graphics processor.

10. The high-precision three-dimensional real-time face reconstruction system based on monocular video according to claim 7, characterized in that, The online tracking module maintains a fixed-capacity keyframe buffer. Every few frames, it calculates the head rotation of the current frame and the minimum geodesic distance between the current frame and the head rotation of existing keyframes in the buffer on the 3D rotation group as the novelty. When the buffer is not full and the novelty exceeds the threshold, the current frame is inserted. When the buffer is full, it is only inserted if replacing a keyframe can improve the pose coverage. Each buffer update adds an identity optimization budget, which is distributed to several subsequent frames. Each frame performs at most one Gauss-Newton iteration to incrementally improve the shared identity parameters.