4D face reconstruction method and device
By combining a 3D face deformation model with a generative adversarial network, the problems of geometric topological inconsistency and texture flicker in 4D face reconstruction are solved, generating a high-fidelity 4D dynamic face mesh sequence, which improves reconstruction efficiency and accuracy.
Patent Information
- Application Number
- CN202511373759.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Existing 4D face reconstruction technologies suffer from problems such as geometric topology inconsistency, texture flickering artifacts, and geometric jitter. They are also characterized by high hardware costs, complex deployment, and the tendency of algorithms to lose facial details. Furthermore, the acquisition process is cumbersome and the degree of automation in post-processing is low, making it difficult to balance accuracy and efficiency.
A 3D human face deformation model is used for fitting and optimization. Combined with generative adversarial network to process texture map, temporal smoothing filtering is performed to generate a high-fidelity, temporally coherent 4D dynamic human face mesh sequence.
It solves the problems of geometric topology inconsistency and texture flicker, improves the fidelity of facial details, achieves efficient 4D face reconstruction, and reduces hardware costs and processing complexity.
Smart Images

Figure CN120876784A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of face reconstruction, and in particular to a 4D face reconstruction method and device. Background Technology
[0002] 4D face capture and reconstruction, by capturing dynamic 3D facial geometry and texture sequences that include a time dimension, overcomes the limitations of traditional static 3D reconstruction and has significant application value. In the medical field, it can record dynamic facial changes, assisting in the assessment of facial paralysis recovery and the planning of craniofacial surgery; in the AR / VR field, it can accurately map facial expressions to virtual avatars, enhancing the realism of interaction; in film and television production, it can directly drive the expressions of digital characters, reducing production costs; in identity authentication, dynamic features can improve anti-spoofing capabilities, providing core technical support for high-fidelity dynamic face applications in multiple fields.
[0003] Current 4D capture technology faces several key bottlenecks, including but not limited to: geometric topology inconsistencies caused by raw scan data, texture flicker artifacts caused by camera exposure variations, and geometric jitter caused by minute inter-frame movements. Furthermore, existing solutions suffer from high hardware costs, complex deployment, multi-view synchronization errors affecting accuracy, algorithms prone to losing facial details or exhibiting poor generalization, cumbersome acquisition processes, and low automation in post-processing, making it difficult to balance accuracy and efficiency. These problems urgently need to be addressed. Summary of the Invention
[0004] The purpose of this application is to provide a 4D face reconstruction method and device that can reconstruct high-fidelity, temporally coherent 4D dynamic face mesh sequences end-to-end.
[0005] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a 4D face reconstruction method, including: Acquire multiple groups of face images within a preset time period; For each group of face images, reconstruct a triangular mesh model of the face; For each of the aforementioned face triangular mesh models, a three-dimensional face deformation model is used for fitting and optimization to obtain multiple corresponding registered face mesh models. For any of the aforementioned face image groups, after flicker removal processing, an initial texture map is extracted, and then the initial texture map is mapped to the corresponding registered face mesh model to obtain a textured mesh. For all the textured meshes corresponding to the face image groups, temporal smoothing filtering is performed to obtain a 4D dynamic face mesh sequence.
[0006] In a second aspect, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of a 4D face reconstruction method.
[0007] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application reconstructs a triangular mesh model of the face based on multiple images in each face image group, thereby obtaining an accurate single-frame facial mesh model. Then, a three-dimensional face deformation model is used for fitting and optimization to obtain multiple corresponding registered face mesh models. Thus, through the processing of the three-dimensional face deformation model, multiple triangular mesh models of the face are transformed into a unified registered face mesh model with a semantically corresponding mesh structure. This processing solves the problem of geometric topological inconsistency in existing data. Secondly, after deflickering processing for any face image group, an initial texture map is extracted and then upsampled using a generative adversarial network. Thus, a texture map with high resolution and no flicker artifacts can be obtained. Then, the texture map is mapped to the corresponding registered face mesh model to obtain a textured mesh. This processing solves the problems of texture flicker and artifacts and improves fidelity. Finally, temporal smoothing filtering is performed on the textured meshes corresponding to all face image groups to eliminate possible minor jitter in the mesh sequence, thereby obtaining a temporally coherent and smoothly transitioned 4D dynamic face mesh sequence. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a flowchart illustrating a 4D face reconstruction method according to an embodiment of this application.
[0010] Figure 2 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0012] This application proposes a complete end-to-end 4D face reconstruction method for reconstructing high-fidelity, temporally coherent 4D facial mesh sequences from multi-view 3D scan data.
[0013] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0014] In one exemplary embodiment, such as Figure 1 As shown, a 4D face reconstruction method is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, it includes the following steps 101 to 105.
[0015] Step 101: Obtain multiple groups of face images within a preset time period.
[0016] In a practical application, a precisely calibrated three-view structured light scanning system is built and electrically connected to the aforementioned computer equipment. The three-view structured light scanning system includes three 720p resolution... The 540 industrial camera simultaneously captures images from the left, front, and right sides of the subject (i.e., the target person), obtaining a left-side facial image, a front-side facial image, and a right-side facial image, which together form a set of facial images.
[0017] During the image capture process of an industrial camera, a camera projection model is established using the Zhang Zhengyou calibration method: .
[0018] in,( u , v ) are pixel coordinates, ( , , () represents the coordinates in the world coordinate system, and the intrinsic parameter matrix. Including focal length ( , ) and principal point ( , ), external reference The solution is obtained through chessboard calibration, where R is the rotation matrix. This is a translation vector. The use of structured light technology can actively project coded patterns onto the surface of a face, and calculate 3D information by decoding the deformation of the pattern. This ensures that high-precision geometric data can be obtained even in areas with fewer texture features (such as the cheek).
[0019] Step 102: For each group of face images, reconstruct the face triangular mesh model; including the following steps: (21) The phase shift method is used to perform three-dimensional reconstruction on the left face image, the front face image and the right face image respectively to obtain left point cloud data, front point cloud data and right point cloud data. These three point cloud data are three independent single-view point clouds.
[0020] (22) The iterative nearest point algorithm based on feature point matching is used to register the left-side point cloud data, the front-side point cloud data, and the right-side point cloud data, and then fuse them to obtain a face point cloud data. The function formula used in the iterative nearest point algorithm based on feature point matching is shown below, and this formula uses camera calibration parameters: .
[0021] in, To achieve the fusion of corresponding points in the two point clouds, corresponding to the left point cloud data and the front point cloud data in this application, the left and front point cloud data after fusion are the right point cloud data; For rigid body transformation, This refers to the number of points in the left-side, front-side, or right-side point cloud data. After precise registration and fusion of the three point clouds using the above function formula, they are stitched together to form a complete, non-overlapping facial point cloud data covering the front of the face. The point cloud is unified under the same world coordinate system, and its data size is usually between 70,000 and 100,000 points, providing a high-quality raw data source for subsequent fine processing.
[0022] Thus, this application obtains accurate, high-density single-frame facial geometric information from a set of three images captured at each time point t, thereby providing a data foundation for subsequent processing.
[0023] (23) The face point cloud data is simplified using the weighted Poisson disk resampling method to obtain simplified face point cloud data. Since the face point cloud data has high redundancy, direct processing is computationally inefficient and unnecessary. Therefore, this step simplifies and standardizes the point cloud and reconstructs it into a more easily manipulated triangular mesh surface. The specific steps are as follows: (1) Voxelization: The disordered facial point cloud data is divided into regular 3D meshes (voxels); the point cloud space is divided into sections with sides of length 1. A voxel mesh is used to calculate the surface area of each voxel. Predict the sampling radius of the Poisson disk ( (This is a scaling factor). This process not only enables spatial indexing of point clouds, but also allows for the estimation of the surface area of a face by statistically analyzing the number and distribution of non-empty voxels, providing an adaptive and reasonable basis for the next step of determining the sampling radius.
[0024] (2) Poisson disk sampling: Based on the Poisson disk sampling radius estimated in the previous step Poisson disk sampling is performed on the surface of the facial point cloud data. This process ensures that the distance between any two sampling points is not less than a specified radius, and a hash nearest neighbor search algorithm is used to accelerate the processing. Compared with random sampling, it can generate sampling results with a more uniform distribution and avoid excessive point clustering, thus better preserving the overall shape of the surface and generating a subset of sampling points. .
[0025] (3) Point Count Control: To ensure consistent processing efficiency across different frames and meet the input size requirements of subsequent algorithms, the number of point clouds is precisely controlled to a preset target value by iteratively deleting the closest point pairs in the current point set. .
[0026] (4) Voronoi Relaxation: To further optimize the distribution quality of the point cloud, the following steps are taken: Iterative relaxation is performed. In each iteration, the local tangent space Voronoi diagram for each point is calculated, and the point is moved to the centroid of its corresponding Voronoi cell. This process allows the point distribution to reach a local optimum, improving the smoothness and uniformity of the point cloud, ultimately resulting in a high-quality simplified point cloud. .
[0027] Through the above steps, the key geometric features of the face can be preserved to the greatest extent while significantly reducing the number of points (e.g., reducing an 80,000-point facial point cloud data to 20,000 points through point cloud simplification).
[0028] (24) The simplified facial point cloud data is reconstructed using the Poisson reconstruction algorithm to obtain a triangular mesh model of the face. That is, by using the Poisson reconstruction algorithm, the discrete and optimized point cloud data is reconstructed into a surface. Connected to form a triangular mesh with a continuous surface This process provides a topological structure for subsequent geometric operations, such as feature point extraction and deformation analysis.
[0029] It should be noted that the grid sequence generated at this time The topology of the mesh is not completely consistent between frames; that is, the number of vertices, indices, and connection relationships of triangles are different in each frame. To solve this problem of topological inconsistency after meshing, this application establishes a unified mesh structure with semantic correspondence for all frames through the following step 103, which is the basis for realizing 4D animation and analysis.
[0030] Step 103: For each of the face triangular mesh models, a three-dimensional face deformation model is used for fitting and optimization to obtain multiple corresponding registered face mesh models.
[0031] In a specific application, taking a triangular mesh of a face as an example, a 3D face deformation model is used for fitting and optimization to obtain the registered face mesh, including the following steps: (31) Project the face triangular mesh model orthogonally onto a two-dimensional plane to obtain the corresponding depth map; to drive the model for accurate alignment, it is necessary to start from the mesh. Semantically consistent feature points are extracted. Therefore, the face triangular mesh is orthogonally projected onto a two-dimensional image plane to generate a depth map.
[0032] (32) Based on a preset set of facial feature points, facial feature points are identified and labeled on the depth map to obtain multiple two-dimensional target facial feature points; specifically, using the powerful dlib face feature detection library, 68 standard facial feature points are automatically located on the depth map. If necessary, 17 points representing the facial contour can be removed, retaining only 51 key feature points inside the face (such as eyebrows, eyes, nose, and mouth). These 51 key feature points are the points in the preset set of facial feature points, and the semantics of these points are strictly aligned with the definitions of the BFM and FLAME three-dimensional face deformation models. The formula is as follows: .
[0033] in, For perspective projection, 51 standard facial feature points were detected.
[0034] (33) The multiple two-dimensional target facial feature points are mapped back to the face triangular mesh model to obtain multiple three-dimensional target facial feature points; specifically, by querying the projection relationship, the coordinates of the above 51 two-dimensional target facial feature points are mapped back to The three-dimensional surface is used to obtain a set of accurate three-dimensional target facial feature points. The data is stored in .npy format. Furthermore, after storage, the 3D target facial feature points can be visualized, with small arrows indicating the orientation of the feature points' normals.
[0035] (34) Based on multiple three-dimensional target facial feature points, with the face triangular mesh model as the target, a three-dimensional face deformation model is used for fitting and optimization to obtain the corresponding registered face mesh model.
[0036] In a specific application, FLAME is a statistically based, differentiable, parameterized head model that can control complex head geometry through a set of low-dimensional parameters. Its model can be represented as follows: The parameterized mesh model has 5023 vertices and 9976 mesh faces. If only the FLAME model is used for mesh fitting, some facial details will inevitably be lost. Compared to the FLAME model, the Basel Face Model (BFM) provides a more detailed description of the face. The standard BFM face model has 53490 vertices, while the model excluding the neck and ears has 35709 vertices and 70789 mesh faces. To fit a more detailed and accurate full-head face mesh model, this application uses a new 3D face deformation model obtained by fusing the above two models for final mesh fitting. That is, the 3D face deformation model is a parameterized face model obtained by fusing the BFM and FLAME models.
[0037] In the step of fitting and optimizing using a 3D face deformation model, the frontal area of the face is fitted using the BFM model, while the non-frontal facial areas (such as the back of the head, ears, and neck) are fitted using the FLAME face model. This results in a 3D face deformation model that combines the advantages of BFM's higher accuracy of the facial area with FLAME's more complete head shape. The final fitted full-head face model has 37,550 vertices and 74,993 faces.
[0038] The final 3D face deformation model obtained through registration is a new parametric model obtained by fusing the frontal face region of BFM and the non-frontal face region of FLAME. Its optimization process uses the following energy function, with the feature points being 51 internal facial feature points; then, the closest surface distance between the original triangular mesh and the fitted optimized model is calculated; the shape and expression parameters are the same for both BFM and FLAME, both being 100-dimensional; the pose parameters use the parameters of the FLAME model, while the BFM model does not include pose parameters, therefore a total of five optimization terms are used: 1) Facial feature points: Both BFM and FLAME have 51 facial feature points, which are the same.
[0039] 2) s2m: is the closest distance between the fitted model and the surface of the face triangular mesh model, i.e., the front is fitted to BFM, and the back of the head and neck are fitted to FLAME.
[0040] 3) Shape parameters: Use 199 dimensions of BFM.
[0041] 4) Pose parameters: Use FLAME's 15-dimensional model because only the frontal BFM of the human face has no pose parameters.
[0042] 5) Expression parameters: Both BFM and FLAME are 100-dimensional, and BFM is actually used.
[0043] The above five optimization terms can be optimized using only one energy function. Specifically, the goal of this application is to optimize the parameters so that the generated mesh fits the input scan mesh as closely as possible. The optimization process minimizes the following energy function. accomplish: .
[0044] in, , These are the weighting coefficients; The feature point alignment term is used to calculate the sum of the three-dimensional Euclidean distances between the facial feature points on the fitted three-dimensional face deformation model and the three-dimensional target facial feature points. This is the key to driving the model to perform macroscopic alignment. This is the dense surface distance term, used to measure the nearest distance from all vertices on the face triangle mesh to the fitted model surface; These are the weighting coefficients for the shape parameters. These are the weighting coefficients for the attitude parameters. These are the weighting coefficients for the facial expression parameters. The notation for norm calculation; , which are the shape parameters that control identity in the 3D face deformation model; , which are the parameters controlling head posture (including global rotation and mandibular movement) in the 3D face deformation model; , which are the parameters that control facial expression changes in a 3D facial deformation model.
[0045] The function formula for the feature point alignment term is: .
[0046] in, These are the facial feature points of the three-dimensional target. The number of facial feature points in the three-dimensional target can be 51; To fit the facial 3D feature points of the model, , which are parameters that control global translation; The symbol for calculating the second norm; To improve robustness, a Geman-McClure robust kernel function is introduced into the dense surface distance term. ,in, It is the value of the surface distance calculated between the two meshes, specifically the three-dimensional Euclidean distance between all the mesh vertices of the fitted three-dimensional face deformation model and all the vertices of the face triangular mesh; It is a parameter with a very small value, controlling the degree of suppression of large errors. The Geman-McClure robust kernel function reduces the weight of outliers (such as noise or occluded parts in the scanned data) in the optimization process, thereby preventing them from having an excessive impact on the fitting results.
[0047] In addition, the regularization term: through , The magnitudes of shape, expression, and pose parameters are penalized separately to constrain the model's output to a "reasonable" range and prevent bizarre geometric shapes that do not conform to human facial anatomy.
[0048] The optimization process based on the above energy function follows a "coarse-to-fine" strategy, consisting of two steps: First, only the rigid transformation parameters (global rotation) are optimized. Peaceful relocation ), complete the initial rigid alignment; then, based on this, adjust all parameters (including shape) ,expression and posture The non-rigid parameters are jointly optimized until the energy function converges. This process ultimately generates a registered face mesh with the same number of vertices and connectivity for each frame in the sequence. .
[0049] Step 104: For any of the face image groups, after performing flicker removal processing, an initial texture map is extracted, and then the initial texture map is mapped to the corresponding registered face mesh model to obtain a textured mesh.
[0050] In a specific application, for any of the aforementioned face image groups, after diffusing processing, an initial texture map is extracted. Through this processing, a set of high-resolution (4K), flawless, and color-realistic texture maps can be generated and bound to the obtained temporally consistent geometry. The specific steps are as follows: (41) Using an image sequence flicker elimination method based on clustering and multi-scale histogram matching, the frontal facial images in multiple face image groups are subjected to flicker removal processing to obtain multiple optimized frontal images.
[0051] The image sequence flicker elimination method based on clustering and multi-scale histogram matching includes the following steps: 1) Robust preprocessing: Non-local means denoising is applied to the frontal face image sequence in the face image group to obtain the denoised frontal face image sequence. The calculation formula is as follows: .
[0052] in, This is a frontal facial image after noise reduction processing. As a weight normalization factor, it is necessary to ensure that the sum of the normalized weights is 1 to avoid image brightness distortion; Pixels in a frontal facial image The intensity value, It is a pixel Search neighborhood, weight Reflects pixels and pixels The similarity between image patches centered on the image is typically measured using Gaussian-weighted Euclidean distance, i.e.: .
[0053] Where h is the filtering parameter. In pixels The image patch centered on In pixels The algorithm uses the self-similarity of images to perform a weighted average of similar neighboring blocks, which can effectively remove noise while preserving edge and texture details to the maximum extent.
[0054] 2) Image Feature Extraction and Clustering: The intrinsic shape feature vector in the color histogram space is extracted from the denoised frontal facial image. Specifically, this application selects a multi-channel normalized color histogram as the intrinsic shape descriptor of the image's color space distribution. The histogram is essentially a discrete approximation of the probability density function of the pixel intensity distribution. For the denoised frontal facial image... Calculate its color histogram features. The histogram features are the normalized histograms of each color channel (e.g., R, G, B channels). The combination of , the formula is: .
[0055] in, Let i be the pixel count of gray level i in the c-th channel. To avoid division by zero errors with extremely small positive numbers, the three-channel histogram features are ultimately concatenated into an feature vector. This vector represents the overall luminous properties of the image and can be viewed as the image's coordinates in the color histogram space.
[0056] The ingenuity of this method lies in the fact that it does not rely on any prior flickering patterns. After obtaining the feature vectors of each denoised frontal facial image, it uses a data-driven approach and the K-Means clustering algorithm to adaptively divide the denoised frontal facial image sequence into a "standard group" with normal visual features and several "abnormal groups" with abnormal visual characteristics in the manifold space according to the distribution of feature vectors.
[0057] 3) Multi-scale histogram matching: For each frame of a frontal facial image in the "outlier group," its color is corrected by performing histogram matching with the "standard group." To balance global tone and local detail, matching is performed on overlapping image blocks at multiple scales. The final color map lookup table is generated by weighted fusion of matching results from different scales using a Gaussian kernel function, which ensures complementary information and effectively avoids "blocking artifacts."
[0058] Specifically, the images in the selected "standard group" are divided into blocks at multiple preset scales, the cumulative distribution function of each color channel is calculated, and stored as a reference CDF (Cumulative Distribution Function) map. The images in the "abnormal group" are also divided into blocks at multiple scales, and the CDF of each color channel is calculated. By matching with the reference CDF image of the corresponding reference image block, a lookup table is generated. The pixel values of the "abnormal group" image blocks are adjusted using the lookup table so that the histogram distribution is close to that of the "standard group" image blocks. Finally, the matching results obtained from different scales are weighted and fused to obtain the matched image.
[0059] 4) Smooth inter-group temporal transition: At the switching boundary between different image groups (e.g., from the standard group to the abnormal group), this application applies a "cross-fading" technique. This technique, within a preset 2N frame transition interval, transitions the last image of the previous image segment... Image matching the current target image segment Perform a weighted average. The fusion result of the nth frame. for: .
[0060] Among them, the fusion coefficient ( For the number of fusions, 0 The value gradually changes from 0 to 1 with the number of frames. This technique ensures that the visual perception is continuous and without jumps at the switching boundary between the abnormal group and the standard group. Finally, the image frames processed by the above steps (including image denoising, grouping, matching, and smooth transition) are arranged in their original time order, and the output is a stable image sequence without obvious flickering, that is, an optimized frontal image sequence.
[0061] The above processing can solve the problem of video texture flickering caused by camera auto-exposure, white balance fluctuations, and periodic changes in structured light illumination modes.
[0062] (42) Using the HRN (High-Resolution Network) method, a texture map is generated based on the optimized frontal image. Specifically, the HRN method is used to generate an initial high-quality texture map with a resolution of 256 using the flicker-free frontal image. 256.
[0063] (43) The texture map is upsampled using a super-resolution model based on generative adversarial networks to obtain an enhanced texture map; wherein, the super-resolution model of the generative adversarial network is: This is to upsample the texture map to achieve a resolution of 4K (3840). 2160) level, and generate more realistic details.
[0064] Since the texture map generated by HRN is based on the UV layout of the BFM face model, and the geometry in this application is based on the new three-dimensional face deformation model after the fusion of BFM and FLAME, in order to solve this problem, this application sets up a texture coordinate mapping program from BFM to FLAME, specifically as follows (44).
[0065] (44) A preset sparse correspondence matrix is used to perform feature transformation on the enhanced texture map to obtain an initial texture map. The following formula is used when performing feature transformation on the enhanced texture map using the preset sparse correspondence matrix: .
[0066] in, As a predefined sparse correspondence matrix, it encodes the linear mapping relationship from the BFM vertex space to the FLAME vertex space; For texture coordinates in BFM space, These are texture coordinates in FLAME space.
[0067] After the above steps (41)-(44), each registered face mesh is processed according to the texture space of the new fused 3DMM model. Bind the corresponding 4K initial texture map and generate the corresponding material library file (.mtl), finally obtaining a textured mesh sequence. .
[0068] Step 105: Perform temporal smoothing filtering on the textured mesh corresponding to all the face image groups to obtain a 4D dynamic face mesh sequence.
[0069] Although mesh model registration ensures topological consistency, parameter optimization for each frame is performed independently during actual fitting. This can lead to minute, high-frequency jitter or "trembling" in vertex positions over time, affecting the smoothness and realism of the final 4D animation. To eliminate the minute jitter that may exist in the final geometric sequence and make the animation transition smoother, this application designs the following anti-flickering algorithm. Step 105 includes the following: (51) Stack the vertex data of the textured mesh corresponding to all the face image groups into a tensor. (T is the number of frames, N is the number of vertices).
[0070] (52) The tensor is moved along the time axis using a preset sliding window, and Gaussian weighted average is performed on the vertex data corresponding to all frames in each window to calculate the vertex position after smoothing the center frame, thus obtaining a 4D dynamic human face mesh sequence.
[0071] Specifically, a size of An odd-numbered preset sliding window moves along the time axis T. For each frame within the window, the smoothed vertex position of its center frame is calculated. The calculation formula is: .
[0072] Where i is the frame offset within the sliding window (e.g., when W=5, i=-2,-1,0,1,2), and n is the vertex index. For the first Frame number The original three-dimensional coordinates of each vertex The Gaussian weighting principle dictates that frames closer to the current center frame should have higher weights and contribute more, and vice versa. This aligns with the inertia of motion, resulting in a more natural smoothing effect. The calculation formula is as follows: .
[0073] in, W is the normalization factor, ensuring that the sum of all weights is 1; W is the preset size of the sliding window. The standard deviation is Gaussian. This step effectively filters out high-frequency noise, ultimately outputting a highly realistic and time-smooth 4D dynamic human face mesh sequence in terms of geometry and texture.
[0074] In summary, this application designs a sophisticated workflow, beginning with high-quality data acquisition. Through innovative point cloud preprocessing and standardized geometric registration, a temporally consistent topology is established. Subsequently, a novel image deflickering algorithm based on cluster analysis is employed, combined with deep learning-based texture generation and remapping techniques, ultimately producing realistic textures with 4K resolution. Finally, temporal smoothing filtering ensures that the final generated 4D dynamic avatar achieves a smooth and natural effect in both geometry and appearance.
[0075] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 2 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), a 4D face acquisition device, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The 4D face acquisition device includes a structured light projection unit and a multi-view image acquisition unit, which establish a data connection with the system bus via the I / O interfaces, serving as the data input source for 4D face reconstruction. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the 4D face reconstruction method.
[0076] Those skilled in the art will understand that Figure 2 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0077] In one exemplary embodiment, a computer device is also provided, including a memory, a processor, a 4D face acquisition device, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the 4D face reconstruction method.
[0078] In one exemplary embodiment, a 4D face acquisition device is provided, including a structured light projection unit (for actively projecting stripes onto the surface of a target person's face and calculating the three-dimensional information of the face through pattern deformation) and a multi-view image acquisition unit (including three industrial cameras for acquiring face images of the target person from three perspectives: left, front, and right). When the computer program is executed by a processor, it implements the steps in the above-described method embodiments.
[0079] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0080] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0081] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0082] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by hardware related to computer program instructions. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0083] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0084] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0085] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A 4D face reconstruction method, characterized in that, The method includes: Acquire multiple groups of face images within a preset time period; For each group of face images, reconstruct a triangular mesh model of the face; For each of the aforementioned face triangular mesh models, a three-dimensional face deformation model is used for fitting and optimization to obtain multiple corresponding registered face mesh models. For any of the aforementioned face image groups, after flicker removal processing, an initial texture map is extracted, and then the initial texture map is mapped to the corresponding registered face mesh model to obtain a textured mesh. For all the textured meshes corresponding to the face image groups, temporal smoothing filtering is performed to obtain a 4D dynamic face mesh sequence.
2. The 4D face reconstruction method according to claim 1, characterized in that, Each facial image group includes the target person's left facial image, frontal facial image, and right facial image; Reconstructing a triangular mesh model of the face for each group of face images, including: The phase shift method is used to perform three-dimensional reconstruction on the left face image, the front face image and the right face image respectively to obtain left point cloud data, front point cloud data and right point cloud data; An iterative nearest point algorithm based on feature point matching is used to perform point cloud registration on the left-side point cloud data, the front-side point cloud data, and the right-side point cloud data, and then fuse them to obtain a human face point cloud data. The face point cloud data is simplified by using a weighted Poisson disk resampling method to obtain simplified face point cloud data. The Poisson reconstruction algorithm is used to reconstruct the simplified facial point cloud data into a surface to obtain a triangular mesh model of the face.
3. The 4D face reconstruction method according to claim 1, characterized in that, For a given triangular mesh model of a face, a 3D face deformation model is used for fitting and optimization to obtain a registered face mesh model, including: The face triangular mesh model is orthogonally projected onto a two-dimensional plane to obtain the corresponding depth map; Based on a preset set of facial feature points, facial feature points are identified and marked on the depth map to obtain multiple two-dimensional target facial feature points; Multiple two-dimensional target facial feature points are mapped back to the face triangular mesh model to obtain multiple three-dimensional target facial feature points. Based on multiple three-dimensional target facial feature points, and taking the face triangular mesh model as the target, a three-dimensional face deformation model is used for fitting and optimization to obtain the corresponding registered face mesh model.
4. The 4D face reconstruction method according to claim 3, characterized in that, The three-dimensional face deformation model is a parametric face model obtained by fusing the BFM model and the FLAME model; the steps of fitting and optimizing using the three-dimensional face deformation model include: The BFM model is used to fit the frontal facial region of the human face, and the FLAME model is used to fit the non-frontal facial region. The optimization process involves minimizing the following energy function. accomplish: ; in, , These are the weighting coefficients; The feature point alignment term is used to calculate the sum of the three-dimensional Euclidean distances between the facial feature points on the fitted three-dimensional face deformation model and the three-dimensional target facial feature points; This is the dense surface distance term, used to measure the nearest distance from all vertices on the face triangle mesh to the fitted model surface; These are the weighting coefficients for the shape parameters. These are the weighting coefficients for the attitude parameters. These are the weighting coefficients for the facial expression parameters. The notation for norm calculation; , which are the shape parameters that control identity in the 3D face deformation model; , which are the parameters controlling head posture in the 3D human face deformation model; , which are the parameters that control facial expression changes in a 3D facial deformation model.
5. The 4D face reconstruction method according to claim 4, characterized in that, The function formula for the feature point alignment term is: ; in, For three-dimensional target facial feature points, The number of facial feature points in the three-dimensional target; To fit the obtained FLAME model, , which are parameters that control global translation; The symbol for calculating the second norm; The dense surface distance term incorporates a Geman-McClure robust kernel function. ;in, It is the value of the surface distance calculated between the two meshes. It is a preset minimum parameter.
6. The 4D face reconstruction method according to claim 1, characterized in that, For any of the aforementioned face image groups, after flicker removal processing, the initial texture map is extracted, including: An image sequence flicker removal method based on clustering and multi-scale histogram matching is used to perform flicker removal processing on frontal facial images in multiple face image groups to obtain multiple optimized frontal images; The HRN method is used to generate a texture map based on the optimized frontal image; An enhanced texture map is obtained by upsampling the texture map using a super-resolution model based on generative adversarial networks. A preset sparse correspondence matrix is used to perform feature transformation on the enhanced texture map to obtain an initial texture map.
7. The 4D face reconstruction method according to claim 6, characterized in that, When performing feature transformation on the enhanced texture map using a preset sparse correspondence matrix, the following formula is used: ; in, This is a pre-defined sparse correspondence matrix; For texture coordinates in BFM space, These are texture coordinates in FLAME space.
8. The 4D face reconstruction method according to claim 1, characterized in that, For all the textured meshes corresponding to the aforementioned face image groups, temporal smoothing filtering is performed to obtain a 4D dynamic face mesh sequence, including: Stack the vertex data of the textured mesh corresponding to all the face image groups into a tensor; The tensor is moved along the time axis using a preset sliding window, and a Gaussian weighted average is performed on the vertex data corresponding to all frames within each window to calculate the vertex position after smoothing the center frame, thus obtaining a 4D dynamic human face mesh sequence.
9. The 4D face reconstruction method according to claim 8, characterized in that, The smoothed vertex position The calculation formula is: ; Where i is the frame offset within the sliding window, and n is the vertex index. To smooth the first 1 Frame number vertex positions, The Gaussian weights are calculated using the following formula: ; in, W is the normalization factor; W is the preset size of the sliding window. The standard deviation is Gaussian.
10. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of the 4D face reconstruction method according to any one of claims 1-9.
Citation Information
Patent Citations
Three-dimensional human face reconstruction method based on single picture
CN108765550A
Digital human synthesis method, computer program product, device and medium
CN118379438A
Face three-dimensional modeling method based on RGBD image
CN120279169A
Moving face model modeling method and device and computer equipment
CN120599172A
Texture mapping method and apparatus based on three-dimensional face reconstruction, device and storage medium
WO2025015834A1
Cited By
Dynamic face reconstruction method and device based on partial differential equation
CN121982224A
A dynamic face reconstruction method and device based on partial differential equations
CN121982224B