Three-dimensional dynamic face reconstruction method based on mixed features and regional expressions
By mixing features and regional expressions, hash conflict and high model complexity problems are solved, and the accuracy and training speed of three-dimensional dynamic face reconstruction are improved, especially in dynamic faces and complex expression changes scenes, achieving higher detail fidelity and color prediction stability.
Patent Information
- Application Number
- CN202510466172.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-25
AI Technical Summary
The existing three-dimensional dynamic face reconstruction method has challenges in hash conflict and high model complexity issues, resulting in a slow reconstruction accuracy and slow training speed, especially in dynamic faces and complex expression changes scenes, and detailed information is seriously lost.
Using a method of mixing features and regional expressions, key expression dimensions are screened through multi-resolution hash coding and analysis of variance, combining view-independent and related color feature branches, color and geometric features of three-dimensional spatial points are generated, reducing hash conflicts and reducing model calculation complexity.
It improves the accuracy and training efficiency of three-dimensional face reconstruction, enhances the stability and detail fidelity of color prediction, and improves the flexibility and speed of the model in dynamic face reconstruction.
Smart Images

Figure CN120374854A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and particularly relates to a three-dimensional dynamic face reconstruction method based on hybrid features and regional expressions. Background Art
[0002] Three-dimensional dynamic face reconstruction is an important research topic in computer vision and computer graphics, and is widely used in fields such as virtual social interaction, film and television production, and medical simulation. The core difficulty of dynamic face reconstruction lies in the fact that the geometric structure and appearance attributes of the human face are extremely complex and delicate, and high-precision and high-resolution reconstruction results are required to achieve a real and vivid visual effect.
[0003] Traditional face reconstruction methods usually rely on high-quality video data from multiple perspectives or use predefined parametric models to limit the range of face variations. Although these methods can provide good reconstruction effects, they have certain limitations in terms of reconstruction quality, generalization ability, and extensiveness in practical applications. For example, the acquisition cost of multi-view data is relatively high, while parametric models may not be able to fully express complex facial expressions and geometric shapes, restricting the authenticity and flexibility of the reconstruction effect. In recent years, Neural Radiance Fields (NeRF) technology has demonstrated its powerful ability to perform high-quality novel view synthesis in the absence of precise geometric information. The NeRF-based dynamic face reconstruction method can directly learn the implicit representation of the human face from monocular video data, significantly improving the quality and flexibility of the reconstruction, especially in dynamic facial expression and detail reconstruction. However, the training and rendering processes of NeRF require millions of queries of deep multi-layer perceptrons, resulting in extremely slow training and rendering speeds, seriously affecting its feasibility in practical applications.
[0004] To improve the inference speed of the NeRF model, some studies have introduced the multi-resolution hash encoding technique. This method reduces the storage requirements and improves the inference efficiency by mapping the coordinates of 3D space points to a hash table of finite size. The multi-resolution hash encoding can capture spatial information at different resolution levels, which can improve the operation efficiency of the model. Especially in dynamic face reconstruction, it can significantly accelerate the inference speed and reduce the computational burden. However, although the multi-resolution hash encoding technique has improved the inference speed, it also brings some new challenges. Due to the limited size of the hash table, multiple 3D space points may be mapped to the same hash location, resulting in hash collisions. Hash collisions make it difficult for the model to distinguish different spatial features. Especially in the scenarios of dynamic faces and complex facial expression changes, detailed information may be lost, leading to a decrease in the reconstruction accuracy. In addition, many current methods still rely on large-scale 3D face deformation meshes to represent facial dynamic deformations, which not only increases the complexity of geometric calculations, significantly increases the computational burden of the model, and further leads to a decrease in the training and inference speeds. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a 3D dynamic face reconstruction method based on hybrid features and regional expressions, aiming to alleviate the hash collision and high model complexity problems in the prior art, and at the same time improve the stability of color prediction and the reconstruction accuracy.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A 3D dynamic face reconstruction method based on hybrid features and regional expressions, comprising the following steps:
[0008] S1: Collect monocular dynamic face video data, preprocess it and obtain the camera internal parameters, the face pose parameters, the face expression parameters, the depth image, and the face semantic segmentation image corresponding to each frame of the image;
[0009] S2: Obtain the multi-resolution hash features and the two-dimensional plane hash features of the 3D space points, and splice the two to construct hybrid features;
[0010] S3: Perform variance analysis on the face expression parameters, screen out the key expression dimensions, and fuse the multi-resolution hash features with the face expression parameters through a multi-layer perceptron to construct regional expression features;
[0011] S4: Input the hybrid features and the regional expression features into a density network to calculate the volume density and geometric features of the 3D space points;
[0012] S5: Introduce a view - independent color feature branch, combine it with the view - related color feature branch to generate the final color of the 3D space points, and through the volume rendering process, combine the calculated volume density and color information to generate the final 3D dynamic face image;
[0013] Further, the specific steps of step S1 are as follows:
[0014] S11: Use a monocular camera to capture a frontal face video data;
[0015] S12: Frame the monocular video data, and use a face detection model to locate the face and crop the image sequence;
[0016] S13: Use a face tracking model to obtain the camera internal parameters, the head pose parameters corresponding to the t - th frame image the facial expression parameters, and the depth image, where t ∈ [0, N - 1] and N is the total number of video frames; the camera internal parameters include the focal lengths f x and f x in the horizontal and vertical directions of the camera, and the principal point coordinates c x and c y ;
[0017] S14: Use a video matting model to remove the background of the t - th frame image, and denote the image after removing the background as
[0018] S15: Use a face parsing framework to perform semantic segmentation on to obtain the face semantic segmentation image, denoted as
[0019] S16: According to the semantic segmentation codes provided by remove the clothing in to obtain the face training image The two - dimensional coordinates of the pixel points of its image are defined as:
[0020] {(x, y)|x ∈ [0, H - 1], y ∈ [0, W - 1]}
[0021] where x and y correspond to the column index and row index of the image respectively; H and W represent the height and width of the image respectively;
[0022] Further, the specific steps of step S2 are as follows:
[0023] S21: According to the camera internal parameters, convert all pixel coordinates on the image plane to camera coordinates The conversion formula is as follows:
[0024]
[0025] S22: According to the image corresponding face pose parameters convert the camera coordinates to world coordinates The conversion formula is as follows:
[0026]
[0027] S23: For the image for a single pixel (x, y) in it, starting from the camera optical center, along the direction of the world coordinate point corresponding to the pixel point, form a sampling ray where k ∈ [0, m - 1], and m is the total number of sampling rays in the image ;
[0028] S24: For the sampling ray According to the bounding box range of the spatial scene, determine the sampling start point and the end point and uniformly sample within the sampling interval to obtain the position coordinates of a set of sampling points as well as the sampling direction where n t,k represents the number of all sampling points under the k-th sampling ray in the t-th frame;
[0029] S25: Perform multi-resolution hash encoding on the position coordinates of any sampling point under the k-th sampling ray of the image to obtain the multi-resolution hash feature of the sampling point;
[0030] Furthermore, the calculation steps of the multi-resolution hash encoding described in step S25 are as follows:
[0031] S251: For the resolution of the l-th layer where l ∈ [0, L - 1], the size of its voxel grid is calculated by the following formula:
[0032]
[0033] where L represents the number of layers of the hash table; N min represents the coarsest resolution; N max represents the finest resolution; b represents the scaling factor for controlling the rate of resolution change;
[0034] S252: For the sampled three-dimensional space point p i , where i ∈ [1, n t,k , normalize it and scale it to the range [0, N l of the l-th layer resolution grid:
[0035]
[0036] S253: Take points The eight nearest voxel corner points form a voxel corner point set:
[0037]
[0038] S254: Map each corner point to a specific index in the hash table through a hash function. The definition of the hash function is as follows:
[0039]
[0040] where represents the exclusive OR operation; 3 is the spatial point dimension; ε k is a unique large prime number used to enhance the hashing performance of the hash function; T represents the size of each layer of the hash table;
[0041] S255: Retrieve the eigenvalue corresponding to the voxel corner point v j in the hash table through the hash function h(v j ).
[0042] S256: Perform trilinear interpolation on to calculate its feature vector at the l-th layer
[0043]
[0044] where weights j is the weight calculated from the spatial distance between and the voxel corner point v j .
[0045] S257: The multi-resolution hash feature F i of the 3D spatial point p H,i is obtained by concatenating the features of all resolution layers:
[0046]
[0047] S26: Perform two-dimensional plane hash encoding on the position coordinates of all sampling points on the k-th sampling ray of the image to obtain the two-dimensional plane hash feature of the sampling points;
[0048] Furthermore, the calculation steps of the two-dimensional plane hash feature described in step S26 are as follows:
[0049] S261: For the 3D spatial point p i =(x i, y i , z i )(where \(i\in[1, n t,k \)), project it onto the two-dimensional planes \((X, Y)\), \((X, Z)\), \((Y, Z)\) to obtain the set \(P\) of projected two-dimensional coordinates 2D,i :
[0050] P 2D,i =\{(x i , y i ), (x i , z i ), (y i , z i )\}
[0051] S262: Independently construct a multi-resolution grid for each two-dimensional plane. For the grid resolution of the \(l\in[0, L 2D - 1]\) layer, the calculation formula is: The calculation formula is:
[0052]
[0053] where \(L 2D represents the number of layers of the two-dimensional plane hash code; represents the coarsest resolution; represents the finest resolution; \(b 2D is a scaling factor used to control the resolution change;
[0054] S263: For the projected coordinate \(p 2D,i =(a i , b i ) on any plane \((A, B)\), scale it according to the grid resolution of level \(l\). The scaled two-dimensional coordinate is:
[0055]
[0056] S264: Take the four nearest voxel corner points to form a voxel corner point set:
[0057]
[0058] S265: Each voxel corner point \(v 2D,j is mapped to the hash table through the following hash function:
[0059]
[0060] where represents the exclusive OR operation; 2 is the dimension of the plane points; \(\epsilon k is a unique large prime number; \(T2D Indicates the size of each layer of hash table for two-dimensional plane hash encoding;
[0061] S266: Through hash mapping, map each corner point v 2D,j to the index in the hash table and obtain the eigenvalue of each corner point from the hash table which represents the spatial feature of the corner point;
[0062] S267: For each two-dimensional coordinate point calculate its hash feature at the resolution of the l-th layer:
[0063]
[0064] where weights j is the weight calculated from the spatial distance between the voxel corner point v 2D,j ;
[0065] S268: The multi-resolution hash feature of point p 2D,i is obtained by concatenating the features of all resolution layers:
[0066]
[0067] where represents the hash feature of the two-dimensional coordinate (a, b) after projecting the three-dimensional space point p i onto the plane (A, B); F 2D represents the number of feature dimensions of each layer;
[0068] S269: Encode the two-dimensional coordinates after mapping the three-dimensional space point p i respectively to obtain three two-dimensional plane hash features, which are Connect them to construct the two-dimensional plane hash feature F 2Dp-H,i :
[0069]
[0070] S27: Connect the multi-resolution hash feature F H,i with the two-dimensional plane hash feature F 2Dp-H,i to construct the hybrid feature F Hy,i :
[0071]
[0072] Furthermore, the step S3 specifically includes the following steps:
[0073] S31: For the 3DMM human face expression parameter matrix E expPerform an analysis of variance, where N is the total number of video frames. Calculate the variance Var of each expression dimension d , and the formula is as follows:
[0074]
[0075] where is the mean of the d-th dimension, d ∈ [1, 100]; represents the value of the d-th dimension of matrix E exp at the t-th frame;
[0076] S32: Calculate the mean Var mean of the variances of all expression dimensions, and the calculation formula is as follows:
[0077]
[0078] S33: Select the expression dimensions with variances greater than the mean High FreqDims = {d|Var d > Var mean}, these dimensions usually contain high-frequency dynamic information of facial expressions, and further introduce some low-frequency dimensions with variances lower than the variance mean, and these dimensions supplement the low-frequency information in facial expressions. The present invention uses the filtered expression parameters E input as input features;
[0079] S34: Use a multi-layer perceptron to reduce the dimension of the expression parameters E input , and the specific formula is as follows:
[0080] f E = W3·ReLU(W2·ReLU(W1·E input + b1)+ b2)+ b3
[0081] where f E is the expression feature after dimension reduction; W3 represents the weight matrix of the third layer in the network; ReLU(·) is the activation function; W2 represents the weight matrix of the second layer; W1 represents the weight matrix of the first layer; b1, b2, b3 represent the bias vectors of each layer in the network;
[0082] S35: Reduce the dimension of the multi-resolution hash feature F H,i of the three-dimensional space points, and use a multi-layer perceptron for calculation. The formula is as follows:
[0083]
[0084] is the space feature after dimension reduction; W3 represents The weight matrix of the third layer in the network; W2 represents the weight matrix of the second layer; W1 represents the weight matrix of the first layer; b1, b2, b3 represent the bias vectors of each layer in the network;
[0085] S36: For the dimensionality-reduced expression feature f E and the spatial feature perform element-wise multiplication to generate the weighted expression feature F E,i , and the formula is as follows:
[0086]
[0087] where the symbol represents the element-wise multiplication operation. Through this operation, the multi-resolution features corresponding to each spatial point are associated with the expression feature. Each dimensional weight in E can provide an adjustment for the local characteristics of the spatial points of the expression feature f
[0088] Furthermore, the step S4 specifically includes the following steps:
[0089] S41: Concatenate the mixed feature F Hy,i and the expression feature F E,i as the input of the density network , and the calculation formula is as follows:
[0090]
[0091] S42: Extract the first component h i,0 in the network output and calculate the volume density σ i of the three-dimensional spatial point p i , and its formula is:
[0092] σ i = ReLU(h i,0 )
[0093] where the volume density represents the transparency of the spatial point and is used for volume rendering;
[0094] S43: Extract the remaining 15 components [h i,1 , h i,2 , …, h i,15 in the network output to form the geometric feature F geo,i , which represents the local geometric information of the spatial point:
[0095] F geo,i = [h i,1 , h i,2 , …, h i,15
[0096] Among them, F geo,i is used as the input of the subsequent color prediction network;
[0097] S44: Determine the local geometric structure and volume attributes of the three-dimensional space points through the calculated bulk density σ i and the geometric feature F geo,i , and provide input data for subsequent color prediction and volume rendering;
[0098] Furthermore, the step S5 specifically includes the following steps:
[0099] S51: Perform spherical harmonic encoding on the viewing direction d i corresponding to the three-dimensional space point p i =(d x , d y , d z ) to obtain the direction feature γ(d).
[0100] S52: Use the multi-layer perceptron of the view-dependent color branch to calculate the view-dependent color information . The input is the geometric feature F geo,i and the direction feature γ(d), and the calculation formula is as follows:
[0101] c vd,i = W3·ReLU(W2·ReLU(W1·[F geo,i , γ(d)] + b1) + b2) + b3
[0102] Among them, W3 represents the weight matrix of the output layer; W2 represents the weight matrix of the second layer; W1 is the weight matrix of the first layer; b1 and b2 respectively represent the bias vectors of the first and second layers; b3 represents the bias vector of the output layer;
[0103] S53: Use the multi-layer perceptron of the view-independent color branch to calculate the view-independent color information c vi,i . The input is the geometric feature F geo,i , and the calculation formula is as follows:
[0104] c vi,i = W2·ReLU(W1·F geo,i + b1) + b2
[0105] Among them, W2 is the weight matrix of the second layer; W1 is the weight matrix of the first layer; b1 represents the bias vector of the first layer; b2 represents the bias vector of the second layer;
[0106] S54: The final color c is obtained by combining view - related color information and view - independent color information and activating it using the sigmoid function. i , and the calculation formula is as follows:
[0107] c i = sigmoid(c vd,i + c vi,i )
[0108] S55: Using the volume density σ i and the color c i , calculate the final color value of each ray through the volume rendering formula:
[0109]
[0110] where T(t) represents the cumulative transmittance;
[0111] S56: Calculate the color values for the positions (u, v) of all image pixels to form the rendered image I render . During the model training process, use the Huber loss function to calculate the RGB loss between the original image I t and the rendered image I render :
[0112] L = Huber(I t , I render )
[0113] The beneficial effects of the present invention are as follows:
[0114] (1) The present invention proposes a three - dimensional face reconstruction method based on hybrid features and regional expressions. By projecting three - dimensional space points onto a two - dimensional plane for hash mapping, the hash feature distribution becomes sparser, thereby reducing hash collisions and improving training efficiency. At the same time, combining multi - resolution hash features with expression parameters to drive three - dimensional face deformation, replacing the traditional deformation method relying on geometric priors, reduces the model calculation complexity, ensuring both high - precision reconstruction and significantly accelerating the model training efficiency.
[0115] (2) The method of the present invention is optimized in the color prediction stage. Specifically for the color modeling problem of dynamic faces, view - independent branch features are introduced to weaken the interference of view changes and complex lighting on color prediction, and a hybrid feature splicing strategy is adopted to enhance the color expression ability. It not only effectively improves the detail fidelity of the reconstructed face but also enhances the modeling ability of the neural radiance field for complex expression changes.
[0116] Other advantages, objectives, and features of the present invention will be further clarified in the following description. To some extent, they will be obvious to those skilled in the art based on the study of the following content, or can be learned from the practice of the present invention. The objectives and technical effects of the present invention can be achieved and embodied through the specific implementation methods described in the following description. Description of the Drawings
[0117] To make the objectives, technical solutions, and advantages of the present invention clearer, the modules of the present invention will be described in detail below in conjunction with the drawings, where:
[0118] Figure 1 It is a flowchart of the three-dimensional dynamic face reconstruction method based on hybrid features and regional expressions according to the embodiment of the present invention;
[0119] Figure 2 It is a schematic diagram of the forward propagation of the model of the present invention;
[0120] Figure 3 It is a schematic diagram of the process of obtaining hybrid features in the model of the present invention;
[0121] Figure 4 It is a comparison diagram of the qualitative effects of the model of the present invention; Detailed Embodiments
[0122] The following uses specific embodiments to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this book. The present invention can also be implemented and applied through other specific implementation manners. The details clearly stated herein can also be modified or changed based on different fields and applications in their respective backgrounds. It should be noted that, without changing the basic concept illustrated by the diagrams and intentions provided in the embodiments, the following embodiments and the features in the embodiments can also be combined with each other.
[0123] Among them, the drawings are used for exemplary illustration, representing only the intention, rather than the actual physical form, and do not understand the limitations of the present invention, nor can they better reflect the embodiments of the present invention. The two parts at the base of the drawings are each omitted, enlarged or reduced, and represent the actual spatial dimensions; those skilled in the art will be able to understand after reading this part.
[0124] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0125] An embodiment of the present invention provides a three-dimensional dynamic face reconstruction method based on hybrid features and regional expressions. Its forward propagation process is as Figure 2 shown, and the method specifically includes the following steps:
[0126] S1: Collect monocular dynamic face video data, preprocess it, and obtain the camera internal parameters, the face pose parameters, face expression parameters, depth image, and face semantic segmentation image corresponding to each frame of the image;
[0127] Specifically, the detailed steps of the preprocessing are as follows:
[0128] S11: Use a nikon Z6Ⅱ camera to shoot a frontal face video data with a duration of 2 - 3 minutes. During the shooting process, the ambient light should be kept constant, and the object being photographed can make changes in expressions, mouth, eyes, etc., and the head can rotate;
[0129] S12: Use the ffmpeg software to perform frame splitting on the monocular video data to obtain an image sequence, use the face detection model ArcFace to locate the face, and then crop each frame of the image in the image sequence to a size of 512 * 512;
[0130] S13: Use the MICA face tracking model to obtain the camera internal parameters, the head pose parameters corresponding to the t-th frame image face expression parameters, and depth image from the image sequence, where t ∈ [0, N - 1], and N is the total number of video frames; the camera internal parameters include the focal lengths f x and f x , and the principal point coordinates c x and c y ;
[0131] S14: Use a video matting model to remove the background of the t-th frame image, and denote the image after removing the background as In this embodiment, the pre-trained RobustVideoMatting video matting model is used to remove the background;
[0132] S15: Use the pre-trained face parsing framework BiSeNetV2 to perform semantic segmentation to obtain a face semantic segmentation image, denoted as
[0133] S16: According to the provided semantic segmentation code, remove the clothing in to obtain a face training image The two-dimensional coordinates of the pixel points of its image are defined as:
[0134] {(x, y)|x ∈ [0, H - 1], y ∈ [0, W - 1]}
[0135] where x and y correspond to the column index and row index of the image respectively; H and W represent the height and width of the image respectively;
[0136] S2: Obtain the multi-resolution hash feature and two-dimensional plane hash feature of the three-dimensional space points, and splice the two to construct a hybrid feature;
[0137] Specifically, the process of obtaining the hybrid feature is as Figure 3 shown, and the calculation process specifically includes the following steps:
[0138] S21: According to the camera internal parameters, convert all pixel coordinates on the image plane to camera coordinates The conversion formula is as follows:
[0139]
[0140] S22: According to the corresponding face pose parameters of the image convert the camera coordinates to world coordinates The conversion formula is as follows:
[0141]
[0142] S23: For a single pixel (x, y) in the image , starting from the camera optical center, along the direction of the world coordinate point corresponding to the pixel point, form a sampling ray where k ∈ [0, m - 1], and m is the total number of sampling rays in the image ;
[0143] S24: For the sampling ray , determine the sampling start point and end point And uniformly sample within the sampling interval to obtain the position coordinates of a set of sampling points and the sampling direction where n t,k represents the number of all sampling points on the k-th sampling ray of the t-th frame image;
[0144] S25: For the image Perform multi-resolution hashing encoding on the position coordinates of any sampling point on the k-th sampling ray to obtain the multi-resolution hash feature of the sampling point.
[0145] The calculation steps of the multi-resolution hashing encoding described in step S25 are as follows:
[0146] S251: For the resolution of the l-th layer where l ∈ [0, L - 1], the size of its voxel grid is calculated by the following formula:
[0147]
[0148] where L represents the number of layers of the hash table, and the value in this embodiment is 16; N min represents the coarsest resolution, with a value of 16; N max represents the finest resolution, with a value of 2048; b represents the scaling factor, which is used to control the rate of resolution change.
[0149] S252: For the sampled three-dimensional space point p i , where i ∈ [1, n t,k , normalize it and scale it to the range [0, N l of the l-th layer resolution grid:
[0150]
[0151] S253: Take the eight voxel corner points closest to the point to form a voxel corner point set:
[0152]
[0153] S254: Map each corner point to a specific index in the hash table through a hash function. The definition of the hash function is as follows:
[0154]
[0155] where represents the exclusive OR operation; 3 is the dimension of the space point; ε k is a unique large prime number used to enhance the hashing performance of the hash function; T represents the size of each layer of the hash table, with a value of 2 18 ;
[0156] S255: Retrieve the eigenvalue corresponding to the voxel corner point v in the hash table by means of the hash function h(v j ). j
[0157] S256: Perform trilinear interpolation on to calculate its feature vector at the l-th layer
[0158]
[0159] where weights j is the weight calculated from the spatial distance between j and the voxel corner point v;
[0160] S257: The multi-resolution hash feature F i of the 3D spatial point p H,i is obtained by concatenating the features of all resolution layers:
[0161]
[0162] where F represents the number of feature dimensions per layer, which takes the value of 8 in the embodiments of the present invention;
[0163] S26: Perform two-dimensional plane hash coding on the position coordinates of all sampling points on the k-th sampling ray of the image to obtain the two-dimensional plane hash feature of the sampling points;
[0164] Furthermore, the calculation steps of the two-dimensional plane hash coding described in step S26 are as follows:
[0165] S261: For the 3D spatial point p i =(x i , y i , z i )(where i ∈ [1, n t,k ), project it onto the two-dimensional planes (X, Y), (X, Z), (Y, Z) to obtain the set of projected two-dimensional coordinates P 2D,i :
[0166] P 2D,i ={(x i , y i ), (x i , z i ), (y i , z i )}
[0167] S262: Independently construct a multi-resolution grid for each two-dimensional plane. For the l ∈ [0, L2D Grid resolution of layer [-1] The calculation formula is as follows:
[0168]
[0169] Where L 2D represents the number of layers of the two-dimensional plane hash code, and the value in this embodiment is 4; represents the coarsest resolution, with a value of 16; represents the finest resolution, with a value of 2048; b 2D is a scaling factor used to control the resolution change;
[0170] S263: For the projection coordinate p 2D,i =(a i , b i ) on any plane (A, B), scale it according to the grid resolution of level l , and the scaled two-dimensional coordinate is:
[0171]
[0172] S264: Take the four nearest voxel corner points to form a voxel corner point set:
[0173]
[0174] S265: Each voxel corner point v 2D,j is mapped to the hash table through the following hash function:
[0175]
[0176] Where represents the exclusive OR operation; 2 is the dimension of the plane point; ε k is a unique large prime number; T 2D represents the size of each layer of the hash table of the two-dimensional plane, with a value of 2 14 ;
[0177] S266: Through hash mapping, map each corner point v 2D,j to the index in the hash table, and obtain the eigenvalue of each corner point from the hash table represents the spatial feature of this corner point;
[0178] S267: Calculate the hash feature of each two-dimensional coordinate point at the resolution of layer l:
[0179]
[0180] Among them, weights j yes With voxel corner v 2D,j The weight is calculated by the spatial distance between them;
[0181] S268: Two-dimensional coordinate set p 2D,i Multi-resolution hashing features The feature concatenation of all resolution layers yields:
[0182]
[0183] in, Represents the hash feature of the two-dimensional coordinates (a, b) on the plane (A, B). 2D Indicates the number of feature dimensions per layer, the value is 2;
[0184] S269: point p i After encoding each plane coordinate, three two-dimensional plane hash features are obtained, which are
[0185] S2610: point p i The three plane features of are connected to form the two-dimensional plane hash feature of the spatial point Its expression is as follows:
[0186]
[0187] S27: Multi-resolution hash feature F H,i With the two-dimensional plane hash feature F 2dP-H,i Connect to form a mixed feature F Hy,i :
[0188] F Hy,i =concat(F H,i , F 2Dp-H,i )
[0189] S3: Perform variance analysis on facial expression parameters to screen out key expression dimensions, and fuse multi-resolution hash features with facial expression parameters through a multi-layer perceptron to form regional expression features;
[0190] S31: 3DMM facial expression parameter matrix Perform variance analysis, where N is the total number of video frames. Calculate the variance Var of each expression dimension d , the formula is as follows:
[0191]
[0192] In the formula, is the mean of the d-th dimension, where d ∈ [1, 100]; represents the matrix E exp the value of the d-th dimension at the T-th frame;
[0193] S32: Calculate the mean variance Var of all expression dimensions mean , and the calculation formula is as follows:
[0194]
[0195] S33: Screen out the expression dimensions with variances greater than the mean High FreqDims = {d|Var d > Var mean}, these dimensions usually contain high-frequency dynamic information of facial expressions, and further introduce some low-frequency dimensions with variances lower than the variance mean, and these dimensions supplement the low-frequency information in facial expressions. The present invention uses the screened expression parameter E input as the input feature. In this embodiment, the first 50 dimensions of the facial expression parameters are selected as the input feature E input .
[0196] S34: Use a multi-layer perceptron to reduce the dimension of the expression parameter E input , and the specific formula is as follows:
[0197] f E = W3·ReLU(W2·ReLU(w1·E input + B1)+ B2)+ B3
[0198] where, is the expression feature after dimension reduction; represents the weight matrix of the third layer in the network; ReLU(·) is the activation function; represents the weight matrix of the second layer; represents the weight matrix of the first layer; b1, b2, represents the bias vector of each layer in the network.
[0199] S35: Reduce the dimension of the multi-resolution hash feature F of the three-dimensional space points H,i , and use a multi-layer perceptron to calculate, and the formula is as follows:
[0200]
[0201] where, is the spatial feature after dimension reduction; represents the weight matrix of the third layer in the network; The weight matrix representing the second layer; The weight matrix representing the first layer; b1, b2, Represents The bias vector for each layer in the network.
[0202] S36: For the dimension-reduced expression feature f E and the spatial feature Multiply element-wise to generate the weighted expression feature F E,i , and the formula is as follows:
[0203]
[0204] Among them, the symbol Represents the element-wise multiplication operation. Through this operation, the multi-resolution features corresponding to each spatial point are associated with the expression feature. Each dimension weight in E Can provide an adjustment function for the local characteristics of the spatial points of the expression feature, enabling the expression feature to adapt to the dynamic changes in different spatial regions.
[0205] S4: Input the mixed feature and the regional expression feature into the density network to calculate the volume density and geometric features of the three-dimensional spatial points;
[0206] Specifically, the step S4 specifically includes the following steps:
[0207] S41: Concatenate the mixed feature F Hy,i and the expression feature F E,i as the input of the density network , and the calculation formula is as follows:
[0208]
[0209] S42: Extract the first component h i,0 in the network output and calculate the volume density σ i of the three-dimensional spatial point p i , and its formula is:
[0210] σ i = ReLU(h i,0 )
[0211] Among them, the volume density Represents the transparency of the spatial point and is used for volume rendering.
[0212] S43: Extract the remaining 15 components [h i,1 , h i,2 , …, h i,15 in the network output to form the geometric feature F geo,i , representing the local geometric information of the spatial point:
[0213] F geo,i = [h i,1 ,h i,2 ,…,h i,15
[0214] Among them, F geo,i is used as the input for the subsequent color prediction network.
[0215] S44: Determine the local geometric structure and volume properties of the 3D space points based on the calculated bulk density σ i and the geometric feature F geo,i , providing input data for subsequent color prediction and volume rendering.
[0216] S5: Introduce a view - independent color feature branch, combine it with the view - dependent color feature branch to generate the final color of the 3D space points, and generate the final 3D dynamic face image through the volume rendering process by combining the calculated bulk density and color information;
[0217] S51: Perform spherical harmonic encoding on the view direction i corresponding to the 3D space point p to obtain the direction feature
[0218] S52: Use the multi - layer perceptron of the view - dependent color branch to calculate the view - dependent color information with the input being the geometric feature F geo,i and the direction feature γ(d):
[0219] c vd,i = W3·ReLU(W2·ReLU(W1·[F geo,i , γ(d)] + b1) + b2) + b3
[0220] Among them, represents the weight matrix of the output layer; represents the weight matrix of the second layer; is the weight matrix of the first layer; b1, respectively represent the bias vectors of the first and second layers; represents the bias vector of the output layer;
[0221] S53: Use the multi - layer perceptron of the view - independent color branch to calculate the view - independent color information with the input being the geometric feature F geo,i , and the calculation formula is as follows:
[0222] c vi,i = W2·ReLU(W1·Fgeo,i +b1)+b2
[0223] Among them, is the weight matrix of the second layer; is the weight matrix of the first layer; represents the bias vector of the first layer; represents the bias vector of the second layer;
[0224] S54: By combining the view - related color information and the view - independent color information and activating with the sigmoid function, the final color c is obtained i , and the calculation formula is as follows:
[0225] c i = sigmoid(c vd,i +c vi,i )
[0226] S55: Using the volume density σ i and the color c i , calculate the final color value of each ray through the volume rendering formula:
[0227]
[0228] Among them, T(t) represents the cumulative transmittance.
[0229] S56: Calculate the color value for the positions (u, v) of all image pixels to form the rendered image I render . During the model training process, use the Huber loss function to calculate the RGB loss between the original image I t and the rendered image I render :
[0230] L = Huber(I t , I render )
[0231] To better illustrate the performance of the present invention, the following specific experimental results are given for reference:
[0232] Dataset: The open - source INSTA dataset is adopted, which contains video data of ten subjects. The dataset contains both the frontal videos of speakers in a calm state and some videos with exaggerated expressions and large - scale head rotations. The original video duration in the dataset is 1 - 3 minutes.
[0233] Evaluation Metrics: To accurately quantify the quality of the images rendered by the model, the following key pixel-level evaluation metrics were adopted in this experiment: Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS). These metrics together constitute a comprehensive and objective quantitative evaluation system for the quality of the images rendered by the model.
[0234] During the optimization process, the Adam optimizer with weighted exponential moving average (weight set to 0.95) was used, and its initial learning rate was set to 0.0025. The learning rate scheduler used ExponentialLR, with a base of 0.1 and an exponent equal to the ratio of the current iteration number to the total iteration number. 2048 rays were sampled per batch, and 1024 spatial points were sampled on each ray. In addition, the last 350 frames of each subject's video were used as the test set, and 1500 frames were randomly selected as the training set in each epoch.
[0235] Experimental Setup: The method of the present invention was compared with the state-of-the-art method (INSTA) in the reconstruction speed direction of current 3D dynamic face reconstruction. The average quantitative experimental results of the dataset are shown in Table 1:
[0236]
[0237] It can be seen that the method of the present invention performs excellently in key metrics such as PSNR, LPIPS, and SSIM on the dataset used by INSTA, which directly reflects the advantages of the model of the present invention in reconstruction quality. More importantly, this method has achieved a significant reduction in training time.
[0238] Qualitative Comparison: Figure 4 This is a comparison diagram of the qualitative effects of the model of the present invention. It can be observed that INSTA has problems of distortion when reconstructing the eye and mouth regions, which affects the overall visual realism. In contrast, the method proposed by the present invention shows higher authenticity when reconstructing the human face, especially when dealing with the eyes, the situation of wearing glasses, and the subtle features of the mouth, demonstrating more excellent detail capture and restoration capabilities.
[0239] In summary, the present invention proposes a dynamic three-dimensional face reconstruction method based on hybrid features and regional expressions to solve the problems of hash collision and high model complexity faced by traditional neural radiance fields in dynamic face reconstruction. The present invention combines two-dimensional planar hash features and multi-resolution hash features to alleviate the problem of hash conflicts in multi-resolution hash features. By associating multi-resolution hash features with facial expression parameters to drive three-dimensional face deformation, the deformation modeling based on geometric priors is replaced, reducing the model calculation complexity. In addition, the method of the present invention adds and fuses view-independent branch features in the face reconstruction color prediction network to reduce the influence of view changes and complex lighting on the inherent color performance of the face and improve the stability of color prediction. Experimental results show that compared with the current SOTA methods, the method of the present invention has improvements in terms of high fidelity, detail fidelity, and efficiency.
[0240] Those of ordinary skill in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, the steps of the present method can be realized. The storage medium, such as: ROM / RAM, magnetic disk, optical disc, etc.
[0241] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A three-dimensional dynamic face reconstruction method based on hybrid features and regional expressions, characterized in that It includes the following steps: S1: Collect monocular dynamic face video data, preprocess it and obtain the camera internal parameters, face pose parameters, face expression parameters, depth images, and face semantic segmentation images corresponding to each frame of the image; S2: Obtain the multi-resolution hash features and two-dimensional plane hash features of the three-dimensional space points, and splice the two to construct a hybrid feature; S3: Conduct an analysis of variance on the face expression parameters, screen out the key expression dimensions, and fuse the multi-resolution hash features with the face expression parameters through a multi-layer perceptron to construct regional expression features; S4: Input the hybrid feature and the regional expression feature into a density network to calculate the volume density and geometric features of the three-dimensional space points; S5: Introduce a view-independent color feature branch, combine it with a view-dependent color feature branch to generate the final color of the three-dimensional space points, and generate the final three-dimensional dynamic face image through a volume rendering process in combination with the calculated volume density and color information; 2. The three-dimensional dynamic face reconstruction method based on hybrid features and regional expressions according to claim 1 of the specification, wherein The specific steps of step S1 include the following steps: S11: Use a monocular camera to shoot a frontal face video data; S12: Perform frame processing on the monocular video data, and use a face detection model to locate the face and crop the image sequence; S13: Obtain the camera internal parameters, the head pose parameters corresponding to the t-th frame image, the facial expression parameters, and the depth image from the image sequence using the face tracking model, where t ∈ [0, N-1] and N is the total number of video frames; the camera internal parameters include the focal lengths f x and f x in the horizontal and vertical directions of the camera, and the principal point coordinates c x and c y ; S14: Use the video matting model to remove the background of the t-th frame image, and denote the image after removing the background as S15: Use a face parsing framework to perform semantic segmentation to obtain a face semantic segmentation image, denoted as S16: According to the provided semantic segmentation code, remove the clothing in to obtain a face training image, and the two-dimensional coordinates of the pixel points of its image are defined as: {(x,y)|x∈[0,H-1],y∈[0,W-1]} where x and y respectively correspond to the column index and row index of the image; H and W respectively represent the height and width of the image; 3. The three-dimensional dynamic face reconstruction method based on hybrid features and regional expressions according to claim 1 of the specification, characterized in that, The specific steps of step S2 include the following steps: S21: According to the camera internal parameters, convert all pixel coordinates on the image plane into camera coordinates The conversion formula is as follows: S22: According to the image corresponding face pose parameters Convert the camera coordinates to world coordinates The conversion formula is as follows: S23: For the image For a single pixel (x, y) in the image, starting from the camera optical center, a sampling ray is formed along the direction of the world coordinate point corresponding to the pixel point where k ∈ [0, m - 1], and m is the total number of sampling rays in the image ; S24: For the sampled light rays Determine the sampling starting point of the ray according to the bounding box range of the spatial scene and the ending point and uniformly sample within the sampling interval to obtain the position coordinates of a set of sampling points as well as the sampling direction where n t,k represents the number of all sampling points under the k-th sampled light ray in the t-th frame; S25: For the image perform multi-resolution hashing encoding on the position coordinates of any sampling point under the k-th sampling ray to obtain the multi-resolution hash feature F of the sampling point H,i ; S26: For the image perform two-dimensional plane hashing encoding on the position coordinates of all sampling points under the k-th sampling ray to obtain the two-dimensional plane hashing feature F of the sampling points 2Dp-H,i ; S27: Connect the multi-resolution hash feature F of the sampling point H,i with the two-dimensional plane hash feature F 2Dp-H,i to construct the hybrid feature F Hy,i : F Hy,i = concat(F H,i , F 2Dp-H,i ) where concat() represents the feature splicing operation; 4. The three-dimensional dynamic face reconstruction method based on hybrid features and regional expressions according to claim 1 of the specification, characterized in that The calculation process of the multi-resolution hash features of the three-dimensional space points in step S25 is as follows: S251: For the resolution of the l-th layer where l∈[0,L-1], the size of its voxel grid is calculated by the following formula: where L represents the number of layers of the hash table; N min represents the coarsest resolution; N max represents the finest resolution; b represents the scaling factor used to control the rate of resolution change; S252: For the sampled three-dimensional spatial point p i , where i ∈ [1, n t,k , normalize it and scale it to the range [0, N l of the resolution grid of the l-th layer: S253: Point selection The eight nearest voxel corner points form a voxel corner point set: S254: Map each corner point to a specific index in the hash table through a hash function, and the definition of the hash function is as follows: Among them represents the exclusive OR operation; 3 is the dimension of the spatial point; ε k is the only large prime number used to enhance the hashing performance of the hash function; T represents the size of each layer of hash table; S255: Retrieve the eigenvalue corresponding to the voxel corner point v j in the hash table through the hash function h(v j ) S256: For perform trilinear interpolation to calculate its feature vector at the l-th layer Among them, weights j is a weight calculated based on the spatial distance between the voxel corner point v j ; S257: 3D spatial point p i The multi-resolution hash feature F H,i is obtained by concatenating the features of all resolution levels:
5. The three-dimensional dynamic face reconstruction method based on hybrid features and regional expressions according to claim 1 of the specification, characterized in that, The calculation process of the two-dimensional plane hash features of the three-dimensional space points in step S26 is as follows: S261: For the three-dimensional space point p i =(x i , y i , z i )(where i ∈ [1, n t,k ), project it onto the two-dimensional planes (X, Y), (X, Z), and (Y, Z) to obtain the projected two-dimensional coordinate set P 2D,i : P 2D,i = {(x i , y i ), (x i , z i ), (y i , z i )} S262: Independently construct a multi-resolution grid for each two-dimensional plane. For the grid resolution of the l-th layer where l ∈ [0, L 2D - 1], the calculation formula is as follows: Calculation formula: Among them, L 2D represents the number of layers of the two-dimensional plane hash code; represents the coarsest resolution; represents the finest resolution; b 2D is a scaling factor used to control the resolution change; S263: The projected coordinates p on any plane (A, B) 2D,i = (a i , b i ), and scale it according to the grid resolution of level l . The scaled two-dimensional coordinates are: S264: Take the four nearest voxel corner points to form a voxel corner point set: S265: Each voxel corner point v 2D,j is mapped to a hash table through the following hash function: Among them, represents an exclusive OR operation; 2 is the dimension of the planar points; ε k is the only large prime number; T 2D represents the size of each hash table in the two-dimensional planar hash code; S266: Map each corner point v through a hash map 2D,j to an index in the hash table and obtain the eigenvalue of each corner point from the hash table representing the spatial feature of the corner point; S267: For each two-dimensional coordinate point calculate its hash feature at the resolution of the l-th layer: Among them, weights j is a weight calculated from the spatial distance between the voxel corner point v 2D,j ; S268: Point p 2D,i 's multi-resolution hash feature is obtained by concatenating features across all resolution levels: Among them, represents the three-dimensional space point p i the hash feature of the two-dimensional coordinates (a, b) after projecting to the plane (A, B); F 2D represents the number of feature dimensions per layer; S269: Encode the two-dimensional coordinates after mapping the three-dimensional space point p i to obtain three two-dimensional plane hash features, which are respectively Connect them to construct the two-dimensional plane hash feature F 2Dp-H,i :
6. The three-dimensional dynamic face reconstruction method based on hybrid features and regional expressions according to claim 1 of the specification, characterized in that The specific steps of step S3 include the following steps: S31: Perform an analysis of variance on the 3DMM facial expression parameter matrix E exp where N is the total number of video frames. Calculate the variance Var d of each expression dimension, and the formula is as follows: In the formula, is the mean value of the d-th dimension, where d ∈ [1, 100]; represents the matrix E exp the value of the d-th dimension at the t-th frame; S32: Calculate the mean value Var of the variances of all expression dimensions mean , and the calculation formula is as follows: S33: Screen out the expression dimensions with variance greater than the mean. High FreqDims = {d|Var d >Var mean}, these dimensions usually contain high-frequency dynamic information of facial expressions, and further introduce some low-frequency dimensions with variance lower than the variance mean. These dimensions supplement the low-frequency information in facial expressions. The present invention uses the screened expression parameters E input as input features; S34: Use a multi-layer perceptron to reduce the dimensionality of the expression parameter E input The specific formula is as follows: f E = W3·ReLU(W2·ReLU(w1·E input + B1)+ B2)+ B3 Among them, f E is the facial expression feature after dimensionality reduction; W3 represents the weight matrix of the third layer in the network; ReLU(·) is the activation function; W2 represents the weight matrix of the second layer; W1 represents the weight matrix of the first layer; b1, b2, b3 represent the bias vectors of each layer in the network; S35: Dimensionality reduction is performed on the multi-resolution hash feature F of the three-dimensional space points, and a multi-layer perceptron is used for calculation. The formula is as follows: H,i The calculation is carried out using a multi-layer perceptron, and the formula is as follows: is the spatial feature after dimensionality reduction; W3 represents the weight matrix of the third layer in the network; W2 represents the weight matrix of the second layer; W1 represents the weight matrix of the first layer; b1, b2, b3 represent the bias vectors of each layer in the network; S36: Element-wise multiply the dimensionality-reduced facial expression feature f E and the spatial feature to generate the weighted facial expression feature F E,i , as shown in the following formula: Among them, the symbol represents an element-wise multiplication operation. Through this operation, the multi-resolution features corresponding to each spatial point are associated with the expression features. Each dimensional weight in E can provide an adjustment function for the local characteristics of the spatial points, enabling the expression features to adapt to the dynamic changes in different spatial regions.
7. The three-dimensional dynamic face reconstruction method based on hybrid features and regional expressions according to claim 1 of the specification, wherein The specific steps of step S4 include the following steps: S41: Concatenate the mixed feature F Hy,i and the expression feature F E,i as the input of the density network The calculation formula is as follows: S42: Extract the first component h from the network output i,0 Calculate the three-dimensional spatial point p i of the bulk density σ i , and its formula is: σ i = ReLU(h i,0 ) wherein, the bulk density σ i represents the transparency of the spatial point and is used for volume rendering. S43: Extract the remaining 15 components [h i,1 , h i,2 , …, h i,15 to form the geometric feature F geo,i , representing the local geometric information of the spatial points: F geo,i = [h i,1 , h i,2 , …, h i,15 Among them, F geo,i is used as the input of the subsequent color prediction network. S44: The bulk density σ obtained by calculation i and the geometric feature F geo,i are used to determine the local geometric structure and volume properties of the three-dimensional space points, providing input data for subsequent color prediction and volume rendering; 8. The three-dimensional dynamic face reconstruction method based on hybrid features and regional expressions according to claim 1 of the specification, characterized in that, The specific steps of step S5 include the following steps: S51: Perform spherical harmonic encoding on the viewing direction corresponding to the three-dimensional spatial point p i to obtain a direction feature S52: Multilayer Perceptron Using Viewpoint-Related Color Branches Calculate viewpoint-related color information The input is geometric feature F geo,i and direction feature γ(d): c vd,i = W3·ReLU(W2·ReLU(W1·[F geo,i ,γ(d)] + B1) + b2) + b3 Among them, represents the weight matrix of the output layer; represents the weight matrix of the second layer; is the weight matrix of the first layer; respectively represent the bias vectors of the first and second layers; represents the bias vector of the output layer; S53: Multilayer perceptron using view-invariant color branches Calculate view-invariant color information The input is geometric feature F geo,i , and the calculation formula is as follows: c vi,i = W2·ReLU(W1·F geo,i + b1)+ b2 Among them, is the weight matrix of the second layer; is the weight matrix of the first layer; represents the bias vector of the first layer; represents the bias vector of the second layer; S54: The final color c is obtained by combining the view-dependent color information and the view-independent color information and activating them using the sigmoid function. i , and the calculation formula is as follows: c i = sigmoid(c vd,i + c vi,i ) S55: Using the bulk density σ i and the color c i , calculate the final color value of each ray through the volume rendering formula: where T(t) represents the cumulative transmittance; S56: Calculate the color values for the positions (u, v) of all image pixels to form the rendered image I render . During the model training process, use the Huber loss function to calculate the original image I t and the rendered image I render for the RGB loss between them: L = Huvber(I t , I render ).
Citation Information
Cited By
Grid ray intersection method based on multi-dimensional Hash acceleration
CN121053306A