Multi-view mark-point-free facial expression capturing method and system
Through the multi-view without marking points facial expression capture method, multi-dimensional camera matrix is used to collect multi-view images of faces for detection and reconstruction, optimize the model coefficients, and generate more accurate and realistic facial expression models of human faces, solving the problem of insufficient accurate and efficient facial expression capture in the existing technology, and achieving accurate recovery of facial details.
Patent Information
- Application Number
- CN202510378381.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
AI Technical Summary
Existing facial expression capture technology cannot accurately restore facial details, such as blinking, detail wrinkles, etc., and it has low accuracy and efficiency.
A multi-view without marker points facial expression capture method is adopted, and a multi-view image of the face is collected through a multi-dimensional camera matrix, two-dimensional key point detection and neutral expression face three-dimensional reconstruction, non-neutral expression depth map is obtained, and the standard face topology model and model coefficient are optimized to generate a real face facial expression model.
It improves image accuracy and completeness, ensures the accuracy of expressions and postures, can reflect facial details, is more similar to real faces, and has high accuracy and high efficiency.
Smart Images

Figure CN120260100A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of facial expression capture, and particularly to a multi-view markerless facial expression capture method and system. Background Art
[0002] With the rise of new services such as virtual idols, virtual anchors, and digital humans, motion capture technology has been more widely applied. Among them, facial expression capture technology has become a key component of motion capture. Through facial expression capture technology, virtual humans can display lifelike expressions comparable to those of real humans. This rich expressiveness enables virtual humans to interact with humans more vividly and realistically. However, existing facial expression capture technologies cannot restore important facial details, such as blinking and detailed wrinkles, and there are also problems of poor accuracy and low efficiency in existing technologies.
[0003] Therefore, there is an urgent need for an efficient and accurate facial expression capture method. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a multi-view markerless facial expression capture method and system, which solves the problem that facial expression capture in the prior art is not accurate and efficient enough.
[0005] To solve the above technical problems, the present invention provides a multi-view markerless facial expression capture method, including:
[0006] Collecting multi-view images of a human face based on a multi-dimensional camera matrix;
[0007] Performing two-dimensional key point detection and three-dimensional reconstruction of a neutral expression human face on the multi-view images of the human face to obtain a standard human face topology model;
[0008] Obtaining a non-neutral expression depth map from the multi-view images of the human face;
[0009] Optimizing model coefficients based on the standard human face topology model, the non-neutral expression depth map, and the multi-view images of the human face; the model coefficients include expression basis coefficients, translation coefficients, and Euler angle rotation coefficients;
[0010] Inputting the optimized model coefficients into the standard human face topology model to obtain a real human face facial expression model.
[0011] Optionally, performing two-dimensional key point detection and three-dimensional reconstruction of a neutral expression human face on the multi-view images of the human face to obtain a standard human face topology model, including:
[0012] Based on each part, using a neural network to perform key point and contour detection on the multi-view images of the human face to obtain target key points and a target contour;
[0013] Based on the neutral facial expression images in the multi-view facial images, a 3D reconstructed facial skin model, reconstructed texture, and internal and external camera parameters are obtained using 3D reconstruction technology;
[0014] Based on the target key points, target contour, and the 3D reconstructed facial skin model, reconstructed texture, and internal and external camera parameters, the standard facial topology model is obtained.
[0015] Optionally, obtaining the standard facial topology model based on the target key points, target contour, and the 3D reconstructed facial skin model, reconstructed texture, and internal and external camera parameters includes:
[0016] Obtain a first standard facial skin topology model, and label the target key points and target contour on the first standard facial skin topology model to obtain a second standard facial skin topology model;
[0017] Based on the second standard facial skin topology model and the 3D reconstructed facial skin model, project the key points using the internal and external camera parameters, solve for the triangle positions using the nricp_sumner algorithm, and solve for the vertex transformations using the nricp_amberg algorithm to obtain a third standard facial skin topology model;
[0018] Based on the first standard texture, use the texture transfer algorithm to transfer the reconstructed texture to the first standard texture to obtain a second standard texture;
[0019] Based on the third standard facial skin topology model and the second standard texture, the standard facial topology model is obtained.
[0020] Optionally, obtaining the standard facial topology model based on the third standard facial skin topology model and the second standard texture includes:
[0021] Obtain an initial standard facial topology model according to the third standard facial skin topology model and the second standard texture;
[0022] Render the initial standard facial topology model using spherical harmonic lighting coefficients to obtain a color map;
[0023] Calculate the loss based on the color map and the neutral facial expression images in the multi-view facial images, backpropagate the loss, and optimize the spherical harmonic lighting coefficients using the adam algorithm;
[0024] Render the initial standard facial topology model using the optimized spherical harmonic lighting coefficients to obtain the standard facial topology model.
[0025] Optionally, optimize the model coefficients based on a standard face topology model, a non-neutral expression depth map, and the multi-view face images, including:
[0026] Obtain a distance loss value according to the distance from the vertices in the standard face topology model to the non-neutral expression depth map, calculate a contour loss value according to the contour in the standard face topology model and the target contour, calculate a key-point loss value according to the key points in the standard face topology model and the target key points, and calculate a total loss value;
[0027] According to the total loss value, use the conjugate gradient algorithm to solve for the optimization direction of the least squares method, and use the line search of the Armijo condition to solve for the optimization step size of the least squares method;
[0028] Based on the optimization direction and the optimization step size, use the least squares method to optimize the expression basis coefficients, translation coefficients, and Euler angle rotation coefficients.
[0029] Optionally, obtain a distance loss value according to the distance from the vertices in the standard face topology model to the non-neutral expression depth map, and calculate a contour loss value according to the contour in the standard face topology model and the target contour, including:
[0030] Project the vertices in the standard face topology model onto the non-neutral expression depth map through the internal and external camera parameters, and calculate the distance and normal angle distance from the vertices to the non-neutral expression depth map;
[0031] Set the initial weight of the distance according to the normal angle distance;
[0032] Set the mask weight according to the distance from the vertices to the non-neutral expression depth map and the distance threshold;
[0033] Obtain the final weight according to the initial weight and the mask weight;
[0034] Obtain the distance loss value according to the final weight and the distance from the vertices to the non-neutral expression depth map;
[0035] Project the points on the contour line of the standard face topology model onto the multi-view face images through the internal and external camera parameters to obtain two-dimensional contour points;
[0036] Determine the target vertical point and the normal of the target vertical point based on the two-dimensional contour points;
[0037] Calculate the contour loss value based on the normal, the target vertical point, and the two-dimensional contour points.
[0038] Optionally, inputting the optimized model coefficients into the standard face topology model to obtain a real face facial expression model, including:
[0039] Inputting the optimized model coefficients into the standard face topology model to obtain an initial face facial expression model;
[0040] Using the distance loss value as supervision information to optimize all vertex offsets in the initial face facial expression model to obtain the real face facial expression model.
[0041] The present invention also provides a multi-view markerless facial expression capture system, including:
[0042] Collecting multi-view face images based on a multi-dimensional camera matrix;
[0043] Performing two-dimensional key point detection and neutral expression face three-dimensional reconstruction on the multi-view face images to obtain a standard face topology model;
[0044] Obtaining a non-neutral expression depth map from the multi-view face images;
[0045] Optimizing model coefficients based on the standard face topology model, the non-neutral expression depth map, and the multi-view face images; the model coefficients include expression basis coefficients, translation coefficients, and Euler angle rotation coefficients;
[0046] Inputting the optimized model coefficients into the standard face topology model to obtain a real face facial expression model.
[0047] The present invention also provides a multi-view markerless facial expression capture device, including:
[0048] A memory for storing a computer program;
[0049] A processor for implementing the above multi-view markerless facial expression capture method when executing the computer program.
[0050] The present invention also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are loaded and executed by a processor, the above multi-view markerless facial expression capture method is implemented.
[0051] It can be seen that the present invention collects multi-view face images based on a multi-dimensional camera matrix; performs two-dimensional key point detection and three-dimensional reconstruction of a neutral-expression face on the multi-view face images to obtain a standard face topology model; obtains a non-neutral-expression depth map from the multi-view face images; optimizes model coefficients based on the standard face topology model, the non-neutral-expression depth map, and the multi-view face images; the model coefficients include expression basis coefficients, translation coefficients, and Euler angle rotation coefficients; inputs the optimized model coefficients into the standard face topology model to obtain a real face facial expression model. The multi-view face images collected by the present invention using the multi-dimensional camera matrix can improve the accuracy and integrity of the obtained images; optimize the expression basis coefficients, translation coefficients, and Euler angle rotation coefficients, so as to obtain a more accurate and realistic face facial expression model based on the standard face topology model, ensure the accuracy of expressions and poses, and be able to reflect facial detail features and be more similar to a real face.
[0052] In addition, the present invention also provides a multi-view markerless facial expression capture system, device, and computer-readable storage medium, which also have the above beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0054] Figure 1 It is a flowchart of a multi-view markerless facial expression capture method provided by an embodiment of the present invention;
[0055] Figure 2 It is an example diagram of a multi-dimensional camera matrix provided by an embodiment of the present invention;
[0056] Figure 3 It is an example diagram of contour line annotation provided by an embodiment of the present invention;
[0057] Figure 4 It is an example diagram of key point annotation provided by an embodiment of the present invention;
[0058] Figure 5 It is a schematic structural diagram of a multi-view markerless facial expression capture system provided by an embodiment of the present invention;
[0059] Figure 6 It is a schematic structural diagram of a multi-view markerless facial expression capture device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0061] Please refer to Figure 1 , Figure 1 which is a flowchart of a multi-view markerless facial expression capture method provided by an embodiment of the present invention. The method may include:
[0062] S101: Collect multi-view face images based on a multi-dimensional camera matrix.
[0063] The execution subject of this embodiment is a terminal. This embodiment does not limit the type of the terminal, as long as it can complete the operations of the multi-view markerless facial expression capture method. In this embodiment, the multi-dimensional camera matrix can be referred to Figure 2 , Figure 2 which is an example diagram of a 4D camera matrix provided by this embodiment. Multi-view face image data is collected by high-definition motion cameras with multiple views. The matrix uses 20 color cameras, and the color cameras can obtain RGB three-channel images. The camera matrix uses white LED lights to evenly illuminate the face, and the obtained face texture is more accurate. Time synchronization operations are performed on each camera, and the frame rate of each camera is 60 frames / s. During the recording stage, the user first shoots for 1 s with a neutral expression (i.e., no expression, mouth closed). Then, various preset expressions (smiling, laughing) or lines are recited, and normal shooting is performed.
[0064] S102: Perform two-dimensional key point detection and three-dimensional reconstruction of a neutral-expression face on the multi-view face images to obtain a standard face topology model.
[0065] Based on the above-collected multi-view face images, two-dimensional key point detection and three-dimensional reconstruction of a neutral-expression face are performed in this embodiment to obtain a standard face topology model. It should be noted that the standard face topology model in this embodiment has the ability to make expressions.
[0066] Furthermore, the above-mentioned performing two-dimensional key point detection and three-dimensional reconstruction of a neutral-expression face on the multi-view face images to obtain a standard face topology model may include the following steps:
[0067] Step 11: Based on each part, use a neural network to perform key point and contour detection on the multi-view face images to obtain target key points and a target contour.
[0068] To further accelerate, only extract the frontal face image from each frame of multi-view images for key point and contour detection. The frontal face image already contains the semantic information of the required facial feature key points, so other side face images are no longer needed. Specifically, first perform face detection on the multi-view face images to obtain the face region, then use full-face key point detection on the face region image to obtain the initial facial features and face edge key points, and then use the full-face key points to calculate and obtain the enclosure and segmentation of each facial feature region to obtain the regional images, such as the left and right eye region images, the left and right eyebrow region images, the nose region image, the lip and tooth region image, the chin region image, and the triangular region wrinkle region image. Input the left eye region (including the left eyebrow) image into the left eye key point detection network to regress the left eye contour, the left eyeball center, and the left eyeball contour key points; input the mirror image of the right eye region (including the right eyebrow) image into the right eye key point detection network to regress the right eye contour, the right eyeball center, and the right eyeball contour key points; similarly, input the left and right eyebrow region images to obtain the left and right eyebrow contours; input the nose region image into the nose key point detection network to regress the nose contour and the nasal bridge key points; input the lip and tooth region image into the mouth key point detection network to regress the inner and outer contour key points of the mouth; input the lip and tooth region image into the tooth key point detection network to regress the upper and lower tooth key points; input the chin region image into the chin key point detection network to regress the chin contour; input the triangular region wrinkle region image into the wrinkle key point detection network to regress the wrinkle contour.
[0069] Using each region for key point detection improves the accuracy of face key point detection compared to directly performing key point detection on the full face. All the above key point detection networks use the Hourglass (also known as the hourglass network, which is a convolutional neural network architecture) neural network. After outputting all the heatmaps, perform softmax to obtain the pixel coordinates of the maximum value as the key point coordinates. At the same time, a confidence value can be output. If the confidence is small, it can be judged that the region is occluded and the key points in this region are not credible. Since the above key point detection is an operation by region, the images of each region are passed through the neural networks corresponding to each region in parallel for prediction, and parallel prediction can speed up the key point detection speed.
[0070] Step 12: Based on the neutral expression image of the face in the multi-view face images, use 3D reconstruction technology to obtain the 3D reconstructed face skin model, the reconstructed texture, and the internal and external camera parameters.
[0071] In this embodiment, one frame of multi-view image can be taken from the neutral expression frame in the first second of the multi-view image frames captured by the camera matrix and input to 3D reconstruction software such as colmap or metashape to reconstruct the 3D reconstructed face skin model of the neutral expression and its reconstructed texture, and output the internal and external camera parameters.
[0072] Step 13: Obtain a standard face topology model based on the target key points, the target contour, the 3D reconstructed face skin model of the human face, the reconstructed texture, and the internal and external camera parameters.
[0073] Based on the target key points and the target contour obtained in the above Step 11, as well as the 3D reconstructed face skin model of the human face, the reconstructed texture, and the internal and external camera parameters obtained in Step 12, obtain a standard face topology model capable of expressing emotions. The main purpose of Step 13 is to convert the reconstructed 3D reconstructed face skin model of the human face (without the ability to express emotions) and the reconstructed texture into a standard face topology model.
[0074] Furthermore, the obtaining of the standard face topology model based on the target key points, the target contour, the 3D reconstructed face skin model of the human face, the reconstructed texture, and the internal and external camera parameters may include the following steps:
[0075] Step 131: Obtain a first standard face skin topology model, and label the target key points and the target contour on the first standard face skin topology model to obtain a second standard face skin topology model.
[0076] It should be noted that the first standard face skin topology model can specifically use the hifi topology, or can be obtained using other open-source topologies, such as the topologies of 3DMM (3D Morphable Model, that is, 3D deformable human face model) models such as hifi\bfm\flame. Label the identified target key points and the target contour on the first standard face skin topology model to obtain the corresponding points of the target key points on the first standard face skin topology model. As Figure 3 、 4 shown are the labeled points and the labeled contour on the first standard face skin topology model. Mark the serial numbers of the geometric points of a total of 11 feature points on the eyes, nose, and mouth on the geometry of the first standard face skin topology model.
[0077] Step 132: Based on the second standard face skin topology model and the 3D reconstructed face skin model of the human face, perform key point projection using the internal and external camera parameters, solve for the triangle positions using the nricp_sumner algorithm, and solve for the transformation of the vertices using the nricp_amberg algorithm to obtain a third standard face skin topology model of the human face.
[0078] It should be noted that the second standard facial skin topology model in this embodiment includes the annotation of target key points and target contours. The internal and external parameters of the camera are used to back-project the facial key points on the image, and the three-dimensional coordinates of the geometric points of each frame are calculated. Specifically, in this embodiment, two non-rigid registration algorithms, nricp-sumner and nricp-amberg, are combined. nricp_amberg (Non-Rigid Iterative Closest Point algorithm) can adapt to the target mesh in fewer steps and can also generate sharp edges (only considering vertices and their neighbors); while nricp_sumner is more inclined to retain the original shape and its parameters are easier to adjust. Therefore, in this embodiment, the nricp_sumner algorithm is used to solve the position of the triangle, and the nricp_amberg algorithm is used to solve the vertex transformation. In this embodiment, the nricp_sumner algorithm is first used to completely align the facial key points to align the semantic information, and then the nricp_amberg algorithm is used to perform fine-grained deformation alignment on the remaining parts, which can give full play to the advantages of the two algorithms and make the result of neutral face registration better. The above algorithms can make the obtained standard topology model fit well in high-frequency information such as wrinkles, and at the same time, the pose is completely matched to the three-dimensional reconstructed facial skin model.
[0079] Step 133: Based on the first standard texture, use the texture transfer algorithm to transfer the reconstructed texture to the first standard texture to obtain the second standard texture.
[0080] Specifically, in order to transfer the relatively fragmented texture obtained by reconstruction to the standard texture of metahuman (an innovative digital character creation technology), the following algorithm needs to be used for texture transfer: find the corresponding pixel of the reconstructed texture for each pixel on the standard texture map, and copy the pixel of the reconstructed texture to the standard texture map (i.e., the first standard texture), and finally obtain the transferred standard texture picture (i.e., the second standard texture). The pseudo-code of the texture transfer algorithm is as follows:
[0081] Traverse the uv triangles of the metahuman source model:
[0082] Traverse the pixels inside the uv triangles of the metahuman source model to calculate the barycentric coordinates of the pixels;
[0083] Obtain the three-dimensional point coordinates of the pixel on the patch of the metahuman source model according to the barycentric coordinates;
[0084] Use the KNN algorithm to search for the nearest triangle id of the target mesh with the three-dimensional point coordinates, obtain the barycentric coordinates on the nearest triangle of the target mesh, and finally obtain the coordinates of the target uv;
[0085] Copy the pixel points on the target uv to the texture coordinates of the metahuman source model. The resulting texture may have gaps, and inpaint can be used to repair the gaps to obtain the final texture.
[0086] Step 134: Obtain the standard face topology model based on the third standard face skin topology model and the second standard texture.
[0087] Specifically, in this embodiment, the standard face topology model can be obtained based on the above-mentioned third standard face skin topology model and the second standard texture. Further, lighting can also be taken into account to obtain a more fitting standard face topology model.
[0088] Further, the obtaining of the standard face topology model based on the third standard face skin topology model and the second standard texture may include the following steps:
[0089] Step 1341: Obtain the initial standard face topology model according to the third standard face skin topology model and the second standard texture;
[0090] Step 1342: Render the initial standard face topology model using the spherical harmonic lighting coefficients to obtain a color map;
[0091] Step 1343: Calculate the loss based on the color map and the face neutral expression image in the face multi-view images, backpropagate the loss, and use the adam algorithm to optimize the spherical harmonic lighting coefficients;
[0092] Step 1344: Render the initial standard face topology model using the optimized spherical harmonic lighting coefficients to obtain the standard face topology model.
[0093] In this embodiment, the coefficients of the spherical harmonic lighting are used as the variables to be optimized. The color map is obtained through differentiable rendering. The L1 loss is calculated for the rgb channels of the color map and the face neutral expression image in the multi-view images and backpropagated. The adam (Adaptive Moment Estimation) optimization algorithm is used to optimize the coefficients of the spherical harmonic lighting. The initial standard face topology model is rendered using the optimized spherical harmonic lighting coefficients to obtain the standard face topology model.
[0094] S103: Obtain the non-neutral expression depth map from the face multi-view images.
[0095] The meaning of the non-neutral expression depth map in this embodiment is: the depth map corresponding to the neutral expression image in the multi-view and multi-image of the human face. In order to speed up the entire system in this embodiment, for all frame images except the neutral expression frame images, only the depth map of each photo needs to be output using the 3D reconstruction software colmap or metashape. That is, for the neutral expression image of the human face in the multi-view image of the human face, 3D reconstruction technology is required to obtain the 3D reconstructed facial skin model, reconstructed texture, and internal and external camera parameters of the human face; for the non-neutral expression image of the human face in the multi-view image of the human face, only the depth map needs to be obtained using 3D reconstruction technology. It can be understood that the depth map is an intermediate product for obtaining the 3D reconstructed facial skin model of the human face using 3D reconstruction technology.
[0096] S104: Optimize the model coefficients based on the standard human face topology model, non-neutral expression depth map, and multi-view human face images; the model coefficients include expression basis coefficients, translation coefficients, and Euler angle rotation coefficients.
[0097] In this embodiment, the standard human face topology model obtained by using step S102 is automatically bound to obtain an expression basis including eyes and teeth. Using this expression basis can fit most of the expressions that people can make, and can fit the movements of teeth and eyeballs. Since the facial pose of the user will change during the performance, the least squares method can be used to optimize the model coefficients, where the model coefficients include expression basis coefficients and 3D translation coefficients (x, y, z), 3D Euler angle rotation (x, y, z) coefficients, so that the standard human face topology model is close to the original image (i.e., the multi-view human face image) and the non-neutral expression depth map.
[0098] Furthermore, the above-mentioned optimization of the model coefficients based on the standard human face topology model, non-neutral expression depth map, and multi-view human face images may include the following steps:
[0099] Step 21: Obtain the distance loss value according to the distance from the vertices in the standard human face topology model to the non-neutral expression depth map, calculate the contour loss value according to the contour in the standard human face topology model and the target contour, calculate the key point loss value according to the key points in the standard human face topology model and the target key points, and calculate the total loss value.
[0100] This step mainly calculates the loss, which mainly includes three aspects: distance loss value, contour loss value, and key point loss value.
[0101] Furthermore, the above-mentioned obtaining the distance loss value according to the distance from the vertices in the standard human face topology model to the non-neutral expression depth map may include the following steps:
[0102] Step 211: Project the vertices in the standard face topology model onto the non-neutral expression depth map through the internal and external camera parameters, and calculate the distance and normal angle distance from the vertices to the non-neutral expression depth map;
[0103] Step 212: Set the initial weight of the distance according to the normal angle distance;
[0104] Step 213: Set the mask weight according to the distance from the vertices to the non-neutral expression depth map and the distance threshold;
[0105] Step 214: Obtain the final weight according to the initial weight and the mask weight;
[0106] Step 215: Obtain the distance loss value according to the final weight and the distance from the vertices to the non-neutral expression depth map.
[0107] The distance loss value (ICP loss) is used to measure the loss of the distance from the model vertices to the depth map. Calculating this loss is to make the model close to the depth map after transformation and applying the expression basis coefficients. For the convenience of expression, here the non-neutral expression depth map is simply referred to as the depth map, and the standard face topology model is simply referred to as the model. Project the model vertices onto the depth map through the internal and external camera parameters, calculate the distance and normal angle distance from the model vertices to the depth map, and set the weight of the distance according to the normal angle distance (calculated by multiplying the normal of the pixel point on the depth map and the normal of the standard topology model vertex). The normal difference threshold is 60 degrees. If it exceeds 60 degrees, it means the correspondence is incorrect and the weight is set to 0. And the model has also defined a facial mask in advance, and the mask weight of the points not in the facial mask is 0; then add a distance constraint to the mask weight, and the mask weight of the facial mask with a distance greater than the distance threshold is set to 0, and the rest are set to (distance threshold - distance) / distance threshold. The specific formula is as follows:
[0108] ;
[0109] ;
[0110] ;
[0111] 。
[0112] Among them, mask_weights is the mask weight, length is the distance, threshold is the distance threshold, weights1 is the initial weight, and weights is the final weight.
[0113] Furthermore, the above-mentioned obtaining the contour loss value according to the contour in the standard face topology model and the target contour may include the following steps:
[0114] Step 226: Project the points on the contour line of the standard face topology model onto the multi-view face image through the internal and external camera parameters to obtain two-dimensional contour points;
[0115] Step 227: Determine the target vertical point and the normal line of the target vertical point based on the two-dimensional contour points;
[0116] Step 228: Calculate the contour loss value based on the normal line, the target vertical point, and the two-dimensional contour points.
[0117] In this embodiment, the contour loss value is used to measure the differences between the upper eyelid, lower eyelid, upper lip, lower lip, upper inner lip, lower inner lip, nasolabial fold contour on the standard face topology and the contour line detected by the face key point detection model, so that the standard face topology contour can approach the image contour and the semantics can be aligned. The points on the contour line of the model are projected onto the multi-view face image (this image is a relatively frontal image) through the internal and external camera parameters to obtain two-dimensional contour points. The nearest vertical point of the curve detected by the two-dimensional projected points is used as the corresponding point, and the normal line of this vertical point is calculated using the control points at both ends of the vertical point. The calculation formula of the contour loss value is as follows:
[0118] ;
[0119] where curve_loss is the contour loss value; normals is the normal line of the corresponding point; targets is the corresponding point, that is, the nearest vertical point / target vertical point; vertices are the two-dimensional projected points / projected points.
[0120] Furthermore, in this embodiment, the key point loss value is used to measure the differences between the corners of the mouth, corners of the eyes, under the nose of the model and the key points obtained by the above key point detection model. Make the corners of the eyes and mouth more aligned. Specifically, the L2 norm is used to measure this loss value, and the calculation formula is as follows:
[0121] 。
[0122] where landmark_loss is the key point loss value.
[0123] The calculation formula of the total loss value is:
[0124] 。
[0125] where loss is the total loss value.
[0126] Step 22: According to the total loss value, use the conjugate gradient algorithm to solve for the optimization direction of the least squares method, and use the line search of the Armijo condition to solve for the optimization step size of the least squares method.
[0127] In this embodiment, the total loss value is used to solve the optimization direction through the conjugate gradient algorithm. The conjugate gradient algorithm has a second-order descent speed, which is faster than the first-order gradient descent speed of Adam. Therefore, this algorithm has a faster optimization speed, and the number of iterations is much smaller than the number of variables. After obtaining the optimization direction, the line search of the Armijo condition is used to solve the optimization step size. This algorithm can ensure that each optimization is correct and can converge stably. This algorithm prevents jitter near the optimal point and can obtain the optimal solution stably, accurately, and quickly. Since the Armijo line search optimization and the least squares optimization are used, the model coefficients obtained are faster and more accurate than those obtained by only using Adam optimization.
[0128] Step 23: Based on the optimization direction and the optimization step size, use the least squares method to optimize the expression basis coefficients, translation coefficients, and Euler angle rotation coefficients.
[0129] In this step, the optimization direction and the optimization step size obtained above are used to optimize the model coefficients by the least squares method.
[0130] S105: Input the optimized model coefficients into the standard face topology model to obtain the real face facial expression model.
[0131] In this embodiment, based on the standard face topology model, the real face facial expression model is obtained by using the optimized model coefficients. This real face facial expression model is almost the same as a real person.
[0132] Furthermore, the above-mentioned step of inputting the optimized model coefficients into the standard face topology model to obtain the real face facial expression model may include the following steps:
[0133] Step 31: Input the optimized model coefficients into the standard face topology model to obtain the initial face facial expression model;
[0134] Step 32: Use the distance loss value as the supervision information to optimize all vertex offsets in the initial face facial expression model to obtain the real face facial expression model.
[0135] Specifically, after completing the optimization stage of the expression basis coefficients and the pose coefficients (translation coefficients and Euler angle rotation coefficients), the vertex offsets offset (n vertices * (x, y, z), a total of 3 * n coefficients) in the frontal face area can be further optimized. At this time, only the ICP loss is used as the supervision information, so that the face can fit high-frequency information such as wrinkles.
[0136] Applying the multi-view markerless facial expression capture method provided by the embodiments of the present invention, multi-view images of a human face are collected based on a multi-dimensional camera matrix; two-dimensional key point detection and three-dimensional reconstruction of a neutral expression human face are performed on the multi-view images of the human face to obtain a standard human face topology model; a non-neutral expression depth map is obtained from the multi-view images of the human face; the model coefficients are optimized based on the standard human face topology model, the non-neutral expression depth map, and the multi-view images of the human face; the model coefficients include expression basis coefficients, translation coefficients, and Euler angle rotation coefficients; the optimized model coefficients are input into the standard human face topology model to obtain a real human face facial expression model. The multi-view images of the human face collected by using the multi-dimensional camera matrix in this method can improve the accuracy and integrity of the acquired images; the expression basis coefficients, translation coefficients, and Euler angle rotation coefficients are optimized, so that a more accurate and real human face facial expression model is obtained based on the standard human face topology model, ensuring the accuracy of expressions and poses, and being able to reflect facial detail features and being more similar to a real human face. Moreover, separately detecting the key points and contours of the human face by parts can completely replace manual marking, and also has the advantages of high accuracy and high efficiency; first, the nricp-sumner algorithm is used to align the key points of the human face, and then the nricp-amberg algorithm is used to deform the vertices of the remaining parts onto the mesh, which has stronger deformation ability compared with the traditional technology of optimizing the shape blendshape coefficients, and the obtained model is more similar to a real person; the model obtained after texture transfer has more textures, making the human face more similar to a real person; and illumination estimation is also added, and finally the image obtained by rendering the human face model is almost the same as a real person's photo; moreover, because armijo line search optimization and least squares optimization are used, the model coefficients obtained are faster and more accurate compared with only using adam optimization.
[0137] The multi-view markerless facial expression capture system provided by the embodiments of the present invention is introduced below. The multi-view markerless facial expression capture system described below can be mutually referred to the multi-view markerless facial expression capture method described above.
[0138] Specifically, please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a multi-view markerless facial expression capture system provided by the embodiments of the present invention and may include:
[0139] An acquisition module 100, configured to collect multi-view images of a human face based on a multi-dimensional camera matrix;
[0140] A reconstruction module 200, configured to perform two-dimensional key point detection and three-dimensional reconstruction of a neutral expression human face on the multi-view images of the human face to obtain a standard human face topology model;
[0141] A depth map acquisition module 300, configured to obtain a non-neutral expression depth map from the multi-view images of the human face;
[0142] The model coefficient optimization module 400 is used to optimize the model coefficients based on the standard face topology model, the non-neutral expression depth map, and the multi-view face images; the model coefficients include expression basis coefficients, translation coefficients, and Euler angle rotation coefficients.
[0143] The facial expression capture module 500 is used to input the optimized model coefficients into the standard face topology model to obtain a real face facial expression model.
[0144] Based on the above embodiments, the reconstruction module 200 may include:
[0145] The detection unit is used to detect key points and contours of the multi-view face images based on each part by using a neural network to obtain target key points and a target contour.
[0146] The reconstruction unit is used to obtain a 3D reconstructed face skin model, reconstructed texture, and internal and external camera parameters based on the face neutral expression image in the multi-view face images by using 3D reconstruction technology.
[0147] The model acquisition unit is used to obtain the standard face topology model based on the target key points and target contour, and the 3D reconstructed face skin model, reconstructed texture, and internal and external camera parameters.
[0148] Based on the above embodiments, the model acquisition unit may include:
[0149] The annotation subunit is used to obtain a first standard face skin topology model and annotate the target key points and target contour on the first standard face skin topology model to obtain a second standard face skin topology model.
[0150] The solution transformation subunit is used to project key points based on the second standard face skin topology model and the 3D reconstructed face skin model by using the internal and external camera parameters, solve the triangle positions by using the nricp_sumner algorithm, and solve the transformation of vertices by using the nricp_amberg algorithm to obtain a third standard face skin topology model.
[0151] The texture migration subunit is used to migrate the reconstructed texture to the first standard texture by using a texture migration algorithm based on the first standard texture to obtain a second standard texture.
[0152] The model subunit is used to obtain the standard face topology model based on the third standard face skin topology model and the second standard texture.
[0153] Based on the above embodiments, the model subunit may include:
[0154] An initial model acquisition subunit, configured to obtain an initial standard face topology model according to the third standard face skin topology model and the second standard texture;
[0155] A rendering subunit, configured to render the initial standard face topology model by using spherical harmonic lighting coefficients to obtain a color map;
[0156] A loss calculation subunit, configured to calculate a loss based on the color map and a face neutral expression image in the multi-view face images, and backpropagate the loss, and optimize the spherical harmonic lighting coefficients by using the adam algorithm;
[0157] A rendering optimization unit, configured to render the initial standard face topology model by using the optimized spherical harmonic lighting coefficients to obtain the standard face topology model.
[0158] Based on the above embodiments, the model coefficient optimization module 400 may include:
[0159] A loss calculation unit, configured to obtain a distance loss value according to the distance from a vertex in the standard face topology model to the non-neutral expression depth map, obtain a contour loss value according to the contour in the standard face topology model and a target contour, obtain a key point loss value calculated according to the key points in the standard face topology model and target key points, and calculate a total loss value;
[0160] A solution unit, configured to solve an optimization direction of the least squares method by using the conjugate gradient algorithm according to the total loss value, and solve an optimization step size of the least squares method by using a line search of the armijo condition;
[0161] A coefficient optimization unit, configured to optimize the expression basis coefficients, translation coefficients, and Euler angle rotation coefficients by using the least squares method based on the optimization direction and the optimization step size.
[0162] Based on the above embodiments, the loss calculation unit may include:
[0163] A distance calculation subunit, configured to project a vertex in the standard face topology model onto the non-neutral expression depth map through internal and external camera parameters, and calculate the distance from the vertex to the non-neutral expression depth map and the normal angle distance;
[0164] An initial weight setting subunit, configured to set an initial weight of the distance according to the normal angle distance;
[0165] A mask weight setting subunit, configured to set a mask weight according to the distance from the vertex to the non-neutral expression depth map and a distance threshold;
[0166] A final weight setting subunit, configured to obtain a final weight according to the initial weight and the mask weight;
[0167] A distance loss calculation subunit, configured to obtain the distance loss value according to the final weight and the distance from the vertex to the non-neutral expression depth map;
[0168] A contour point acquisition subunit, configured to project points on the contour line of the standard face topology model onto the multi-view face image through the internal and external parameters of the camera to obtain two-dimensional contour points;
[0169] A vertical point and normal acquisition subunit, configured to determine a target vertical point and the normal of the target vertical point based on the two-dimensional contour points;
[0170] A contour loss calculation subunit, configured to calculate the contour loss value based on the normal, the target vertical point, and the two-dimensional contour points.
[0171] Based on the above embodiments, the facial expression capture module 500 may include:
[0172] An initial human face facial expression model acquisition unit, configured to input the optimized model coefficients into the standard face topology model to obtain an initial human face facial expression model;
[0173] A model optimization unit, configured to use the distance loss value as supervision information to optimize all vertex offsets in the initial human face facial expression model to obtain the real human face facial expression model.
[0174] It should be noted that the order of the modules and units in the above multi-view markerless facial expression capture system can be changed before and after without affecting the logic.
[0175] Applying the multi-view markerless facial expression capture system provided by the embodiments of the present invention, through the acquisition module 100, which is used to acquire multi-view images of a human face based on a multi-dimensional camera matrix; the reconstruction module 200, which is used to perform two-dimensional key point detection and three-dimensional reconstruction of a neutral expression human face on the multi-view images of the human face to obtain a standard human face topology model; the depth map acquisition module 300, which is used to obtain a non-neutral expression depth map from the multi-view images of the human face; the model coefficient optimization module 400, which is used to optimize the model coefficients based on the standard human face topology model, the non-neutral expression depth map and the multi-view images of the human face; the model coefficients include expression basis coefficients, translation coefficients and Euler angle rotation coefficients; the facial expression capture module 500, which is used to input the optimized model coefficients into the standard human face topology model to obtain a real human face facial expression model. The multi-view images of the human face collected by this system using a multi-dimensional camera matrix can improve the accuracy and integrity of the acquired images; optimize the expression basis coefficients, translation coefficients and Euler angle rotation coefficients, so as to obtain a more accurate and realistic human face facial expression model based on the standard human face topology model, ensure the accuracy of expressions and poses, and can reflect facial detail features and be more similar to a real human face. Moreover, detecting the key points and contours of the human face separately by part can completely replace manual marking, and also has the advantages of high accuracy and high efficiency; first use the nricp-sumner algorithm to align the key points of the human face, and then use the nricp-amberg algorithm to deform the vertices of the remaining parts onto the mesh, which has stronger deformation ability compared with the traditional technology of optimizing the shape blendshape coefficients, and the obtained model is more similar to a real person; the model obtained after texture transfer has more textures, making the human face more similar to a real person; and illumination estimation is also added, and finally the image obtained by rendering the human face model is almost the same as a real person's photo; moreover, due to the use of the armijo line search optimization and the least squares optimization, the model coefficients obtained are faster and more accurate compared with only using the adam optimization.
[0176] The multi-view markerless facial expression capture device provided by the embodiments of the present invention will be introduced below. The multi-view markerless facial expression capture device described below can be mutually corresponded and referred to the multi-view markerless facial expression capture method described above.
[0177] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a multi-view markerless facial expression capture device provided by the embodiments of the present invention, and can include:
[0178] A memory 10, which is used to store computer programs;
[0179] A processor 20, which is used to execute the computer program to implement the above-mentioned multi-view markerless facial expression capture method.
[0180] The memory 10, the processor 20, and the communication interface 31 all communicate with each other via the communication bus 32.
[0181] In an embodiment of the present invention, the memory 10 is used to store one or more programs. The program may include program code, and the program code includes computer operation instructions. In an embodiment of the present invention, the memory 10 may store a program for implementing the following functions:
[0182] Collect multi-view face images based on a multi-dimensional camera matrix;
[0183] Perform two-dimensional key point detection and three-dimensional reconstruction of a neutral-expression face on the multi-view face images to obtain a standard face topology model;
[0184] Obtain a non-neutral-expression depth map from the multi-view face images;
[0185] Optimize the model coefficients based on the standard face topology model, the non-neutral-expression depth map, and the multi-view face images; the model coefficients include expression basis coefficients, translation coefficients, and Euler angle rotation coefficients;
[0186] Input the optimized model coefficients into the standard face topology model to obtain a real face facial expression model.
[0187] In a possible implementation, the memory 10 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function, etc.; the data storage area may store data created during use.
[0188] In addition, the memory 10 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include NVRAM. The memory stores an operating system and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof. Among them, the operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic tasks and processing hardware-based tasks.
[0189] The processor 20 may be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic devices. The processor 20 may be a microprocessor or any conventional processor, etc. The processor 20 may call the program stored in the memory 10.
[0190] The communication interface 31 may be an interface of a communication module for connecting to other devices or systems.
[0191] Of course, it should be noted that Figure 6 the structure shown does not constitute a limitation on the multi-view markerless facial expression capture device in the embodiments of the present invention. In actual applications, the multi-view markerless facial expression capture device may include more or fewer components than Figure 6 those shown, or combine certain components.
[0192] Next, the readable storage medium provided by the embodiments of the present invention will be introduced. The computer-readable storage medium described below can be correspondingly referred to the multi-view markerless facial expression capture method described above.
[0193] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above multi-view markerless facial expression capture method are implemented.
[0194] The computer-readable storage medium may include: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0195] In this specification, the various embodiments are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, please refer to the description in the method part.
[0196] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in the form of hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0197] Finally, it should also be noted that in this text, relationships such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0198] The above has introduced in detail a multi-viewpoint markerless facial expression capture method, system, device and computer-readable storage medium provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A multi-view markerless facial expression capture method, characterized in that, Including: Collecting multi-view face images based on a multi-dimensional camera matrix; Performing two-dimensional key point detection and three-dimensional reconstruction of a neutral expression face on the multi-view face images to obtain a standard face topology model; Obtaining a non-neutral expression depth map from the multi-view face images; Optimizing model coefficients based on the standard face topology model, the non-neutral expression depth map, and the multi-view face images; the model coefficients include expression basis coefficients, translation coefficients, and Euler angle rotation coefficients; Inputting the optimized model coefficients into the standard face topology model to obtain a real face facial expression model.
2. The multi-view markerless facial expression capture method according to claim 1, characterized in that Performing two-dimensional key point detection and three-dimensional reconstruction of a neutral expression face on the multi-view face images to obtain a standard face topology model, including: Based on each part, using a neural network to perform key point and contour detection on the multi-view face images to obtain target key points and a target contour; Based on the neutral expression face image in the multi-view face images, using three-dimensional reconstruction technology to obtain a three-dimensional reconstructed face skin model, reconstructed texture, and internal and external camera parameters; Obtaining the standard face topology model based on the target key points and target contour, as well as the three-dimensional reconstructed face skin model, reconstructed texture, and internal and external camera parameters.
3. The multi-view markerless facial expression capture method according to claim 2, wherein, Obtaining the standard face topology model based on the target key points and target contour, as well as the three-dimensional reconstructed face skin model, reconstructed texture, and internal and external camera parameters, including: Obtaining a first standard face skin topology model, and annotating the target key points and target contour on the first standard face skin topology model to obtain a second standard face skin topology model; Based on the second standard face skin topology model and the three-dimensional reconstructed face skin model, performing key point projection using the internal and external camera parameters, solving for triangle positions using the nricp_sumner algorithm, and solving for vertex transformations using the nricp_amberg algorithm to obtain a third standard face skin topology model; Based on a first standard texture, using a texture transfer algorithm to transfer the reconstructed texture to the first standard texture to obtain a second standard texture; Obtaining the standard face topology model based on the third standard face skin topology model and the second standard texture.
4. The multi-view markerless facial expression capture method according to claim 3, wherein Obtaining the standard face topology model based on the third standard face skin topology model and the second standard texture, including: Obtaining an initial standard face topology model according to the third standard face skin topology model and the second standard texture; Rendering the initial standard face topology model using spherical harmonic lighting coefficients to obtain a color map; Calculating a loss based on the color map and the neutral expression face image in the multi-view face images, backpropagating the loss, and using the adam algorithm to optimize the spherical harmonic lighting coefficients; Rendering the initial standard face topology model using the optimized spherical harmonic lighting coefficients to obtain the standard face topology model.
5. The multi-view markerless facial expression capture method according to any one of claims 1 to 4, characterized in that Optimizing model coefficients based on the standard face topology model, the non-neutral expression depth map, and the multi-view face images, including: A distance loss value is obtained according to the distance from the vertices in the standard face topology model to the non-neutral expression depth map, a contour loss value is calculated according to the contour in the standard face topology model and the target contour, a key-point loss value is calculated according to the key points in the standard face topology model and the target key points, and a total loss value is calculated; According to the total loss value, the optimization direction of the least squares method is obtained by using the conjugate gradient algorithm, and the optimization step size of the least squares method is obtained by using the line search of the Armijo condition; Based on the optimization direction and the optimization step size, the expression basis coefficient, the translation coefficient and the Euler angle rotation coefficient are optimized by using the least squares method.
6. The multi-view markerless facial expression capture method according to claim 5, wherein, A distance loss value is obtained according to the distance from the vertices in the standard face topology model to the non-neutral expression depth map, and a contour loss value is calculated according to the contour in the standard face topology model and the target contour, including: The vertices in the standard face topology model are projected onto the non-neutral expression depth map through the internal and external parameters of the camera, and the distance and the normal angle distance from the vertices to the non-neutral expression depth map are calculated; The initial weight of the distance is set according to the normal angle distance; The mask weight is set according to the distance from the vertices to the non-neutral expression depth map and the distance threshold; The final weight is obtained according to the initial weight and the mask weight; The distance loss value is obtained according to the final weight and the distance from the vertices to the non-neutral expression depth map; The points on the contour line of the standard face topology model are projected onto the multi-view face image through the internal and external parameters of the camera to obtain two-dimensional contour points; Based on the two-dimensional contour points, the target vertical point and the normal line of the target vertical point are determined; The contour loss value is calculated based on the normal line, the target vertical point and the two-dimensional contour points.
7. The multi-view markerless facial expression capture method according to claim 6, characterized in that, The optimized model coefficients are input into the standard face topology model to obtain a real face facial expression model, including: The optimized model coefficients are input into the standard face topology model to obtain an initial face facial expression model; Taking the distance loss value as supervision information, all vertex offsets in the initial face facial expression model are optimized to obtain the real face facial expression model.
8. A multi-view markerless facial expression capture system, characterized in that Including: An acquisition module for acquiring multi-view face images based on a multi-dimensional camera matrix; A reconstruction module for performing two-dimensional key-point detection and three-dimensional reconstruction of a neutral-expression face on the multi-view face images to obtain a standard face topology model; A depth map acquisition module for acquiring a non-neutral expression depth map from the multi-view face images; A model coefficient optimization module for optimizing model coefficients based on a standard face topology model, a non-neutral expression depth map and the multi-view face images; the model coefficients include an expression basis coefficient, a translation coefficient and an Euler angle rotation coefficient; A facial expression capture module for inputting the optimized model coefficients into the standard face topology model to obtain a real face facial expression model.
9. A multi-view markerless facial expression capture device, characterized in that, Including: A memory for storing a computer program; A processor, which is configured to implement the multi-view markerless facial expression capture method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are loaded and executed by a processor, the multi-view markerless facial expression capture method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Facial motion capture method based on deep learning
CN120976992A
A deep learning-based facial motion capture method
CN120976992B
VR virtual character interaction system and method based on user-defined expression and mouth shape synchronization
CN121742641A