Three-dimensional face model reconstruction method, system, intelligent terminal and storage medium
By combining aesthetic, identity, and 3D supervised training, the problem of insufficient aesthetic perspective in 3D face model reconstruction is solved, achieving high aesthetic consistency and high accuracy in 3D face model reconstruction, which is suitable for applications such as virtual makeup.
Patent Information
- Application Number
- CN202411353182.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-09-26
AI Technical Summary
Existing 3D face model reconstruction technology does not adequately consider aesthetics, resulting in inconsistent visual appeal between the reconstructed 3D face model and the original image. Furthermore, it lacks robustness and reconstruction accuracy, making it difficult to meet the application requirements with high aesthetic demands.
By acquiring two-dimensional face images and facial key point information, and combining them with the key point information of the initial three-dimensional face model and facial projection image, a three-dimensional face model reconstruction network is used for reconstruction. A beauty scoring network is introduced as part of the reconstruction loss. Combined with aesthetic, identity and three-dimensional supervised losses, hybrid supervised training is carried out until the reconstruction termination condition is met.
It achieves visual consistency between the reconstructed 3D face model and the original image, improves the robustness and accuracy of reconstruction, and is suitable for applications with high aesthetic requirements such as virtual makeup and digital entertainment.
Smart Images

Figure CN119295658B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and particularly relates to a three-dimensional face model reconstruction method and system, an intelligent terminal and a storage medium. BACKGROUND
[0002] With the development of science and technology, especially the development of computer vision and image processing technology, users have higher and higher requirements for image processing. For example, it is necessary to realize three-dimensional face model reconstruction based on two-dimensional images to meet the subsequent three-dimensional data processing requirements.
[0003] In the related art, when performing three-dimensional face model reconstruction, a deep learning-based scheme can be used to utilize the powerful feature extraction capability of a deep network to extract high-dimensional features from a two-dimensional image and generate a corresponding three-dimensional model. However, the related art has a problem that when performing three-dimensional face model reconstruction, only the geometric features of the image are considered, and the aesthetic aspect is not considered, which is not conducive to ensuring that the reconstructed three-dimensional face model has consistent visual aesthetics with the original image.
[0004] Therefore, the related art still needs to be improved and developed. SUMMARY
[0005] The main purpose of the present application is to provide a three-dimensional face model reconstruction method and system, an intelligent terminal and a storage medium, which aims to solve the technical problem that in the related art, when performing three-dimensional face model reconstruction, only the geometric features of the image are considered, and the aesthetic aspect is not considered, which is not conducive to ensuring that the reconstructed three-dimensional face model has consistent visual aesthetics with the original image.
[0006] To achieve the above purpose, the first aspect of the present application provides a three-dimensional face model reconstruction method, wherein the three-dimensional face model reconstruction method comprises:
[0007] obtaining a two-dimensional face image corresponding to a target object and first facial key point information corresponding to the two-dimensional face image;
[0008] obtaining an initial three-dimensional face model, obtaining a face projection image based on the initial three-dimensional face model, and obtaining second facial key point information corresponding to the face projection image;
[0009] performing three-dimensional face model reconstruction through a pre-set three-dimensional face model reconstruction network according to the two-dimensional face image, the first facial key point information, the face projection image and the second facial key point information, to update the initial three-dimensional face model;
[0010] determine a reconstruction loss according to the updated initial three-dimensional face model and the two-dimensional face image, wherein the reconstruction loss comprises a beauty loss determined based on a preset beauty score network;
[0011] If the reconstruction loss does not satisfy a preset reconstruction termination condition, return to perform the steps of obtaining the face projection image based on the initial three-dimensional face model and the second face key point information corresponding to the face projection image until the reconstruction loss satisfies the preset reconstruction termination condition, so as to obtain a target three-dimensional face model matched with the target object after reconstruction is completed.
[0012] Optionally, the obtaining of the two-dimensional face image corresponding to the target object and the first face key point information corresponding to the two-dimensional face image comprises:
[0013] obtaining an initial face image corresponding to the target object;
[0014] preprocessing the initial face image according to a preset preprocessing operation to obtain the two-dimensional face image corresponding to the target object, wherein the preprocessing operation comprises at least one of standardization, cropping and scaling;
[0015] performing face key point positioning on the two-dimensional face image through a preset face key point positioning network to obtain the first face key point information corresponding to the two-dimensional face image.
[0016] Optionally, the obtaining of the initial three-dimensional face model, the obtaining of the face projection image based on the initial three-dimensional face model and the second face key point information corresponding to the face projection image comprise:
[0017] obtaining an initial three-dimensional face model;
[0018] projecting a plurality of different face poses of the initial three-dimensional face model to obtain face projection images corresponding to the plurality of different face poses;
[0019] performing face key point positioning on the face projection images through a preset face key point positioning network to obtain the second face key point information corresponding to the face projection images.
[0020] Optionally, the three-dimensional face model reconstruction according to the two-dimensional face image, the first face key point information, the face projection image and the second face key point information through a preset three-dimensional face model reconstruction network to update the initial three-dimensional face model comprises:
[0021] inputting the two-dimensional face image, the first facial key point information, the face projection image, and the second facial key point information into the three-dimensional face model reconstruction network to obtain a rendering control vector generated by the three-dimensional face model reconstruction network, wherein the rendering control vector is used to represent a camera pose, a lighting parameter, a face shape parameter, and a face texture parameter;
[0022] The three-dimensional face model reconstruction network is used to reconstruct a three-dimensional face model according to the rendering control vector, and the reconstructed three-dimensional face model is taken as an updated initial three-dimensional face model.
[0023] Optionally, the three-dimensional face model reconstruction network further generates a rendering image corresponding to the updated initial three-dimensional face model.
[0024] The reconstruction loss is determined according to the updated initial three-dimensional face model and the two-dimensional face image, and the reconstruction loss includes:
[0025] The first beauty score and the second beauty score are obtained by using a pre-trained beauty score network.
[0026] The beauty loss is determined according to the first beauty score and the second beauty score.
[0027] Optionally, the reconstruction loss further includes an identity loss and a perception loss, and the perception loss includes a pixel loss and a key point loss.
[0028] The reconstruction loss is determined according to the updated initial three-dimensional face model and the two-dimensional face image, and the reconstruction loss further includes:
[0029] The first identity encoding vector corresponding to the two-dimensional face image and the second identity encoding vector corresponding to the rendering image are obtained by using a pre-trained face recognition network, and the identity loss is determined according to a cosine distance between the first identity encoding vector and the second identity encoding vector.
[0030] The pixel loss is determined according to a color error of a pixel corresponding to the two-dimensional face image and the rendering image.
[0031] The third facial key point information corresponding to the rendering image is obtained, and the key point loss is determined according to the first facial key point information and the third facial key point information.
[0032] In the calculation process of the key point loss, the calculation weight of a key point in a facial contour region and a facial feature region in the two-dimensional face image and the rendering image is higher than the calculation weight of a key point in other regions.
[0033] Optionally, the reconstruction loss further comprises a three-dimensional supervision loss.
[0034] The reconstruction loss is determined according to the updated initial three-dimensional face model and the two-dimensional face image, and further comprises:
[0035] A standard three-dimensional face model is obtained.
[0036] The standard three-dimensional face model is calibrated to be consistent with the pose of the updated initial three-dimensional face model.
[0037] The three-dimensional supervision loss is determined according to the Euclidean distance between the corresponding points of the standard three-dimensional face model and the updated initial three-dimensional face model.
[0038] The second aspect of the present application provides a three-dimensional face model reconstruction system, wherein the three-dimensional face model reconstruction system comprises:
[0039] A first data acquisition module is configured to acquire a two-dimensional face image corresponding to a target object and first facial key point information corresponding to the two-dimensional face image.
[0040] A second data acquisition module is configured to acquire an initial three-dimensional face model, a facial projection image corresponding to the initial three-dimensional face model, and second facial key point information corresponding to the facial projection image.
[0041] A three-dimensional model reconstruction module is configured to perform three-dimensional face model reconstruction through a preset three-dimensional face model reconstruction network according to the two-dimensional face image, the first facial key point information, the facial projection image, and the second facial key point information, to update the initial three-dimensional face model.
[0042] A loss determination module is configured to determine a reconstruction loss according to the updated initial three-dimensional face model and the two-dimensional face image, wherein the reconstruction loss comprises a beauty degree loss determined based on a preset beauty score network.
[0043] A reconstruction control module is configured to return to trigger the second data acquisition module to perform the step of acquiring the facial projection image based on the initial three-dimensional face model and the second facial key point information corresponding to the facial projection image until the reconstruction loss satisfies the preset reconstruction termination condition, if the reconstruction loss does not satisfy the preset reconstruction termination condition, to obtain a target three-dimensional face model matched with the target object after reconstruction is completed.
[0044] The third aspect of the present application provides an intelligent terminal, which comprises a memory, a processor, and a three-dimensional face model reconstruction program stored in the memory and executable on the processor. The three-dimensional face model reconstruction program, when executed by the processor, implements the steps of any one of the three-dimensional face model reconstruction methods.
[0045] The fourth aspect of the present application provides a computer-readable storage medium, which stores a three-dimensional face model reconstruction program. The three-dimensional face model reconstruction program, when executed by a processor, implements the steps of any one of the three-dimensional face model reconstruction methods.
[0046] As can be seen from the above, in the present application, a two-dimensional face image corresponding to a target object and first facial key point information corresponding to the two-dimensional face image are obtained. An initial three-dimensional face model is obtained, and a face projection image and second facial key point information corresponding to the face projection image are obtained based on the initial three-dimensional face model. A three-dimensional face model reconstruction network is used to perform three-dimensional face model reconstruction based on the two-dimensional face image, the first facial key point information, the face projection image, and the second facial key point information, so as to update the initial three-dimensional face model. A reconstruction loss is determined based on the updated initial three-dimensional face model and the two-dimensional face image, wherein the reconstruction loss comprises a beauty loss determined based on a preset beauty score network. If the reconstruction loss does not satisfy a preset reconstruction termination condition, the step of obtaining the face projection image based on the initial three-dimensional face model and the second facial key point information corresponding to the face projection image is returned to be executed until the reconstruction loss satisfies the preset reconstruction termination condition, so as to obtain a target three-dimensional face model matching the target object after reconstruction.
[0047] Compared with the prior art, in the present application, the three-dimensional face model reconstruction network is used to perform three-dimensional face model reconstruction based on the two-dimensional face image corresponding to the target object and the first facial key point information corresponding thereto, the initial three-dimensional face model, the face projection image corresponding to the initial three-dimensional face model, and the second facial key point information corresponding to the face projection image. In the reconstruction process, not only the geometric features in the image are considered, but also the beauty loss determined based on the preset beauty score network is taken as part of the reconstruction loss. Thus, in the process of determining whether further reconstruction is needed based on the reconstruction loss, an aesthetic judgment is also made, which is beneficial to ensuring that the three-dimensional face model after reconstruction has consistent visual aesthetics with the original image. BRIEF DESCRIPTION OF DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description only constitute some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort.
[0049] Figure 1 is a flow diagram of a three-dimensional face model reconstruction method provided by an embodiment of the present application;
[0050] Figure 2 is a specific flow diagram of a three-dimensional face model reconstruction method provided by an embodiment of the present application;
[0051] Figure 3 is a data processing diagram of a three-dimensional face model reconstruction network provided by an embodiment of the present application;
[0052] Figure 4 is a component module diagram of a three-dimensional face model reconstruction system provided by an embodiment of the present application;
[0053] Figure 5 is an internal structure principle diagram of a smart terminal provided by an embodiment of the present application. DETAILED DESCRIPTION
[0054] In the following description, specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, persons skilled in the art will understand that the present application can be practiced without these specific details. In other instances, well-known systems, methods, circuits, and devices are not shown or described in detail in order to avoid obscuring the present application.
[0055] It should be understood that, when used in the specification and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0056] It should also be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms, unless the context clearly indicates otherwise.
[0057] It should also be further understood that the term "and / or" as used in the specification and in the claims, means any one of the items, or combinations of items, listed are possible and includes all possible combinations.
[0058] As used in the specification and in the claims, the term "if can be interpreted as meaning "when," or "upon," or "in response to determining," or "in response to classifying," depending on the context. Similarly, the phrase "if it is determined" or "if it is classified [that a described condition or event]" can be interpreted as meaning "upon determining," or "in response to determining," or "upon classifying [a described condition or event]," or "in response to classifying [a described condition or event]," depending on the context.
[0059] The technical solutions in the embodiments of the present application are described clearly and completely below in combination with the drawings of the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0060] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present application, but the present application can also be implemented in other manners different from those described herein, and those of ordinary skill in the art can make similar generalizations without departing from the connotation of the present application, therefore the present application is not limited to the specific embodiments disclosed below.
[0061] At present, users have higher and higher requirements for image processing, for example, it is required to implement reconstruction of a three-dimensional face model based on a two-dimensional image to meet the needs of subsequent three-dimensional data processing. Existing three-dimensional face reconstruction technologies include a geometry-based method and a deep learning-based method.
[0062] Commonly used geometry-based methods include Multi-view Stereo, Shape from Contour, Photometric Stereo, etc. Multi-view Stereo method calculates the disparity of each pixel through multiple images taken from different angles, and then estimates the three-dimensional shape of the object. The advantage of this method is its high geometric accuracy, but it requires high requirements for hardware equipment and shooting conditions. In order to achieve high-precision reconstruction, multiple cameras are needed to shoot at the same time, and accurate camera calibration and image matching algorithms are required. Shape from Contour method is based on the geometric characteristics of the object contour, which analyzes the changes of the object contour under different angles to estimate the three-dimensional shape of the object. This method works well when dealing with simple geometric shapes, but has limited reconstruction ability for complex shapes and textures. Photometric Stereo method calculates the normal vector of the object surface through multiple images taken under different lighting conditions, and then estimates the three-dimensional shape of the object. Photometric Stereo method can handle surface texture to some extent, but has great challenges in dealing with complex lighting conditions and smooth surfaces. Geometry-based methods usually require complex hardware settings and strict shooting conditions, and have great limitations in dealing with complex scenes. Especially when generating high-quality face models, these methods often fail to meet the needs of practical applications.
[0063] With the rise of deep learning technology, Convolutional Neural Networks (CNN) and Generative Adversarial Networks (GAN) have gradually become the mainstream methods in the field of three-dimensional face reconstruction. These methods use the powerful feature extraction ability of deep networks to extract high-dimensional features from two-dimensional images and generate corresponding three-dimensional models. Deep learning-based methods: These methods usually use convolutional neural networks to extract facial features from a single or multiple two-dimensional images, and then generate three-dimensional shapes through regression or classification models.
[0064] Although existing technologies have achieved some success in the field of three-dimensional face reconstruction, there are still several key problems:
[0065] Lack of aesthetic consistency: existing methods mainly focus on geometric accuracy, lack of consideration of aesthetic consistency, resulting in generated three-dimensional models that are visually inconsistent with the aesthetics of the original image. This is particularly evident in applications that require high aesthetic requirements, such as virtual makeup, digital entertainment, and plastic surgery recommendations.
[0066] Lack of robustness: Existing methods exhibit low robustness in handling face occlusion, expression changes, and complex lighting conditions, leading to inconsistent three-dimensional models across different scenarios. This makes it difficult for these methods to be widely deployed in practical applications, especially those that require handling a large number of complex scenarios.
[0067] Insufficient reconstruction accuracy: Many existing methods lack real 3D face data supervision, making it difficult for trained models to accurately reconstruct the true scale of a face. However, accurate size information is crucial for tasks such as VR, AR, and plastic surgery recommendations, and will significantly impact the effectiveness of these downstream tasks.
[0068] Specifically, current research on facial aesthetics has gradually shifted from traditional subjective evaluation to objective analysis based on data and algorithms. Two-dimensional aesthetic analysis and beautification techniques based on images have been fully developed. The three-dimensional facial data used in current three-dimensional aesthetic research can be divided into three-dimensional faces collected by 3D cameras and three-dimensional faces obtained through reconstruction algorithms. The 3D camera acquisition requires high time and economic costs, and it is difficult to obtain enough beautiful sample data. Three-dimensional face reconstruction technology can accurately reconstruct the corresponding three-dimensional face from a single or a group of images, and it plays an important role in virtual reality, film production, game development, medical diagnosis, and aesthetic analysis.
[0069] Traditional three-dimensional face reconstruction techniques mainly rely on multi-view stereo vision or geometric methods such as contour-based shape recovery. These methods usually require multiple images taken from different angles and generate three-dimensional models by calculating the three-dimensional shape of the object. However, such methods have high requirements for hardware devices and shooting conditions, and show great limitations in handling complex scenarios. With the development of deep learning technology, three-dimensional face reconstruction methods based on convolutional neural networks and generative adversarial networks have attracted attention. These methods can extract high-dimensional features from a single or a small number of two-dimensional images and generate high-quality three-dimensional models. However, most existing methods mainly focus on geometric accuracy and pay less attention to facial aesthetic consistency, resulting in differences in aesthetics between the reconstructed results and the original images, making it difficult to meet the needs of three-dimensional face aesthetic research and downstream applications.
[0070] To solve at least one of the above technical problems, in the scheme, a two-dimensional face image corresponding to a target object and first facial key point information corresponding to the two-dimensional face image are obtained; an initial three-dimensional face model is obtained, a face projection image is obtained based on the initial three-dimensional face model, and second facial key point information corresponding to the face projection image is obtained; three-dimensional face model reconstruction is performed through a preset three-dimensional face model reconstruction network based on the two-dimensional face image, the first facial key point information, the face projection image, and the second facial key point information, to update the initial three-dimensional face model; a reconstruction loss is determined based on the updated initial three-dimensional face model and the two-dimensional face image, wherein the reconstruction loss includes a beauty loss determined based on a preset beauty score network; if the reconstruction loss does not satisfy a preset reconstruction termination condition, the step of obtaining the face projection image based on the initial three-dimensional face model and the second facial key point information corresponding to the face projection image is returned to be executed until the reconstruction loss satisfies the preset reconstruction termination condition, to obtain a target three-dimensional face model matched with the target object after reconstruction.
[0071] Compared with the prior art, in the scheme, three-dimensional face model reconstruction is performed based on the two-dimensional face image corresponding to the target object and the first facial key point information corresponding thereto, the initial three-dimensional face model, the face projection image corresponding to the initial three-dimensional face model, and the second facial key point information corresponding to the face projection image through the three-dimensional face model reconstruction network. In the reconstruction process, not only the geometric features in the image are considered, but also the beauty loss determined based on the preset beauty score network is taken as part of the reconstruction loss, so that, in the process of determining whether further reconstruction is needed based on the reconstruction loss, the judgment is also made from the aesthetic point of view, which is beneficial to guarantee that the three-dimensional face model after reconstruction has consistent visual beauty with the original image.
[0072] Exemplary method
[0073] As Figure 1 shown, the embodiment of the present application provides a three-dimensional face model reconstruction method, specifically, the method includes the following steps:
[0074] In step S100, a two-dimensional face image corresponding to a target object and first facial key point information corresponding to the two-dimensional face image are obtained.
[0075] The target object is a user who needs to be reconstructed into a three-dimensional face model. In the embodiment of the present application, the three-dimensional face model of the target object is reconstructed based on the two-dimensional face image corresponding to the target object to obtain a target three-dimensional face model capable of representing the facial features of the target object. The first facial key point information is used to represent the positions of the facial key points in the two-dimensional face image. The first facial key point information can be automatically identified from the two-dimensional face image or specified by the user, which is not limited here.
[0076] In the embodiment of the present application, the two-dimensional face image corresponding to the target object and the first facial key point information corresponding to the two-dimensional face image are obtained as follows:
[0077] An initial face image corresponding to the target object is obtained.
[0078] The initial face image is preprocessed according to a preset preprocessing operation to obtain a two-dimensional face image corresponding to the target object. The preprocessing operation includes at least one of standardization, cropping, and scaling.
[0079] The two-dimensional face image is positioned by a preset facial key point positioning network to obtain first facial key point information corresponding to the two-dimensional face image.
[0080] The initial face image is an image containing the face of the target object, and the size of the initial face image may not be consistent. Preprocessing can improve the accuracy and efficiency of subsequent processing. It should be noted that in the process of three-dimensional face model reconstruction, facial segmentation technology can also be used to process the key areas of the face, pay more attention to the key areas such as facial features, and reduce the attention weight of occlusion, decoration, etc., thereby enhancing the robustness of the model in processing complex scenes (such as occlusion and poor lighting).
[0081] In a specific application scenario, a two-dimensional initial face image is received and preprocessed, such as standardization, cropping, scaling, etc. The preprocessed image is unified to 400*400 pixels (the specific size can be set and adjusted according to actual needs), and then the facial key points of each image are detected and stored as a reference by the MTCNN facial key point positioning network. The input image is segmented by a pre-trained facial segmentation network, a higher weight coefficient is applied to the regions such as facial features and facial contours, a lower weight is applied to glasses, hair, beard, accessories and other occlusions, the facial weight is stored as a gray image and used as the above-mentioned two-dimensional face image for subsequent calculation process.
[0082] In step S200, an initial three-dimensional face model is obtained, a face projection image is obtained based on the initial three-dimensional face model, and second facial key point information corresponding to the face projection image is obtained.
[0083] The initial three-dimensional face model is a pre-set three-dimensional face model, and the initial three-dimensional face model is adjusted in the embodiments of the present application until a target three-dimensional face model matching the target object is obtained.
[0084] Specifically, the initial three-dimensional face model is obtained, the face projection image is obtained based on the initial three-dimensional face model, and the second face key point information corresponding to the face projection image.
[0085] The initial three-dimensional face model is obtained.
[0086] The initial three-dimensional face model is projected for a plurality of different face poses to obtain face projection images corresponding to the plurality of different face poses.
[0087] The face key points of the face projection image are located by a pre-set face key point positioning network to obtain second face key point information corresponding to the face projection image.
[0088] Specifically, the initial three-dimensional face model is projected for a plurality of different face poses to obtain face projection images corresponding to a plurality of different angles, thereby improving the accuracy of three-dimensional face model reconstruction. In one application scenario, the face poses include a front view and left and right side views of the three-dimensional face model, and the corresponding face projection images include a front view and left and right side views (the left and right side views can be left and right head turns of 30° to 45°), thereby improving the standardization of the subsequent processing process. In another application scenario, the corresponding face poses can be determined according to the angle of the face in the two-dimensional face image, thereby projecting the initial three-dimensional face model for the corresponding poses to obtain a face projection image with the same angle as the two-dimensional face image, thereby improving the accuracy of three-dimensional face model reconstruction.
[0089] In one specific application scenario, the initial three-dimensional face model is obtained based on a 3D face dataset LYHM dataset and used as a 3D supervision input. The dataset contains 3D scanned face models of volunteers as real references, and face images (i.e., face projection images) of different angles are used as inputs of a reconstruction network (i.e., a pre-set three-dimensional face model reconstruction network). Here, MTCNN is also used to detect and record the key points of the face image as a reference. The real three-dimensional face (i.e., the initial three-dimensional face model) and the reconstructed three-dimensional face (i.e., the reconstructed three-dimensional face model) are all unified into the format of a three-dimensional variable face model FLAME, represented by a same number of corresponding points.
[0090] Step S300, according to the above-mentioned two-dimensional face image, the above-mentioned first face key point information, the above-mentioned face projection image and the above-mentioned second face key point information, the three-dimensional face model reconstruction is carried out through the preset three-dimensional face model reconstruction network to update the initial three-dimensional face model.
[0091] In an application scenario, the three-dimensional face model reconstruction network is constructed based on ResNet50, and after inputting the image and the corresponding face key point information, the reconstruction network will output a vector, which contains camera pose, lighting parameters, face shape parameters and face texture parameters to render and generate a reconstructed 3D face, a rendered image and the face key points corresponding to the rendered image.
[0092] Specifically, the three-dimensional face model reconstruction is carried out according to the above-mentioned two-dimensional face image, the above-mentioned first face key point information, the above-mentioned face projection image and the above-mentioned second face key point information through the preset three-dimensional face model reconstruction network to update the initial three-dimensional face model, comprising:
[0093] The above-mentioned two-dimensional face image, the above-mentioned first face key point information, the above-mentioned face projection image and the above-mentioned second face key point information are input into the above-mentioned three-dimensional face model reconstruction network to obtain a rendering control vector generated by the above-mentioned three-dimensional face model reconstruction network, wherein the rendering control vector is used to represent camera pose, lighting parameters, face shape parameters and face texture parameters;
[0094] The three-dimensional face model reconstruction is carried out according to the above-mentioned rendering control vector through the above-mentioned three-dimensional face model reconstruction network, and the three-dimensional face model obtained by the reconstruction is taken as the updated initial three-dimensional face model.
[0095] Step S400, according to the updated initial three-dimensional face model and the above-mentioned two-dimensional face image, a reconstruction loss is determined, wherein the reconstruction loss includes a beauty loss determined based on a preset beauty score network.
[0096] In an application scenario, the function of the mixed supervision model loss is designed in the embodiment of the application to calculate the corresponding reconstruction loss, and the mixed supervision is carried out based on the reconstruction loss to obtain a high-quality three-dimensional face reconstruction model with high precision, high beauty consistency and high robustness.
[0097] Specifically, the three-dimensional face model reconstruction network also generates a rendered image corresponding to the updated initial three-dimensional face model.
[0098] The reconstruction loss is determined according to the updated initial three-dimensional face model and the above-mentioned two-dimensional face image, comprising:
[0099] The first beauty score corresponding to the two-dimensional face image and the second beauty score corresponding to the rendered image are obtained through a pre-trained beauty score network.
[0100] The beauty loss is determined according to the first beauty score and the second beauty score.
[0101] Further, the reconstruction loss further includes an identity loss and a perception loss, and the perception loss includes a pixel loss and a key point loss.
[0102] The reconstruction loss is determined according to the updated initial three-dimensional face model and the two-dimensional face image, and further includes:
[0103] The first identity encoding vector corresponding to the two-dimensional face image and the second identity encoding vector corresponding to the rendered image are obtained through a pre-trained face recognition network, and the identity loss is determined according to the cosine distance between the first identity encoding vector and the second identity encoding vector.
[0104] The pixel loss is determined according to the color error of the corresponding pixels between the two-dimensional face image and the rendered image.
[0105] The third facial key point information corresponding to the rendered image is obtained, and the key point loss is determined according to the first facial key point information and the third facial key point information.
[0106] In the calculation process of the key point loss, the calculation weight of the key points in the facial contour region and the facial feature region in the two-dimensional face image and the rendered image is higher than the calculation weight of the key points in other regions.
[0107] Further, the reconstruction loss further includes a three-dimensional supervision loss.
[0108] The reconstruction loss is determined according to the updated initial three-dimensional face model and the two-dimensional face image, and further includes:
[0109] A preset standard three-dimensional face model is obtained.
[0110] The standard three-dimensional face model is calibrated to make the pose of the standard three-dimensional face model consistent with that of the updated initial three-dimensional face model.
[0111] The three-dimensional supervision loss is determined according to the Euclidean distance between the corresponding points of the standard three-dimensional face model and the updated initial three-dimensional face model.
[0112] In the embodiments of the present application, the set reconstruction loss can be divided into a two-dimensional supervised loss (i.e., 2D supervised loss) and a three-dimensional supervised loss (i.e., 3D supervised loss). Among them, the two-dimensional supervised loss includes beauty loss, identity loss and perception loss.
[0113] Specifically, an aesthetic feature can be extracted from a two-dimensional image by a pre-trained beauty scoring network, and a beauty consistency loss can be calculated, which is used to guide the three-dimensional reconstruction process to ensure that the generated target three-dimensional face model maintains the same beauty as the input face image. Among them, the aesthetic feature is a high-dimensional feature vector extracted by the pre-trained beauty scoring network, which can be mapped to a beauty score (for example, 1 to 5 points, the higher the score, the more beautiful) through a fully connected layer. The above pre-trained beauty scoring network can be AlexNet, ResNet, ResNeXt, etc., which is not limited here. In the embodiments of the present application, the consistency of the face beauty before and after reconstruction is ensured by approximating the beauty scores of the two-dimensional face image I and the rendered image , and the specific beauty loss function is shown in the following formula (1):
[0114] ;
[0115] Among them, represents the beauty loss, represents the second beauty score corresponding to the rendered image, represents the first beauty score corresponding to the two-dimensional face image.
[0116] Further, the identity feature is extracted from the input image by using a convolutional neural network, and is mapped to an identity hidden space to ensure that the generated three-dimensional face model has consistent identity information with the original image. In the embodiments of the present application, the face recognition network ArcFace (denoted as ) is used to extract features from the original two-dimensional face image and the rendered image , calculate the extracted identity encoding vectors before and after reconstruction, and calculate the cosine distance as the identity loss. The identity loss function is shown in the following formula (2):
[0117] ;
[0118] Among them, represents the identity loss, represents the second identity encoding vector corresponding to the rendered image, represents the first identity encoding vector corresponding to the two-dimensional face image.
[0119] Furthermore, the pixel loss and key facial defects of each pixel in the rendered image and the 2D face image, determined based on facial segmentation weights, can be calculated to determine the perceptual loss. In this embodiment, the perceptual loss includes pixel loss and key point loss. Pixel loss in the perceptual loss... Grayscale image based on facial weights Color error of each pixel of the face before and after reconstruction The following weighted calculation is performed to obtain:
[0120] ;
[0121] Keypoint loss in perceptual loss involves detecting original facial keypoints during the 2D supervised input stage. Reconstructed facial key points from network output The Euclidean distance is calculated as shown in the following formula (4):
[0122] ;
[0123] in, Represents key point losses, Indicates the first The weights of facial key points. It should be noted that, when calculating here, the weights of key points in the facial contour and feature areas can also be set to be higher.
[0124] Furthermore, after aligning and calibrating the updated initial 3D face model and the preset standard 3D face model through facial key points, the Euclidean distance between them is calculated as the 3D supervised loss. In this embodiment, the aforementioned preset standard 3D face model is 3D face data from the publicly available dataset LYHM, specifically obtained by pre-collecting standard user head data using a 3D imaging device. 3D reconstruction is an ill-posed problem, meaning that multiple possible 3D face models may exist for the same face image, all of which can produce the same rendered image. Therefore, introducing 3D supervised loss improves the accuracy of the model in face reconstruction, aiming to obtain a more suitable and accurate reconstruction model.
[0125] In this application, pre-positioning and calibration are first performed using 68 facial key points located in the standard 3D face model and 68 facial key points extracted from the reconstructed (updated) initial 3D face model to ensure that the positions and poses of the two face models are consistent in 3D space. Then, the position and pose of each point in the standard 3D face model are calculated. Corresponding points in the reconstructed (updated) initial 3D face model The Euclidean distance between them is then used as the average error of all points as the 3D supervision loss, as shown in the following formula (5):
[0126] ;
[0127] wherein, represents a three-dimensional supervision loss, represents the total number of facial key points, which is 68 in the present application, but is not limited specifically.
[0128] Thus, based on the above hybrid supervision training strategy, the reconstruction network is trained by 2D images and 3D face models, so that the reconstruction network can fully learn facial aesthetic information and identity information, and the above various loss functions are used for hybrid supervision in the three-dimensional reconstruction process, and finally a high-quality three-dimensional face reconstruction model with high precision, high aesthetic consistency and high robustness is realized, which is used to obtain three-dimensional face data.
[0129] Step S500, if the above reconstruction loss does not satisfy the preset reconstruction termination condition, return to execute the above steps of obtaining the facial projection image based on the initial three-dimensional face model and the second facial key point information corresponding to the facial projection image until the reconstruction loss satisfies the preset reconstruction termination condition, so as to obtain the target three-dimensional face model matched with the target object after reconstruction.
[0130] wherein, the preset reconstruction termination condition can be preset and adjusted according to actual needs, for example, it can be set to be less than the preset loss threshold, and / or the reconstruction number reaches the preset reconstruction number threshold, and other conditions can also be set, which are not limited specifically. After satisfying the preset reconstruction termination condition, the initial three-dimensional face model updated finally at this time is taken as the target three-dimensional face model.
[0131] Thus, in order to solve the problems of insufficient geometric accuracy, low aesthetic consistency and poor robustness in the prior art, the aesthetic information, identity information, perception information and 3D information supervision training process are introduced in the present application, and the face segmentation technology is combined to realize accurate capture and reconstruction of facial aesthetics, thereby enhancing the effect of reconstruction technology in geometric accuracy, aesthetic consistency and robustness. In the three-dimensional face reconstruction process, aesthetic supervision, identity supervision and 3D supervision are introduced, and deep neural network is used for end-to-end hybrid supervision training to realize three-dimensional model generation.
[0132] As can be seen from the above, in the three-dimensional face model reconstruction method provided in the embodiments of the present application, a two-dimensional face image corresponding to a target object and first facial key point information corresponding to the two-dimensional face image are obtained; an initial three-dimensional face model is obtained, a face projection image is obtained based on the initial three-dimensional face model, and second facial key point information corresponding to the face projection image is obtained; three-dimensional face model reconstruction is performed through a preset three-dimensional face model reconstruction network according to the two-dimensional face image, the first facial key point information, the face projection image, and the second facial key point information, so as to update the initial three-dimensional face model; a reconstruction loss is determined according to the updated initial three-dimensional face model and the two-dimensional face image, wherein the reconstruction loss includes a beauty loss determined based on a preset beauty score network; if the reconstruction loss does not satisfy a preset reconstruction termination condition, the step of obtaining the face projection image based on the initial three-dimensional face model and the second facial key point information corresponding to the face projection image is returned to be executed until the reconstruction loss satisfies the preset reconstruction termination condition, so as to obtain a target three-dimensional face model matched with the target object after reconstruction.
[0133] Compared with the prior art, in the scheme of the present application, three-dimensional face model reconstruction is performed through a three-dimensional face model reconstruction network based on the two-dimensional face image corresponding to the target object and the first facial key point information corresponding thereto, the initial three-dimensional face model, the face projection image corresponding to the initial three-dimensional face model, and the second facial key point information corresponding to the face projection image. In the reconstruction process, not only the geometric features in the image are considered, but also the beauty loss determined based on the preset beauty score network is taken as part of the reconstruction loss, so that, in the process of determining whether further reconstruction is needed based on the reconstruction loss, an aesthetic judgment is also made, which is beneficial to guarantee that the three-dimensional face model after reconstruction has consistent visual beauty with the original image.
[0134] In the embodiments of the present application, the three-dimensional face model reconstruction method is also described in detail based on a specific application scenario. Figure 2 is a specific flowchart of a three-dimensional face model reconstruction method provided in the embodiments of the present application, as shown in Figure 2 Face segmentation and key point detection are performed on the two-dimensional face image, projection is performed on the initial three-dimensional face model to obtain a multi-view face projection image, and then joint training and three-dimensional face model reconstruction are performed based on a three-dimensional face model reconstruction network and a hybrid supervision model loss function, so as to obtain a final target three-dimensional face model. Figure 3 is a data processing schematic diagram of a three-dimensional face model reconstruction network provided in the embodiments of the present application, as shown in Figure 3As shown, the three-dimensional face model reconstruction network reconstructs a three-dimensional face model based on 2D supervised input and 3D supervised input, and determines the corresponding 2D supervision loss and 3D supervision loss based on the reconstructed 3D face output by the network, so as to determine whether the next round of three-dimensional face model reconstruction is needed.
[0135] In the embodiments of the present application, by introducing the aesthetic information supervision loss, the aesthetic features of the original image can be effectively captured and reconstructed, ensuring that the generated three-dimensional model is consistent with the original image in terms of facial aesthetics. The application of face segmentation technology makes the method of the present application more robust in handling complex scenes (such as face occlusion, pose change, etc.), ensuring that the obtained three-dimensional model maintains high quality in complex scenes. At the same time, the training method of mixed supervision of 2D images and 3D models fully learns the mapping relationship between face images and three-dimensional faces, has high reconstruction accuracy, and can meet the needs of most application scenarios. Using a large amount of two-dimensional image data and a small amount of three-dimensional data for mixed training, a large amount of two-dimensional images are used to extract facial features and reconstruct them, so that the model has better generalization. Using a small amount of three-dimensional data for auxiliary supervision can make full use of high-quality three-dimensional face data in the case of data scarcity, so that the model can learn three-dimensional face features more accurately and achieve high-quality three-dimensional reconstruction. This way of mixed supervision to improve reconstruction quality is conducive to improving the applicability of the method of the present application in practical applications.
[0136] Further, the three-dimensional face model reconstruction method in the embodiments of the present application is also tested on multiple public data sets, and significant results are achieved on the SCUT-FBP5500 data set and the NOW benchmark test set. The experimental results show that compared with traditional methods, the method of the present application has significantly improved in terms of aesthetic consistency and geometric accuracy. The test results on the SCUT-FBP5500 data set show that the method of the present application achieves high accuracy in facial aesthetic scoring, and the generated three-dimensional model has good aesthetic consistency. The results on the NOW benchmark test set show that the method of the present application can maintain high robustness and consistency when handling complex scenes (such as face occlusion, expression change).
[0137] It should be noted that the beauty score network used in the embodiments of the present application can be replaced or adjusted according to actual needs to adapt to different requirements. For example, when reconstructing a three-dimensional face model for a certain age group, a beauty score network corresponding to the age group can be used to improve the applicability of the reconstruction result.
[0138] Meanwhile, in addition to using a single two-dimensional face image as input, multi-modal data (such as multi-view images, video sequences, or depth maps) can be introduced to further improve the accuracy and aesthetic consistency of three-dimensional reconstruction. The fusion of multi-modal data can provide more rich information, thus generating more accurate and aesthetically pleasing three-dimensional models.
[0139] By optimizing algorithms and utilizing hardware acceleration techniques, the method of the present application can be applied to real-time three-dimensional reconstruction scenarios, such as virtual makeup, real-time beautification, etc., expanding its application range.
[0140] Further, the face segmentation technique can be further optimized to handle more diverse facial expressions and poses, adapt to more complex lighting and background conditions, and improve the robustness and applicability of the model. For example, deep learning-based segmentation algorithms can be introduced to handle different regions more finely, thus improving the overall quality of the reconstructed model.
[0141] Exemplary device
[0142] As shown in Figure 4 Corresponding to the three-dimensional face model reconstruction method described above, the embodiments of the present application also provide a three-dimensional face model reconstruction system, which comprises:
[0143] The first data acquisition module 410 is configured to acquire a two-dimensional face image corresponding to a target object and first facial key point information corresponding to the two-dimensional face image;
[0144] The second data acquisition module 420 is configured to acquire an initial three-dimensional face model, a facial projection image corresponding to the initial three-dimensional face model, and second facial key point information corresponding to the facial projection image;
[0145] The three-dimensional model reconstruction module 430 is configured to perform three-dimensional face model reconstruction through a pre-set three-dimensional face model reconstruction network according to the two-dimensional face image, the first facial key point information, the facial projection image, and the second facial key point information, to update the initial three-dimensional face model;
[0146] The loss determination module 440 is configured to determine a reconstruction loss according to the updated initial three-dimensional face model and the two-dimensional face image, wherein the reconstruction loss comprises a beauty degree loss determined based on a pre-set beauty score network;
[0147] The reconstruction control module 450 is configured to return to triggering the second data acquisition module 420 to perform the steps of acquiring the facial projection image based on the initial three-dimensional face model and acquiring the second facial key point information corresponding to the facial projection image until the reconstruction loss meets the preset reconstruction termination condition, so as to obtain the target three-dimensional face model matched with the target object after reconstruction, if the reconstruction loss does not meet the preset reconstruction termination condition.
[0148] In this way, the three-dimensional face model reconstruction is performed by the three-dimensional face model reconstruction network based on the two-dimensional face image corresponding to the target object and the first facial key point information corresponding to the two-dimensional face image, the initial three-dimensional face model, the facial projection image corresponding to the initial three-dimensional face model, and the second facial key point information corresponding to the facial projection image. In the reconstruction process, the beauty loss determined by the preset beauty score network is taken as part of the reconstruction loss in addition to the geometric features in the image, so that the judgment from the aesthetic perspective is also made in the process of determining whether further reconstruction is needed based on the reconstruction loss, which is beneficial to guarantee that the three-dimensional face model after reconstruction has consistent visual beauty with the original image.
[0149] It should be noted that the specific structure and implementation manner of the three-dimensional face model reconstruction system and each module or unit thereof can be referred to the corresponding description in the above method embodiments, which will not be described here.
[0150] It should be noted that the division manner of each module of the three-dimensional face model reconstruction system is not unique, and is not specifically limited here.
[0151] Based on the above embodiments, the application further provides an intelligent terminal, and a principle block diagram thereof can be as shown in Figure 5 The intelligent terminal includes a processor, a memory, a network interface, and a display screen connected through a system bus. The processor of the intelligent terminal is configured to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a three-dimensional face model reconstruction program. The internal memory provides an environment for the operating system and the three-dimensional face model reconstruction program in the non-volatile storage medium. The network interface of the intelligent terminal is configured to communicate with an external terminal through a network connection. The three-dimensional face model reconstruction program is executed by the processor to implement the steps of any one of the three-dimensional face model reconstruction methods. The display screen of the intelligent terminal can be a liquid crystal display screen or an electronic ink display screen.
[0152] Those skilled in the art can understand that Figure 5The principle block diagram shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the intelligent terminal to which the scheme of the present application is applied. The specific intelligent terminal can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0153] In an embodiment, an intelligent terminal is provided, which includes a memory, a processor, and a three-dimensional face model reconstruction program stored in the memory and executable on the processor. The three-dimensional face model reconstruction program, when executed by the processor, implements the steps of any three-dimensional face model reconstruction method provided in the embodiments of the present application.
[0154] The embodiments of the present application also provide a computer readable storage medium, which stores a three-dimensional face model reconstruction program. The three-dimensional face model reconstruction program, when executed by a processor, implements the steps of any three-dimensional face model reconstruction method provided in the embodiments of the present application.
[0155] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of the processes should be determined according to their functions and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0156] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is taken as an example for illustration. In actual application, the above functions can be completed by different functional units and modules according to needs, i.e., the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the device can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0157] In the above embodiments, the description of each embodiment has its own emphasis. The parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0158] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different ways to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0159] In the embodiments provided in the present application, it should be understood that the disclosed system / terminal device and method can be implemented in other ways. For example, the above-described system / terminal device embodiments are only schematic, and the division of the above modules or units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0160] The integrated modules / units described above, if realized in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, all or part of the flow of the above-described embodiment methods can also be completed by computer programs instructing related hardware, and the above computer programs can be stored in a computer readable storage medium. The computer programs are executed by the processor, and the steps of the above various method embodiments can be implemented. The computer programs include computer program codes, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium can include any entity or device capable of carrying the above computer program codes, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the computer readable storage medium contains content which can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.
[0161] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand; it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements, do not deviate from the spirit and scope of the corresponding technical solutions of the present application, and should be included in the protection scope of the present application.
Claims
1. A method for three-dimensional face model reconstruction, characterized in that, The method comprises: obtaining a two-dimensional face image corresponding to a target object and first facial key point information corresponding to the two-dimensional face image; obtaining an initial three-dimensional face model, obtaining a face projection image based on the initial three-dimensional face model, and obtaining second facial key point information corresponding to the face projection image; inputting the two-dimensional face image, the first facial key point information, the face projection image, and the second facial key point information into a three-dimensional face model reconstruction network to obtain a rendering control vector generated by the three-dimensional face model reconstruction network, wherein the rendering control vector is used to represent a camera pose, a lighting parameter, a face shape parameter, and a face texture parameter; performing three-dimensional face model reconstruction according to the rendering control vector through the three-dimensional face model reconstruction network, and taking a three-dimensional face model obtained by the reconstruction as an updated initial three-dimensional face model; determining a reconstruction loss according to the updated initial three-dimensional face model and the two-dimensional face image, wherein the reconstruction loss comprises a beauty loss determined based on a preset beauty score network; if the reconstruction loss does not satisfy a preset reconstruction termination condition, returning to perform the step of obtaining the face projection image based on the initial three-dimensional face model and the second facial key point information corresponding to the face projection image until the reconstruction loss satisfies the preset reconstruction termination condition, so as to obtain a target three-dimensional face model matched with the target object after the reconstruction is completed.
2. The three-dimensional face model reconstruction method of claim 1, wherein, The method comprises: obtaining an initial face image corresponding to the target object; preprocessing the initial face image according to a preset preprocessing operation to obtain a two-dimensional face image corresponding to the target object, wherein the preprocessing operation comprises at least one of standardization, cropping, and scaling; performing facial key point positioning on the two-dimensional face image through a preset facial key point positioning network to obtain first facial key point information corresponding to the two-dimensional face image.
3. The method of claim 1, wherein, The method comprises: obtaining an initial three-dimensional face model; performing projection on the initial three-dimensional face model for multiple different face poses to obtain face projection images corresponding to the multiple different face poses; performing facial key point positioning on the face projection images through a preset facial key point positioning network to obtain second facial key point information corresponding to the face projection images.
4. The method of claim 1, wherein, The three-dimensional face model reconstruction network further generates a rendering image corresponding to the updated initial three-dimensional face model; The method comprises: obtaining a first beauty score corresponding to the two-dimensional face image and a second beauty score corresponding to the rendering image through a pre-trained beauty score network; determining the beauty loss according to the first beauty score and the second beauty score.
5. The method of claim 4, wherein, The reconstruction loss further comprises an identity loss and a perception loss, and the perception loss comprises a pixel loss and a key point loss; The method further comprises: The first identity encoding vector corresponding to the two-dimensional face image and the second identity encoding vector corresponding to the rendered image are obtained through a pre-trained face recognition network, and the identity loss is determined according to the cosine distance between the first identity encoding vector and the second identity encoding vector; The pixel loss is determined according to the color error of the corresponding pixels between the two-dimensional face image and the rendered image; The third facial key point information corresponding to the rendered image is obtained, and the key point loss is determined according to the first facial key point information and the third facial key point information. In the calculation process of the key point loss, the calculation weight of the key points in the facial contour region and the facial feature region in the two-dimensional face image and the rendered image is higher than the calculation weight of the key points in other regions.
6. The method of claim 4, wherein, The reconstruction loss further comprises a three-dimensional supervision loss; The method further comprises: A standard three-dimensional face model is obtained; The standard three-dimensional face model is calibrated to make the pose of the standard three-dimensional face model consistent with that of the updated initial three-dimensional face model; The three-dimensional supervision loss is determined according to the Euclidean distance between the corresponding points of the standard three-dimensional face model and the updated initial three-dimensional face model.
7. A three-dimensional face model reconstruction system, characterized by, The system comprises: A first data acquisition module is configured to acquire a two-dimensional face image corresponding to a target object and first facial key point information corresponding to the two-dimensional face image; A second data acquisition module is configured to acquire an initial three-dimensional face model, a facial projection image corresponding to the initial three-dimensional face model, and second facial key point information corresponding to the facial projection image; A three-dimensional model reconstruction module is configured to input the two-dimensional face image, the first facial key point information, the facial projection image, and the second facial key point information into a three-dimensional face model reconstruction network to obtain a rendered control vector generated by the three-dimensional face model reconstruction network, wherein the rendered control vector is used to represent a camera pose, a lighting parameter, a facial shape parameter, and a facial texture parameter; the three-dimensional face model reconstruction network is used to perform three-dimensional face model reconstruction according to the rendered control vector, and a three-dimensional face model obtained through the reconstruction is used as an updated initial three-dimensional face model; A loss determination module is configured to determine a reconstruction loss according to the updated initial three-dimensional face model and the two-dimensional face image, wherein the reconstruction loss comprises a beauty degree loss determined based on a preset beauty score network. The reconstruction control module is configured to return and sequentially trigger the second data acquisition module, the three-dimensional model reconstruction module and the loss determination module to re-execute corresponding steps until the reconstruction loss meets the preset reconstruction termination condition, so as to obtain a target three-dimensional face model matched with the target object after reconstruction is completed.
8. A smart terminal, characterized by The intelligent terminal comprises a memory, a processor, and a three-dimensional face model reconstruction program stored in the memory and executable on the processor. When the three-dimensional face model reconstruction program is executed by the processor, the steps of the three-dimensional face model reconstruction method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a three-dimensional face model reconstruction program. When the three-dimensional face model reconstruction program is executed by the processor, the steps of the three-dimensional face model reconstruction method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Intelligent face beautifying method and system based on aesthetic guidance
CN115862111A
Reconstruction method and device of three-dimensional face model, storage medium and electronic equipment
CN117011449A