Method and device for reconstructing three-dimensional face model, storage medium and electronic device
Patent Information
- Application Number
- CN202210641294.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-08
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2042-06-08
AI Technical Summary
[0005]本发明实施例提供了一种三维面部模型的重构方法和装置、存储介质及电子设备,以至少解决现有三维面部模型的重构方法得到的面部模型的准确性较低的技术问题
[0010]根据本发明实施例的又一方面,还提供了一种电子设备,包括存储器和处理器,上述存储器中存储有计算机程序,上述处理器被设置为通过所述计算机程序执行上述的三维面部模型的重构方法。
Smart Images

Figure CN117011449B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computers, and more specifically, to a method and apparatus for reconstructing a three-dimensional facial model, a storage medium, and an electronic device. Background Technology
[0002] In many applications today, to enrich the user experience, 3D facial models can be reconstructed from real user facial images. For example, in the metaverse, the reconstructed 3D facial model can be used to generate a controllable head for the metaverse; and in some 3D games, facial reconstruction can be used to obtain the face of a virtual game character that closely resembles the facial features of a real user.
[0003] The most common reconstruction method currently is to acquire a user's facial image, extract key point features from the facial image, and generate corresponding 3D facial model parameters based on these key point features. Then, according to the definition of the 3D facial model parameters, the aforementioned real facial features are transferred to the face of the corresponding virtual character. This 3D facial model reconstruction method is limited by fixed parameter definition standards, making it difficult to achieve refined expressions in the transferred virtual character, thus resulting in low accuracy of the 3D facial model reconstruction results.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This invention provides a method, apparatus, storage medium, and electronic device for reconstructing a three-dimensional facial model, to at least solve the technical problem that the facial model obtained by existing three-dimensional facial model reconstruction methods has low accuracy.
[0006] According to one aspect of the present invention, a method for reconstructing a three-dimensional facial model is provided, comprising: acquiring a target facial image of a target object; extracting features from the target facial image to obtain expression features of the target object; determining three-dimensional reconstruction parameters of the target facial image using an expression representation vector matching the expression features and a reconstruction reference representation vector corresponding to the target facial image, wherein the reconstruction reference representation vector includes: a pixel representation vector matching the pixel features of the target facial image and a facial representation vector matching the facial features of the target object; and reconstructing the target facial image in three dimensions according to the three-dimensional reconstruction parameters to obtain a three-dimensional facial model matching the target object.
[0007] According to another aspect of the present invention, a three-dimensional facial model reconstruction apparatus is also provided, comprising: an acquisition unit for acquiring a target facial image of a target object; an extraction unit for extracting features from the target facial image to obtain expression features of the target object; a determination unit for determining three-dimensional reconstruction parameters of the target facial image using an expression representation vector matching the expression features and a reconstruction reference representation vector corresponding to the target facial image, wherein the reconstruction reference representation vector includes: a pixel representation vector matching the pixel features of the target facial image and a facial representation vector matching the facial features of the target object; and a reconstruction unit for performing three-dimensional reconstruction of the target facial image according to the three-dimensional reconstruction parameters to obtain a three-dimensional facial model matching the target object.
[0008] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to execute the above-described method for reconstructing a three-dimensional facial model at runtime.
[0009] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the three-dimensional facial model reconstruction method as described above.
[0010] According to another aspect of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to execute the above-described method for reconstructing a three-dimensional facial model through the computer program.
[0011] In this embodiment of the invention, a target facial image of a target object is acquired; features are extracted from the target facial image to obtain the expression features of the target object; using an expression representation vector matching the expression features and a reconstruction reference representation vector corresponding to the target facial image, the three-dimensional reconstruction parameters of the target facial image are determined. The reconstruction reference representation vector includes: a pixel representation vector matching the pixel features of the target facial image and a facial representation vector matching the facial features of the target object; then, the target facial image is reconstructed in three dimensions according to the three-dimensional reconstruction parameters to obtain a three-dimensional facial model matching the target object. In this embodiment, by determining the three-dimensional reconstruction parameters based on the expression representation vector and the reconstruction reference representation vector of the target object, the expression features of the target object are emphasized during the acquisition of the three-dimensional reconstruction parameters, thereby improving the accuracy of the generated three-dimensional facial model in terms of expression features and solving the technical problem of low accuracy of facial models obtained by existing three-dimensional facial model reconstruction methods. Attached Figure Description
[0012] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the hardware environment for an optional three-dimensional facial model reconstruction method according to an embodiment of the present invention; Figure 2 This is a flowchart of an optional method for reconstructing a three-dimensional facial model according to an embodiment of the present invention; Figure 3 This is a schematic diagram of an optional three-dimensional facial model reconstruction method according to an embodiment of the present invention; Figure 4 This is a schematic diagram of another optional method for reconstructing a three-dimensional facial model according to an embodiment of the present invention; Figure 5 This is a schematic diagram of another optional method for reconstructing a three-dimensional facial model according to an embodiment of the present invention; Figure 6 This is a schematic diagram of another optional method for reconstructing a three-dimensional facial model according to an embodiment of the present invention; Figure 7 This is a schematic diagram of another optional method for reconstructing a three-dimensional facial model according to an embodiment of the present invention; Figure 8 This is a flowchart of another optional method for reconstructing a three-dimensional facial model according to an embodiment of the present invention; Figure 9 This is a schematic diagram of an optional three-dimensional facial model reconstruction device according to an embodiment of the present invention; Figure 10 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present invention. Detailed Implementation
[0013] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0014] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0015] The following explains the terminology used in this application: Deep learning is a branch of machine learning that is based on neural network architecture and learns representations of data. It is divided into unsupervised, semi-supervised, and fully supervised learning and has been widely used in fields such as computer vision, speech recognition, and natural language processing. 3D face reconstruction: refers to the task of reconstructing a 3D face from a single image or multiple images.
[0016] According to one aspect of the present invention, a method for reconstructing a three-dimensional facial model is provided. As an optional implementation, the above-described method for reconstructing a three-dimensional facial model can be applied, but is not limited to, to applications such as... Figure 1 The system shown is a 3D facial model reconstruction system consisting of server 102 and terminal device 104. (As shown...) Figure 1As shown, server 102 is connected to terminal device 104 via network 110. This network may include, but is not limited to, wired networks and wireless networks. The wired network includes local area networks (LANs), metropolitan area networks (MANs), and wide area networks (WANs). The wireless network includes Bluetooth, Wi-Fi, and other networks enabling wireless communication. The terminal device may include, but is not limited to, at least one of the following: mobile phones (such as Android phones, iOS phones, etc.), laptops, tablets, PDAs, MIDs (Mobile Internet Devices), PADs, desktop computers, smart TVs, in-vehicle devices, etc. The terminal device may have a client installed, such as a 3D facial model generation client or a game client. The terminal device also includes a display, a processor, and a memory. The display can be used to show the program interface of the 3D facial model generation client or the game client, as well as the facial expression image of the target object uploaded to the server. The processor can be used to preprocess the video file to be uploaded before transmission, for example, by compressing the acquired image file. The memory is used to store the image file to be uploaded. It is understood that after obtaining the facial image of the target object to be uploaded in the aforementioned terminal device 104, the terminal device 104 can send the facial image to the server 102 via network 110. Upon receiving the facial image, the server 102 generates 3D reconstruction parameters matching the target object based on the facial image uploaded by the terminal device 104, and generates a corresponding 3D facial model based on the target 3D reconstruction parameters. The terminal device 104 can receive the 3D facial model returned by the server 102 via network 110. The server 102 can be a single server, a server cluster consisting of multiple servers, or a cloud server. The aforementioned server includes a database and a processing engine. The database may include a basic facial model used to reconstruct the 3D facial model for the user object; the processing engine is used to reconstruct the corresponding 3D facial model using the 3D reconstruction parameters.
[0017] According to one aspect of the present invention, the above-described three-dimensional facial model reconstruction system may further perform the following steps: The terminal device 104 performs step S102 to acquire a facial image of the target object; then, in step S104, the terminal device 104 sends the facial image to the server 102 via network 110; the server 102 performs steps S106 to S112 to acquire a target facial image of the target object; features are extracted from the target facial image to obtain the expression features of the target object; using an expression representation vector matching the expression features and a reconstruction reference representation vector corresponding to the target facial image, the three-dimensional reconstruction parameters of the target facial image are determined, wherein the reconstruction reference representation vector includes: a pixel representation vector matching the pixel features of the target facial image and a facial representation vector matching the facial features of the target object; the target facial image is reconstructed in three dimensions according to the three-dimensional reconstruction parameters to obtain a three-dimensional facial model matching the target object; then, in step S114, the server 102 sends the three-dimensional facial model to the terminal device 104 via network 110; finally, in step S116, the three-dimensional facial model can be displayed on the terminal device 104. It is understandable that if the terminal device 104 is a device with sufficient computing power, the above steps S106 to S112 can also be performed in the terminal device 104.
[0018] In this embodiment of the invention, a target facial image of a target object is acquired; features are extracted from the target facial image to obtain the expression features of the target object; using an expression representation vector matching the expression features and a reconstruction reference representation vector corresponding to the target facial image, the three-dimensional reconstruction parameters of the target facial image are determined. The reconstruction reference representation vector includes: a pixel representation vector matching the pixel features of the target facial image and a facial representation vector matching the facial features of the target object; then, the target facial image is reconstructed in three dimensions according to the three-dimensional reconstruction parameters to obtain a three-dimensional facial model matching the target object. In this embodiment, by determining the three-dimensional reconstruction parameters based on the expression representation vector and the reconstruction reference representation vector of the target object, the expression features of the target object are emphasized during the acquisition of the three-dimensional reconstruction parameters, thereby improving the accuracy of the generated three-dimensional facial model in terms of expression features and solving the technical problem of low accuracy of facial models obtained by existing three-dimensional facial model reconstruction methods.
[0019] The above is merely an example, and no limitation is made in this embodiment.
[0020] As an optional implementation method, such as Figure 2 As shown, the above-mentioned method for reconstructing a 3D facial model includes the following steps: S202, Obtain the target facial image of the target object; It is understandable that the method for obtaining the facial image of the target object can be either to obtain an image file containing the target object's face that is already stored in the mobile terminal, or to call the camera function of the mobile terminal to obtain the facial image of the target object in real time. The obtained facial image can be one or multiple images.
[0021] As an alternative, the target can be instructed to activate the terminal camera to obtain a facial image via text or voice prompts on the terminal interface, or to upload a facial image via touch operation. When instructing the target to activate the terminal camera to obtain a facial image via text or voice prompts on the terminal interface, the target can be further instructed to make different facial expressions via voice or text to extract facial expression features from the image. For example, voice prompts such as "Please open your mouth as wide as possible," "Please open your eyes as wide as possible," "Please close your eyes," and "Please laugh" can be played to instruct the target to make different expressions.
[0022] S204, Extract features from the target facial image to obtain the facial expression features of the target object; The following explains the aforementioned facial expression features. Existing methods typically utilize multiple images of the same object for regression learning to obtain reconstruction parameters for generating a 3D facial model. For example, a reconstructed image is obtained based on a single facial image of the same object, and regression learning is then performed on this facial image and the reconstructed image. If multiple facial images of the same object exist, this learning process is repeated multiple times. However, although different images of the same person may contain different expressions, the reconstructed images of different expressions usually exhibit strong correlations. That is, the facial expression features between the reconstructed images of different facial expressions of the same object also contain relevant features that indicate the object. By analyzing the facial expression features between different images of the same object, more accurate facial features of the same object can be obtained, leading to a more accurate 3D facial model.
[0023] In one optional function of facial expression feature indication, the aforementioned facial expression features can be used to indicate fine eye features of the same object. For example, if the target object is in a closed-eye state in an input facial image, the closed-eye state of the target object can be accurately reconstructed in a 3D facial model based on eye expression features.
[0024] S206, using the expression representation vector that matches the expression features and the reconstruction reference representation vector corresponding to the target facial image, the three-dimensional reconstruction parameters of the target facial image are determined. The reconstruction reference representation vector includes: a pixel representation vector that matches the pixel features of the target facial image and a facial representation vector that matches the facial features of the target object.
[0025] Understandably, after obtaining the expression representation vector matching the facial features, it is also necessary to obtain the reconstructed reference expression vector corresponding to the facial image, thereby accurately determining the reconstruction parameters used for 3D facial model reconstruction. As an optional approach, the reconstructed reference expression vector may include pixel representation vectors indicating the fine similarity between the original and reconstructed images, facial contour representation vectors indicating the perceptual hierarchical similarity between the original and reconstructed images, and facial keypoint representation vectors indicating the keypoint position features between the original and reconstructed images. It is understood that the aforementioned facial representation vectors may include the aforementioned facial contour representation vectors and facial keypoint representation vectors.
[0026] S208, Perform three-dimensional reconstruction on the target facial image according to the three-dimensional reconstruction parameters to obtain a three-dimensional facial model that matches the target object.
[0027] The following explains the aforementioned 3D reconstruction parameters. In this embodiment, these 3D reconstruction parameters can be used to linearly combine with an average 3D facial model to obtain a 3D facial model that matches the target object. For example, the aforementioned 3D reconstruction parameters may include identity (…). ),expression( ), texture ( ),attitude( ),illumination( Five types of reconstruction parameters are linearly combined with an average 3D facial model to obtain the reconstructed target 3D facial model.
[0028] In this embodiment of the invention, a target facial image of a target object is acquired; features are extracted from the target facial image to obtain the expression features of the target object; using an expression representation vector matching the expression features and a reconstruction reference representation vector corresponding to the target facial image, the three-dimensional reconstruction parameters of the target facial image are determined. The reconstruction reference representation vector includes: a pixel representation vector matching the pixel features of the target facial image and a facial representation vector matching the facial features of the target object; then, the target facial image is reconstructed in three dimensions according to the three-dimensional reconstruction parameters to obtain a three-dimensional facial model matching the target object. In this embodiment, by determining the three-dimensional reconstruction parameters based on the expression representation vector and the reconstruction reference representation vector of the target object, the expression features of the target object are emphasized during the acquisition of the three-dimensional reconstruction parameters, thereby improving the accuracy of the generated three-dimensional facial model in terms of expression features and solving the technical problem of low accuracy of facial models obtained by existing three-dimensional facial model reconstruction methods.
[0029] As an optional implementation, the above-mentioned method of determining the three-dimensional reconstruction parameters of the target facial image using an expression representation vector that matches the expression features and a reconstruction reference representation vector corresponding to the target facial image includes: S1, In the 3D reconstruction parameter prediction network, the expression representation vector, pixel representation vector, and facial contour representation vector and facial key point representation vector in the facial representation vector are fused to obtain a multi-mode representation vector. The 3D reconstruction parameter prediction network is a deep neural network obtained by deep learning the sample pixel features of the sample facial image, the sample facial features of the sample object in the sample facial image, and the sample expression features. The sample facial image includes different expression images that match the same sample object. S2 predicts the 3D reconstruction parameters of the target facial image based on the multimodal representation vector.
[0030] It is understood that the sample images used for training the deep neural network can include different facial expression images matching the same sample object, or multiple different facial expression images of multiple different sample objects, in order to improve the learning effect of the deep learning process.
[0031] As an alternative approach, assume that the above facial expression representation vector can be represented as The pixel representation vector mentioned above can be expressed as: The facial contour representation vector in the above facial representation vectors can be represented as follows: The above facial key point representation vector can be expressed as: The above fusion operation can be performed by obtaining the above facial expression representation vector. Pixel representation vector Facial contour representation vector and key point representation vectors sum vector In another alternative approach, the aforementioned multimodal representation vector can also be a weighted sum of the multiple representation vectors, i.e.:
[0032] The above , , as well as These are the weight coefficients corresponding to the aforementioned pixel representation vector, expression representation vector, facial contour representation vector, and keypoint representation vector, respectively. In a preferred embodiment, the aforementioned... , , , .
[0033] In one alternative approach, the above , , as well as The multi-mode representation vector can be used to represent the loss function generated during feature extraction in the training process. It can also represent the overall loss function obtained during deep learning training by combining multiple loss functions. Understandably, in one optional approach, if the value of the overall loss function reaches a target threshold, it indicates that the performance of the parameter prediction network trained using the overall loss function meets the requirements, and it can be used to reconstruct a 3D facial model based on the facial image of the target object.
[0034] Through the above-described embodiments of this application, a multi-modal representation vector is obtained by fusing the expression representation vector, the pixel representation vector, and the facial contour representation vector and facial key point representation vector in the facial representation vector in a 3D reconstruction parameter prediction network. Based on the multi-modal representation vector, the 3D reconstruction parameters of the target facial image are predicted. Thus, based on the expression representation vector used to represent the expression features of the target object and the representation vectors of multiple other feature dimensions, the multi-modal representation vector is used to predict the 3D reconstruction parameters, thereby making the 3D facial model obtained based on the above-described 3D reconstruction parameters closer to the facial image of the target object, solving the technical problem of low accuracy of facial models obtained by existing reconstruction methods.
[0035] As an optional approach, the above process, before acquiring the target facial image of the target object, also includes: S1, acquire multiple sample facial images; S2, input multiple sample facial images into the initialized 3D reconstruction parameter prediction network for training until the convergence condition is reached. The convergence condition is used to indicate that multiple consecutive multi-mode loss values of the training output are less than the target threshold. The multi-mode loss value is obtained by weighted summation of pixel loss value, facial loss value, key point loss value and expression loss value.
[0036] In this embodiment, multiple sample facial images can be used to train the initialized 3D reconstruction parameter prediction network. It is understood that these multiple sample facial images include facial images of the same object with different expressions, in order to extract accurate facial expression features.
[0037] The following explains the specific methods for obtaining the aforementioned pixel loss values and facial loss values.
[0038] like Figure 3 As shown, first, the facial image The input is used to initialize the R-net network to obtain the initialized 3D reconstruction parameters. In this embodiment, the above-mentioned 3D reconstruction parameters may include identity ( ),expression( ), texture ( ),attitude( ),illumination( ) and pupil ( Six types of 3D reconstruction parameters are used, and differentiable rendering is performed based on these six types of 3D reconstruction to obtain the reconstructed image. As an alternative approach, the above-mentioned reconstructed image It can be based on the above identity ( ),expression( ), texture ( ),attitude( ),illumination( ) and pupil ( Six types of 3D reconstruction parameters are used to render a 3D facial model, which is then compared with the facial image. Projection mapping is performed along the same projection direction to obtain the reconstructed image corresponding to the reconstructed model. .
[0039] Simultaneously, a facial mask is obtained by processing the aforementioned facial images using a SEG network. Specifically, to make the network more robust to face occlusion and other facial changes such as beards or heavy makeup, the aforementioned seg network can be a simple Bayesian classifier based on a Gaussian mixture model, classifying each pixel... Predict the probability of a skin color Generate the above facial mask The method can be as follows:
[0040] It should be noted that the above method of generating facial masks can be considered a binary classification problem. Each pixel of the input image is classified into two categories (background or person), thus extracting the person from the background. For example... Figure 3 In the diagram, the white area represents the region that needs to be learned and reconstructed.
[0041] After obtaining the above facial image Reconstructing the image and facial mask In the case of the above pixel loss value It can be obtained in the following ways:
[0042] in, Indicates pixel index, It is the projected face area. Represents facial images With reconstructed image The distance between corresponding pixels.
[0043] Facial loss value The calculation method is as follows:
[0044] As an optional approach, the above It can be used to identify images Deep feature encoding, This represents the inner product of the deep feature encoding results of the facial image and the reconstructed image. In this embodiment, by using signals from a pre-trained face recognition network as weak supervision, the facial feature distance between the source image and the reconstructed image is calculated, thereby further guiding the training process through loss at the perceptual level.
[0045] After obtaining the above pixel loss values and facial loss value Based on this, combined with facial expression loss value and key point loss value The multimodal loss value can be obtained. The multimodal loss value mentioned above can also be a weighted sum of the multiple loss values mentioned above, that is:
[0046] The above , , as well as These are the weighting coefficients corresponding to the pixel loss value, expression loss value, facial loss value, and keypoint loss value mentioned above. In a preferred embodiment, the above... , , , .
[0047] Through the above-described embodiments of this application, multiple sample facial images are obtained; these multiple sample facial images are input into an initialized 3D reconstruction parameter prediction network for training until convergence is achieved, thereby determining the final value based on the total loss. The initial 3D reconstruction parameter prediction network is trained to improve the accuracy of the 3D reconstruction parameters output by the network.
[0048] As an alternative approach, training the 3D reconstruction parameter prediction network initialized with multiple sample facial images includes: Extract a subset of facial expression images corresponding to the same sample object from multiple sample facial images. This subset includes different facial expression images of the sample object. Then, sequentially use each subset of facial expression images as the current subset of facial expression images and perform the following operations: S1, Obtain the first expression image and the second expression image from the current subset of expression images; S2, replace the first expression parameter of the first expression image with the second expression parameter of the second expression image to obtain the first reference expression image, and obtain the first distance between the first expression image and the first reference expression image; S3, replace the second expression parameter of the second expression image with the first expression parameter of the first expression image to obtain the second reference expression image, and obtain the second distance between the second expression image and the second reference expression image; S4, based on the first distance and the second distance, determines the expression loss value.
[0049] It should be noted that obtaining the first and second expression images from the current subset of expression images can mean obtaining two or more images of different expressions of the same object. Different expressions can be represented by different facial angles or by different details of facial features.
[0050] In this embodiment, an inter-image perceptual loss function constraint is applied to the reconstruction of different facial expressions of the same person, thereby improving the facial expression fitting ability of the reconstruction model. Specifically, as follows... Figure 4 As shown, for the two portrait images of the object with ID A ( , The corresponding reconstruction parameters obtained from the output of the initialized R-net network ( , And the facial expression parameters of these two images. Replacement, after rendering, forms a... emoticons new Image (texture, pose, lighting and) Consistent), regarding the new emoticons created by fan artists. and At the sensory level, a face recognition loss function is used for constraint, and the expression loss value is calculated as follows:
[0051] It is understandable that the above This represents the dot product of vectors.
[0052] In this embodiment, The function can use ArcFace as the face recognition network. The main ideas of the ArcFace face recognition network include: ArcFace loss: Additive Angular Margin Loss, which normalizes the feature vector and weights, adding an angular margin m to θ. The angular margin has a more direct impact on the angle than the cosine margin. Geometrically, there is a constant linear angle margin; ArcFace directly maximizes the classification boundary in the angle space θ, while CosFace maximizes the classification boundary in the cosine space cos(θ); Preprocessing (face alignment): Facial key points are detected by MTCNN, and then the cropped aligned face is obtained through similarity transformation; Training (face classifier): ResNet50 + ArcFace loss; Testing: 512-dimensional embedded features are extracted from the output of the FC1 layer of the face classifier, the cosine distance between the two input features is calculated, and then face verification and face recognition are performed; In the actual code, training is divided into ResNetModel + ArcHead + Softmax loss. The ResNet model outputs features; ArcHead adds an angular interval between the features and weights, and then outputs the predicted label, which is used to calculate the ACC; Softmax loss calculates the error between the predicted label and the actual label.
[0053] Through the above-described embodiments of this application, a first expression image and a second expression image are obtained from the current subset of expression images; the first expression parameter of the first expression image is replaced with the second expression parameter of the second expression image to obtain a first reference expression image, and a first distance between the first expression image and the first reference expression image is obtained; the second expression parameter of the second expression image is replaced with the first expression parameter of the first expression image to obtain a second reference expression image, and a second distance between the second expression image and the second reference expression image is obtained; based on the first distance and the second distance, the expression loss value is determined, thereby employing a multi-image reconstruction method, further determining the relationship between different expressions of the same object, applying perceptual loss function constraints to the image reconstruction graphs of different expressions of the same person, improving the expression fitting ability of the reconstruction model, and solving the technical problem of low reconstruction accuracy in existing facial model reconstruction methods.
[0054] As an alternative approach, training the 3D reconstruction parameter prediction network initialized with multiple sample facial images includes: Use multiple sample facial images sequentially as the current sample facial image, and perform the following operations: S1, determine the location of facial key points of the sample object in the current sample facial image, wherein the location of facial key points includes the location of the pupils of the sample object in the sample facial image; S2, based on the position of each facial key point, the corresponding reconstruction reference key point position is predicted; S3, based on the location of facial key points and the location of reconstruction reference key points, determines the key point loss value.
[0055] In one alternative approach, the aforementioned set of facial key points can be a two-dimensional point set, treating the target object's face as a plane to determine the position coordinates of each point on that plane. For example, assuming the area containing the target object's face is 500px... A rectangle of 800px is defined with the left pupil's two-dimensional coordinates at (200px, 600px), and the right pupil's coordinates at (400px, 600px). Based on this method, the coordinates of each key point on the target object's face are determined, thus extracting the target object's facial key point set.
[0056] In another alternative approach, the aforementioned set of facial key points can be a three-dimensional point set, that is, reconstructing the facial spatial structure of the target object using the three-dimensional coordinates of different points. By determining the three-dimensional spatial coordinates of each key point on the target object's face, the set of facial key points of the target object is extracted.
[0057] As a specific method, such as Figure 5 As shown, the facial key points in this embodiment are defined using a 70-key-point definition. Figure 5 In the diagram, points 1 to 17 indicate the location of the facial contours; points 18 to 22 indicate the location of the left eyebrow; points 23 to 27 indicate the location of the right eyebrow; points 28 to 31 indicate the location of the bridge of the nose; points 32 to 36 indicate the location of the tip of the nose; points 37 to 42 and points 43 to 48 indicate the location of the eyes; points 49 to 68 indicate the location of the mouth; key point 69 indicates the location of the left pupil; and key point 70 indicates the location of the right pupil.
[0058] As an optional approach, the determination of keypoint loss values based on facial keypoint locations and reconstruction reference keypoint locations includes: S1, obtain the first coordinates corresponding to the positions of each facial key point, and obtain the first coordinate set; S2, obtain the second coordinates corresponding to the positions of each reconstruction reference key point, and obtain the second coordinate set; S3, determine the mean square variance of the first coordinate set and the second coordinate set as the key point loss value.
[0059] Assuming the first 68 key points of the target object's facial image are used It indicates that the key point of the left pupil is... It indicates that the key points of the right pupil are used The calculation method for the above keypoint loss values is as follows:
[0060] It is understandable that the above as well as The location parameters of 68 key points on the facial image of the target object and the pupil key points are determined. as well as The location parameters of 68 key points and pupil key points were obtained by projecting the 3D key points of the reconstructed face into the image space. and These are the weighting coefficients corresponding to other facial key points and pupil key points, respectively.
[0061] Understandably, the positional relationship between the pupil keypoint and the periorbital keypoints can reveal the details of the target object's eyes. For example, the positional relationship between keypoint 69 and points 37 to 42 can reveal details such as the openness and closing of the eyes. Furthermore, because the details of the target's eyes can further accurately display the target object's emotional characteristics, using the pupil keypoint can accurately express the target object's emotional characteristics in the reconstructed 3D model. Figure 6 As shown, the upward arrows indicate the prediction results of the prediction network trained without using pupil keypoints. It can be seen that the reconstructed image cannot accurately reconstruct the closed-eye features because pupil keypoints were not used for training. The downward arrows indicate the prediction results of the prediction network trained with pupil keypoints. It can be seen that because pupil keypoint features were used for training constraints, the reconstructed image can well show the feature that the pupil keypoints are on the same straight line as other periorbital keypoints, that is, it can accurately show the closed-eye features of the target object.
[0062] Through the above-described embodiments of this application, the facial key point positions of a sample object in a current sample facial image are determined, wherein the facial key point positions include the pupil positions of the sample object in the sample facial image; the corresponding reconstruction reference key point positions are predicted based on each facial key point position; and the key point loss value is determined based on the facial key point positions and the reconstruction reference key point positions, thereby constraining the facial reconstruction model with key point information including the pupil key position. Subsequently, the subtle facial expression features around the eyes can be reflected through the facial reconstruction model, thus solving the technical problem of low accuracy in existing facial reconstruction methods.
[0063] As an optional implementation, acquiring multiple sample facial images includes: S1, Obtain a set of real facial images; S2, Obtain a real face image from the real face image set as the current real face image, and perform the following operations: S3, obtain facial angle information of the current real facial image; S4, Obtain a reference occluder from the candidate occluder set; S5, Based on facial angle information and the type information of the reference occluder, determine the addition position information of the reference occluder on the current real facial image; S6, add a reference occluder to the current real face image at the position indicated by the added position information to obtain a sample face image corresponding to the current real face image.
[0064] It is understood that in this embodiment, by artificially constructing occluded images and using these artificially constructed occluded images to train the face reconstruction parameter prediction network, the prediction network's ability to reconstruct occluded face images can be improved. For example... Figure 7 As shown, the upward arrows indicate the reconstruction results of the prediction network trained without occlusion data. It is evident that the reconstructed image cannot reconstruct eye features, and severe shadows still exist in the eye area. The downward arrows indicate the reconstruction results of the prediction network trained with occlusion data. It is evident that the reconstructed image can well display eye features, i.e., the influence of occlusions on the reconstructed image of the target object is eliminated.
[0065] It is understandable that, in this embodiment, the artificially constructed occlusion data still needs to be reasonable to improve training results. When the reference occlusion includes a hand or sunglasses, the specific form and position of the occlusion can be determined based on the facial angle information of the acquired facial image. For example, if the acquired reference occlusion is sunglasses, and the facial angle of the acquired facial image is determined to be frontal, then both lenses of the sunglasses can be normally occluded to cover the eyes of the current real facial image; if the acquired reference occlusion is sunglasses, and the facial angle of the acquired facial image is determined to be sideways, then one lens of the sunglasses can be normally occluded to cover the eyes of the current real facial image from the side; if the acquired reference occlusion is a hand, the occlusion can be randomly applied to the face.
[0066] The present application employs the above-described embodiments to obtain facial angle information of a current real facial image; obtain a reference occluder from a set of candidate occluders; determine the addition position information of the reference occluder on the current real facial image based on the facial angle information and the type information of the reference occluder; and add the reference occluder to the current real facial image at the position indicated by the addition position information to obtain a sample facial image similar to the current real facial image. Furthermore, a manually constructed occluded image is used to train the prediction network. Because the occlusion is manually constructed, it possesses a ground truth for reconstruction, enabling supervised constraints on the reconstructed image. Ultimately, this allows the network to make reasonable predictions about the occluded region, thereby improving the accuracy of the reconstruction results of the facial model and solving the technical problem of low accuracy in facial models obtained by existing 3D facial model reconstruction methods.
[0067] The following combination Figure 8 A complete embodiment of this application will be described.
[0068] S802, construct occlusion data; Specifically, when the reference occlusion includes a hand or sunglasses, the specific form and location of the occlusion can be determined based on the facial angle information of the acquired facial image. For example, if the acquired reference occlusion is sunglasses, and the facial angle of the acquired facial image is determined to be frontal, then both lenses of the sunglasses can be normally occluded to cover the eyes of the current real facial image; if the acquired reference occlusion is sunglasses, and the facial angle of the acquired facial image is determined to be sideways, then one lens of the sunglasses can be normally occluded to cover the eyes of the current real facial image from the side; if the acquired reference occlusion is a hand, the occlusion can be randomly applied to the face. S804, uses occlusion data to train the initial network; During training, the following loss function is used to constrain the initial prediction network:
[0069] The above , , as well as These are the weighting coefficients corresponding to the aforementioned pixel loss values, expression loss values, facial contour loss values, and keypoint loss values, respectively. In a preferred embodiment, the aforementioned... , , , .
[0070] in:
[0071] Among them, after obtaining the above facial images Reconstructing the image and facial mask In this case, Indicates pixel index, It is the projected face area. Represents facial images With reconstructed image The distance between corresponding pixels.
[0072]
[0073] In the above formula, an inter-image perceptual loss function is applied to the reconstruction of different expressions of the same person, thereby improving the expression fitting ability of the reconstruction model. Specifically, as follows... Figure 4 As shown, for the two portrait images of the object with ID A ( , The corresponding reconstruction parameters obtained from the output of the initialized R-net network ( , And the facial expression parameters of these two images. Replacement, after rendering, forms a... emoticons new Image (texture, pose, lighting and) Consistent), regarding the new emoticons created by fan artists. and Constraints are applied at the sensory level using a facial recognition loss function.
[0074]
[0075] The above It can be used to identify images Deep feature encoding, This represents the inner product of the deep feature encoding results of the facial image and the reconstructed image. In this embodiment, by using signals from a pre-trained face recognition network as weak supervision, the facial feature distance between the source image and the reconstructed image is calculated, thereby further guiding the training process through loss at the perceptual level.
[0076]
[0077] It is understandable that the above as well as The location parameters of 68 key points on the facial image of the target object and the pupil key points are determined. as well as The location parameters of 68 key points and pupil key points were obtained by projecting the 3D key points of the reconstructed face into the image space. and These are the weighting coefficients corresponding to other facial key points and pupil key points, respectively.
[0078] S806, Obtain the facial image of the target object; S808 inputs the facial image into the trained prediction network; S810 generates a 3D facial model based on the input reconstruction parameters.
[0079] The aforementioned 3D reconstruction parameters can be used to linearly combine with an average 3D facial model to obtain a 3D facial model that matches the target object. For example, these 3D reconstruction parameters can include identity (…). ),expression( ), texture ( ),attitude( ),illumination( ), pupil point ( By linearly combining six types of reconstruction parameters with an average 3D facial model, the reconstructed target 3D facial model can be obtained.
[0080] Through the above-described embodiments of this application, a multi-modal representation vector is obtained by fusing the expression representation vector, the pixel representation vector, and the facial contour representation vector and facial key point representation vector in the facial representation vector in a 3D reconstruction parameter prediction network. Based on the multi-modal representation vector, the 3D reconstruction parameters of the target facial image are predicted. Thus, based on the expression representation vector used to represent the expression features of the target object and the representation vectors of multiple other feature dimensions, the multi-modal representation vector is used to predict the 3D reconstruction parameters, thereby making the 3D facial model obtained based on the above-described 3D reconstruction parameters closer to the facial image of the target object, solving the technical problem of low accuracy of facial models obtained by existing reconstruction methods.
[0081] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0082] According to another aspect of the present invention, a three-dimensional facial model reconstruction apparatus is also provided for implementing the above-described three-dimensional facial model reconstruction method. For example... Figure 9 As shown, the device includes: Acquisition unit 902 is used to acquire the target facial image of the target object; The extraction unit 904 is used to extract features from the target facial image to obtain the expression features of the target object; The determining unit 906 is used to determine the three-dimensional reconstruction parameters of the target facial image using an expression representation vector that matches the expression features and a reconstruction reference representation vector corresponding to the target facial image. The reconstruction reference representation vector includes: a pixel representation vector that matches the pixel features of the target facial image and a facial representation vector that matches the facial features of the target object. The reconstruction unit 908 is used to perform three-dimensional reconstruction of the target facial image according to the three-dimensional reconstruction parameters to obtain a three-dimensional facial model that matches the target object.
[0083] Optionally, in this embodiment, the implementation of each of the above-mentioned unit modules can be referred to the above-mentioned method embodiments, which will not be repeated here.
[0084] According to another aspect of the present invention, an electronic device for implementing the above-described method for reconstructing a three-dimensional facial model is also provided. This electronic device may be... Figure 10 The terminal device or server shown. This embodiment uses this electronic device as an example for illustration. Figure 10 As shown, the electronic device includes a memory 1002 and a processor 1004. The memory 1002 stores a computer program, and the processor 1004 is configured to execute the steps of any of the above method embodiments via the computer program.
[0085] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0086] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program: S1, Obtain the target facial image of the target object; S2, extract features from the target facial image to obtain the facial expression features of the target object; S3. Using the expression representation vector that matches the expression features and the reconstruction reference representation vector that corresponds to the target facial image, the three-dimensional reconstruction parameters of the target facial image are determined. The reconstruction reference representation vector includes: pixel representation vector that matches the pixel features of the target facial image and facial representation vector that matches the facial features of the target object. S4. Perform 3D reconstruction on the target facial image according to the 3D reconstruction parameters to obtain a 3D facial model that matches the target object.
[0087] Alternatively, as those skilled in the art will understand, Figure 10The structure shown is for illustrative purposes only. Electronic devices can also be in-vehicle terminals, smartphones (such as Android phones, iOS phones, etc.), tablets, handheld computers, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 10 This does not limit the structure of the aforementioned electronic devices. For example, the electronic device may also include components that are more... Figure 10 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 10 The different configurations shown.
[0088] The memory 1002 can be used to store software programs and modules, such as the program instructions / modules corresponding to the three-dimensional facial model reconstruction method and apparatus in this embodiment of the invention. The processor 1004 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002, thereby realizing the aforementioned three-dimensional facial model reconstruction method. The memory 1002 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1002 may further include memory remotely located relative to the processor 1004, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 1002 may be used, but is not limited to, to store information such as various elements in the viewing perspective image and the reconstruction information of the three-dimensional facial model. As an example, such as Figure 10 As shown, the memory 1002 may include, but is not limited to, the acquisition unit 902, extraction unit 904, determination unit 906, and reconstruction unit 910 in the three-dimensional facial model reconstruction device. Furthermore, it may include, but is not limited to, other module units in the three-dimensional facial model reconstruction device, which will not be elaborated upon in this example.
[0089] Optionally, the transmission device 1006 described above is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 1006 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 1006 is a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0090] In addition, the aforementioned electronic device also includes a display 1008 and a connection bus 1010 for connecting various module components in the aforementioned electronic device.
[0091] In other embodiments, the aforementioned terminal device or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer (P2P) network, and any form of computing device, such as a server, terminal, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.
[0092] According to one aspect of this application, a computer program product is provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions provided in embodiments of this application.
[0093] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0094] According to one aspect of this application, a computer-readable storage medium is provided, wherein a processor of a computer device reads computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the above-described method for reconstructing a three-dimensional facial model.
[0095] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps: S1, Obtain the target facial image of the target object; S2, extract features from the target facial image to obtain the facial expression features of the target object; S3. Using the expression representation vector that matches the expression features and the reconstruction reference representation vector that corresponds to the target facial image, the three-dimensional reconstruction parameters of the target facial image are determined. The reconstruction reference representation vector includes: pixel representation vector that matches the pixel features of the target facial image and facial representation vector that matches the facial features of the target object. S4. Perform 3D reconstruction on the target facial image according to the 3D reconstruction parameters to obtain a 3D facial model that matches the target object.
[0096] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0097] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0098] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0099] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.
[0100] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0101] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0102] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for reconstructing a three-dimensional facial model, characterized in that, include: Acquire multiple sample facial images; Training a 3D reconstruction parameter prediction network initialized with multiple sample facial images includes: obtaining a subset of expression images corresponding to the same sample object from the multiple sample facial images, wherein the subset of expression images includes different expression images of the sample object; sequentially using each subset of expression images as the current expression image subset, and performing the following operations: obtaining a first expression image and a second expression image from the current expression image subset; replacing the first expression parameter of the first expression image with the second expression parameter of the second expression image to obtain a first reference expression image, and obtaining a first distance between the first expression image and the first reference expression image; replacing the second expression parameter of the second expression image with the first expression parameter of the first expression image to obtain a second reference expression image, and obtaining a second distance between the second expression image and the second reference expression image; and determining an expression loss value based on the first distance and the second distance. If multiple consecutive multimodal loss values output during training are all less than the target threshold, the convergence condition is determined to be met. The multimodal loss value is obtained by weighted summation of pixel loss value, face loss value, key point loss value and expression loss value. Obtain the target facial image of the target object; Feature extraction is performed on the target facial image to obtain the expression features of the target object; Using the expression representation vector that matches the expression features and the reconstruction reference representation vector that corresponds to the target facial image, the three-dimensional reconstruction parameters of the target facial image are determined. The reconstruction reference representation vector includes: a pixel representation vector that matches the pixel features of the target facial image and a facial representation vector that matches the facial features of the target object. The target facial image is reconstructed in three dimensions according to the three-dimensional reconstruction parameters to obtain a three-dimensional facial model that matches the target object.
2. The method according to claim 1, characterized in that, The step of determining the three-dimensional reconstruction parameters of the target facial image by using an expression representation vector that matches the expression features and a reconstruction reference representation vector corresponding to the target facial image includes: In the 3D reconstruction parameter prediction network, the expression representation vector, the pixel representation vector, and the facial contour representation vector and facial key point representation vector in the facial representation vector are fused to obtain a multi-mode representation vector. The 3D reconstruction parameter prediction network is a deep neural network obtained by performing deep learning on the sample pixel features of the sample facial image, the sample facial features of the sample object in the sample facial image, and the sample expression features. The sample facial image includes different expression images that match the same sample object. The three-dimensional reconstruction parameters of the target facial image are predicted based on the multimodal representation vector.
3. The method according to claim 1, characterized in that, The step of training a 3D reconstruction parameter prediction network by inputting multiple sample facial images into the initialization process includes: The multiple sample facial images are sequentially used as the current sample facial image, and the following operations are performed: Determine the positions of facial key points of the sample object in the current sample facial image, wherein the positions of the facial key points include the positions of the pupils of the sample object in the sample facial image; The corresponding reconstruction reference key point positions are predicted based on the positions of each of the facial key points. The key point loss value is determined based on the location of the facial key points and the location of the reconstruction reference key points.
4. The method according to claim 3, characterized in that, The step of determining the keypoint loss value based on the facial keypoint location and the reconstruction reference keypoint location includes: Obtain the first coordinates corresponding to the positions of each of the facial key points to obtain the first coordinate set; Obtain the second coordinates corresponding to the positions of each of the reconstructed reference key points to obtain a set of second coordinates; The mean square variance of the first coordinate set and the second coordinate set is determined as the key point loss value.
5. The method according to claim 1, characterized in that, The acquisition of multiple sample facial images includes: Obtain a set of real facial images; Obtain a real face image from the set of real face images as the current real face image, and perform the following operations: Obtain the facial angle information of the current real facial image; Obtain a reference occluder from the set of candidate occluders; Based on the facial angle information and the type information of the reference occluder, the position information for adding the reference occluder on the current real facial image is determined; The reference occluder is added to the current real face image at the position indicated by the added position information to obtain a sample face image corresponding to the current real face image.
6. A device for reconstructing a three-dimensional facial model, characterized in that, include: The acquisition unit is used to acquire multiple sample facial images; Training a 3D reconstruction parameter prediction network initialized with multiple sample facial images includes: obtaining a subset of expression images corresponding to the same sample object from the multiple sample facial images, wherein the subset of expression images includes different expression images of the sample object; sequentially using each subset of expression images as the current expression image subset, and performing the following operations: obtaining a first expression image and a second expression image from the current expression image subset; replacing the first expression parameter of the first expression image with the second expression parameter of the second expression image to obtain a first reference expression image, and obtaining a first distance between the first expression image and the first reference expression image; replacing the second expression parameter of the second expression image with the first expression parameter of the first expression image to obtain a second reference expression image, and obtaining a second distance between the second expression image and the second reference expression image; and determining an expression loss value based on the first distance and the second distance. If multiple consecutive multimodal loss values output during training are all less than the target threshold, the convergence condition is determined to be met. The multimodal loss value is obtained by weighted summation of pixel loss value, face loss value, key point loss value and expression loss value. Obtain the target facial image of the target object; The extraction unit is used to extract features from the target facial image to obtain the expression features of the target object; The determining unit is used to determine the three-dimensional reconstruction parameters of the target facial image using an expression representation vector that matches the expression features and a reconstruction reference representation vector that corresponds to the target facial image. The reconstruction reference representation vector includes: a pixel representation vector that matches the pixel features of the target facial image and a facial representation vector that matches the facial features of the target object. The reconstruction unit is used to perform three-dimensional reconstruction of the target facial image according to the three-dimensional reconstruction parameters to obtain a three-dimensional facial model that matches the target object.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method described in any one of claims 1 to 5.
8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 5.
9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 5 through the computer program.