Facial modeling model training method and device, modeling method and device, electronic device, and computer program

The facial modeling model training method addresses inaccuracies in existing face reconstruction technologies by using feature encoding and neural radiance fields for 3D reconstruction, enabling decoupled control of facial attributes and improving the realism of 3D facial simulations.

JP2025535488APending Publication Date: 2025-10-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025523809
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-02-09
Filing Date
2023-11-08
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing face reconstruction technologies, whether voice-driven or text-driven, suffer from inaccuracies and limitations in generating realistic 3D facial simulations, with voice-driven methods producing varied results due to voice differences among individuals and text-driven methods only capable of generating 2D facial images.

Method used

A facial modeling model training method that utilizes feature encoding on sample face images from different viewpoints to obtain face region and action latent codes, combined with a neural radiance field for 3D reconstruction, allowing decoupling control of facial attributes and enabling long-sequence 3D facial modeling.

Benefits of technology

Enables realistic 3D facial modeling with improved simulation effectiveness by allowing individual control over different face regions and actions, enhancing the realism of facial animations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025535488000001_ABST
    Figure 2025535488000001_ABST
Patent Text Reader

Abstract

This application discloses a method and apparatus for training a facial modeling model, a modeling method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which relate to the field of artificial intelligence. The method includes the steps of: acquiring sample facial images, where the sample facial images include facial images of the same object at the same time and from different viewpoints; performing feature encoding on the sample facial images by an encoder of the facial modeling model to obtain facial region latent codes corresponding to different facial regions and facial action latent codes corresponding to facial actions; performing 3D facial reconstruction on the object based on the facial region latent codes and the facial action latent codes via a neural radiation field of the facial modeling model to obtain a 3D facial image of the object; obtaining a 3D reconstruction loss between the 3D facial image and the sample facial image; and training the facial modeling model based on the 3D reconstruction loss.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application is based on a Chinese patent application bearing application number 202310136271.9, filed with the China Patent Office on February 9, 2023, and claims priority to that Chinese patent application, the entire contents of which are incorporated herein by reference.

[0002] TECHNICAL FIELD Embodiments of the present application relate to the field of artificial intelligence, and in particular to a facial modeling model training method and apparatus, a modeling method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product. [Background technology]

[0003] In the field of computer vision, generating drivable faces has a wide range of applications, where the reconstruction of a face requires changing the viewpoint, facial expression, etc. based on input signals, and is required to have a natural speaking appearance and be synchronized with the driving signals.

[0004] Related technologies mainly use voice-driven and text-driven face reconstruction methods for face reconstruction. In the voice-driven face reconstruction method, significant differences exist between different people's voices, which can lead to different generated results from the trained network and inaccurate driving, thereby reducing the simulation effectiveness during face reconstruction. Meanwhile, text-driven face reconstruction methods can only generate two-dimensional facial images, resulting in reduced simulation effectiveness when driving three-dimensional faces. As such, the overall simulation effectiveness of related technologies during face reconstruction is low. Summary of the Invention [Problem to be solved by the invention]

[0005] Embodiments of the present application provide a facial modeling model training method and apparatus, a modeling method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which realize 3D facial modeling through facial decoupling control and improve the realism of the simulation. [Means for solving the problem]

[0006] An embodiment of the present application provides a method for training a face modeling model, the method comprising: acquiring sample face images, the sample face images including face images of the same object taken at the same time and from different viewpoints; performing feature encoding on the sample face images by an encoder of the face modeling model to obtain face region latent codes corresponding to different face regions and face action latent codes corresponding to different face actions; performing 3D face reconstruction on the object based on the face region latent code and the facial action latent code via a neural radiance field (NeRF) of the face modeling model to obtain a 3D face image of the object; obtaining a 3D reconstruction loss between the 3D face image and the sample face image, and training the face modeling model based on the 3D reconstruction loss.

[0007] An embodiment of the present application provides a method for modeling a three-dimensional face, the method comprising: when receiving input text, determining text phonemes corresponding to said input text; querying a target region latent code and a target action latent code corresponding to the text phoneme based on a phoneme latent code index, where the phoneme latent code index indicates a correspondence between a phoneme and a latent code sequence, and the latent code sequence is obtained by performing feature encoding on a face image corresponding to the phoneme based on an encoder of a face model; performing 3D reconstruction based on the target region latent code and the target action latent code through the neural radiation field of the face modeling model to obtain a facial action image corresponding to the text phoneme; and generating a three-dimensional facial movement animation corresponding to the input text based on the facial movement image of each frame.

[0008] An embodiment of the present application provides a facial modeling model training device, the device comprising: an image acquisition module configured to acquire sample facial images, the sample facial images including facial images of the same object at the same time but taken from different viewpoints; a feature encoding module configured to perform feature encoding on the sample face image by an encoder of the face modeling model to obtain face region latent codes corresponding to different face regions and face action latent codes corresponding to different facial actions; a 3D reconstruction module configured to perform 3D face reconstruction on the object based on the face region latent code and the facial action latent code via a neural radiation field of the face modeling model to obtain a 3D face image of the object; and a model training module configured to obtain a 3D reconstruction loss between the 3D face image and the sample face image, and train the face modeling model based on the 3D reconstruction loss.

[0009] An embodiment of the present application provides a 3D face modeling device, the device comprising: a phoneme determination module configured to, when receiving an input text, determine text phonemes corresponding to said input text; a latent code lookup module configured to look up a target region latent code and a target action latent code corresponding to the text phoneme based on a phoneme latent code index, wherein the phoneme latent code index indicates a correspondence between a phoneme and a latent code sequence, and the latent code sequence is obtained by performing feature encoding on a face image corresponding to the phoneme based on an encoder of a face model; a 3D reconstruction module configured to perform 3D reconstruction based on the target region latent code and the target action latent code through a neural radiation field of the face modeling model to obtain a facial action image corresponding to the text phoneme; and an animation generation module configured to generate a three-dimensional facial movement animation corresponding to the input text based on the facial movement image of each frame.

[0010] An embodiment of the present application provides an electronic device, the electronic device comprising: a memory for storing executable instructions; and a processor that executes computer-executable instructions stored in the memory to perform the facial modeling model training method described in the above aspect or the three-dimensional facial modeling method described in the above aspect.

[0011] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, cause the processor to perform the facial modeling model training method described in the above aspect or the three-dimensional face modeling method described in the above aspect.

[0012] An embodiment of the present application provides a computer program product including a computer program or computer-executable instructions stored in a computer-readable storage medium, wherein a processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium and executes the computer program or computer-executable instructions to cause the electronic device to perform the facial modeling model training method provided in the above embodiment or the 3D face modeling method described in the above embodiment. [Effects of the Invention]

[0013] The technical solutions according to the embodiments of the present application may include the following beneficial effects:

[0014] In an embodiment of the present application, feature coding is performed on sample face images from multiple viewpoints to obtain face region latent codes and face action latent codes corresponding to the sample face images, which are used to represent the region features and face action features of different face regions. Then, 3D reconstruction is performed using the face region latent codes and face action latent codes to obtain a 3D face image. A face modeling model capable of face decoupling control is obtained by training based on the difference between the 3D face image and the sample face image. When the model is used for 3D face modeling, different face regions and facial actions can be decoupling controlled individually, so that face images with combinations of different face regions and different facial actions can be generated. In the 3D face modeling process, by adjusting only the latent codes that need to be changed, a long sequence of 3D face modeling can be achieved, and the realism of face simulation in the text-driven face-driven process can be improved. [Brief explanation of the drawings]

[0015] [Figure 1] 1 shows a schematic diagram of an implementation environment according to one exemplary embodiment of the present application; [Figure 2] 1 shows a flowchart of a method for training a facial modeling model according to one exemplary embodiment of the present application. [Figure 3] 1 shows a flowchart of a method for determining 3D reconstruction loss according to another exemplary embodiment of the present application; [Figure 4] FIG. 1 illustrates a block diagram of an operational neural radiation field according to one exemplary embodiment of the present application. [Figure 5] FIG. 1 illustrates a diagram of a regional neural radiation field according to one exemplary embodiment of the present application. [Figure 6] FIG. 1 illustrates a block diagram of a rendering neural radiation field according to one exemplary embodiment of the present application. [Figure 7] 1 shows a flowchart of a method for training a facial modeling model according to another exemplary embodiment of the present application. [Figure 8] 1 shows a schematic diagram of an encoder according to one exemplary embodiment of the present application; [Figure 9] FIG. 1 shows a schematic diagram of a facial modeling model configuration according to one exemplary embodiment of the present application. [Figure 10] 1 shows a flowchart of a method for modeling a three-dimensional face according to one exemplary embodiment of the present application. [Figure 11] 1 illustrates a flowchart of generating a phoneme latent code index according to one exemplary embodiment of the present application. [Figure 12] 1 shows a block diagram of the structure of a facial modeling model training device according to one exemplary embodiment of the present application; [Figure 13] 1 shows a block diagram of the structure of a 3D face modeling device according to one exemplary embodiment of the present application; [Figure 14] 1 illustrates an exemplary structural diagram of an electronic device according to one exemplary embodiment of the present application; DETAILED DESCRIPTION OF THE INVENTION

[0016] To make the objectives, technical solutions and advantages of the present application clearer, the following describes the embodiments of the present application in more detail with reference to the drawings.

[0017] Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate and extend human intelligence, sense the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology in computer science that aims to understand the nature of intelligence and create new intelligent machines that can respond in a manner similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, equipping them with the capabilities of sensing, reasoning, and decision-making.

[0018] Artificial intelligence technology is a comprehensive academic field that covers a wide range of fields, including both hardware and software technologies. The basic technologies of artificial intelligence typically include sensors, artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.

[0019] Computer vision (CV) is the science that studies how machines can "see." More specifically, it refers to machine vision, which uses cameras and computers to identify and measure targets in place of the human eye, and then performs graphic processing to enable computers to obtain images suitable for human observation and machine detection. As a scientific field, CV studies related theories and technologies to establish artificial intelligence systems that can extract information from images or multidimensional data. Computer vision technologies typically include image processing, image identification, image segmentation, image semantic understanding, image retrieval, video processing, video semantic understanding, video content / action recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and map building, and other technologies, including general facial recognition and biometric recognition technologies such as fingerprint recognition. The method according to the present embodiment, i.e., computer vision technology, is applied to 3D facial reconstruction.

[0020] Related technologies primarily use voice-driven facial control to achieve 3D facial control. However, with voice-driven facial control, significant differences exist between different people's voices during the training process, which can lead to different training results. Furthermore, in actual use, if the recorded voice contains pronunciation errors, the voice must be re-recorded, which impacts facial control efficiency. Meanwhile, with text-driven facial control, related technologies can only drive 2D facial images, resulting in poor simulation results.

[0021] In an embodiment of the present application, a method for training a facial modeling model is proposed, which can train a model for reconstructing a 3D face, and the model can realize decoupling control of different facial attributes, for example, decoupling control of different facial regions and facial movements, which is useful for realizing long-sequence 3D facial modeling. After the model is trained, an encoder within the model can be used to extract latent code sequences of facial images corresponding to different phonemes. In application, the latent code sequences corresponding to the text phonemes can be queried, and long-sequence 3D facial modeling can be performed using the neural radiation field within the model to obtain 3D facial movement animation corresponding to the text, thereby realizing a text-based 3D facial driving method and improving the applicability of face driving.

[0022] The method according to the embodiment of the present application can be applied to a text-based facial driving scenario. For example, during a video conference, participants can input text through a program. When a computer device receives the input text, it can query the corresponding latent code sequence based on the text phonemes of the input text, thereby performing 3D reconstruction on the latent code sequence through a neural radiation field, obtaining a 3D facial movement animation, realizing simulated conversation, and improving the realism of the simulation. Of course, the method can also be applied to other scenarios requiring facial driving, such as network education, virtual anchors, and virtual social situations, but the embodiment is not limited thereto.

[0023] 1 shows a schematic diagram of an implementation environment according to one exemplary embodiment of the present application, including a terminal 110 and a server 120. Here, the terminal 110 and the server 120 perform data communication via a communication network, which may optionally be a wired network or a wireless network, and which may be at least one of a local area network, a metropolitan area network, and a wide area network.

[0024] The terminal 110 may be an electronic device with a text-based face activation function. The electronic device may be a mobile terminal such as a smartphone, a tablet, or a portable laptop computer, or may be a terminal such as a desktop computer or a projection computer, although the embodiments of the present application are not limited thereto. The terminal may also provide the text-based face activation function through a conference program, a live streaming program, an educational program, etc., although the embodiments of the present application are not limited thereto.

[0025] The server 120 may be an independent physical server, a server cluster or a distributed system configured by multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), big data, and artificial intelligence platforms. In an embodiment of the present application, the server 120 is a back-end server that provides the terminal 110 with a text-based face activation function, and can generate a 3D facial movement video based on text entered by a user and return it to the terminal 110.

[0026] As shown in FIG. 1, the terminal 110 can receive an input text input by a user, and then transmit the input text to the server 120, which stores a phoneme latent code index 122. The server 120 determines a text phoneme 121 corresponding to the input text, queries the phoneme latent code index 122 for a corresponding target latent code sequence 123, and then inputs the target latent code sequence 123 into a neural radiation field 124 for three-dimensional reconstruction to obtain a facial action image 125 corresponding to each text phoneme, thereby generating a three-dimensional facial action animation based on the facial action image 125 and feeding it back to the terminal 110.

[0027] Here, the phoneme latent code index is established based on the encoding results of different facial images by the encoder in the facial modeling model, and the encoder and neural radiation field of the facial modeling model are pre-trained using sample facial images, where the training process can be performed by the server 120, and subsequently, the server 120 uses the trained encoder to establish the phoneme latent code index and the neural radiation field to perform 3D reconstruction.

[0028] In another possible embodiment, the above 3D reconstruction process can also be performed by the terminal 110. The server 120 sends the trained model to the terminal 110, so that the terminal 110 can realize the 3D reconstruction locally without relying on the server 120. Alternatively, the face modeling model can be trained on the terminal 110 side, and the 3D reconstruction process can be performed by the terminal 110. The embodiments of the present application are not limited thereto.

[0029] For ease of explanation, the following embodiments will be described with reference to an example in which the method is performed by an electronic device.

[0030] 2, a flowchart of a method for training a face model according to an exemplary embodiment of the present application is shown. In this embodiment, the method is performed by a computer device as an example, and includes the following steps:

[0031] In step 201, sample face images are acquired, and the sample face images include face images of the same object at the same time and from different viewpoints.

[0032] In some embodiments, the electronic device can first acquire a set of multi-view sample images. In one possible embodiment, a multi-view camera shooting system can be used to collect facial images of the same object from different viewpoints to obtain facial images from different viewpoints at the same time. For example, when a reader is given a text in any language and the reader recites the text, the multi-view camera shooting system can capture the reader to obtain a video of the reader reciting the text from multiple viewpoints, where only the reader's face is captured during the recording to obtain a multi-view facial video. Then, the electronic device processes the multi-view facial video to obtain facial images of the reader from different viewpoints at the same time, and the facial images at the same time become a set of sample facial images, and the electronic device can obtain multiple sets of sample facial images through multi-view video processing.

[0033] Illustratively, the electronic device can acquire facial images of the reader from three viewpoints, i.e., a set of sample facial images includes facial images from a frontal viewpoint, a left-side viewpoint, and a right-side viewpoint, thereby including comprehensive facial features.

[0034] In step 202, the encoder of the face model performs feature encoding on the sample face image to obtain face region latent codes corresponding to different face regions and face action latent codes corresponding to different face actions.

[0035] In an embodiment of the present application, a facial modeling model is provided, where the model is configured with an encoder-neural radiation field structure. In one possible embodiment, the encoder is used to encode facial images from multiple viewpoints to obtain facial region latent codes corresponding to different facial regions and facial action latent codes corresponding to facial actions. That is, the electronic device inputs a set of sample facial images to the encoder, and the encoder of the facial model performs feature encoding on the sample facial images to obtain facial region latent codes and facial action latent codes corresponding to facial images at corresponding times. That is, the sample facial images are input to the encoder of the facial modeling model and perform feature encoding on them to obtain facial region latent codes corresponding to different facial regions and facial action latent codes corresponding to facial actions.

[0036] Here, the latent code is a feature vector for representing facial features. The facial region latent code is used to represent regional features of local facial regions, and the facial action latent code is used to represent facial action features such as mouth action, eye action, and nose action.

[0037] Different facial regions have different states. In some embodiments, the facial region may be divided into upper and lower regions, i.e., a facial region latent code corresponding to the upper half of the face and a facial region latent code corresponding to the lower half of the face may be coded, or a left and right region may be divided, i.e., a facial region latent code corresponding to the left half of the face and a facial region latent code corresponding to the right half of the face may be coded. This embodiment is not limited to this facial division method. Similarly, the number of regions to be divided is also not limited, and may be, for example, two, three, or four regions. This embodiment achieves decoupling control of the face through the region latent codes of different facial regions and the action latent codes of facial actions.

[0038] For example, a front image of a reader's face, a left image of the face, and a right image of the face at the same time can be input to an encoder to obtain a facial region latent code and a facial action latent code for the reader's face at that time.

[0039] In step 203, a 3D face reconstruction is performed on the object based on the face region latent code and the face action latent code through the neural radiation field of the face modeling model to obtain a 3D face image of the object.

[0040] Neural Radiance Field (NeRF) is a model for 3D reconstruction. NeRF represents a 3D scene as a neural network-approximated radiance field, which describes the color and volumetric density at each point in the scene and at each viewing direction. In the present embodiment, the neural radiation field includes multiple neural radiation fields, achieving decoupling control of each part of the face.

[0041] In one possible embodiment, the neural radiation fields include a neural radiation field for extracting face region features and a neural radiation field for extracting face action features, and finally, the multiple neural radiation fields are used to perform 3D face reconstruction, obtain pixel colors and pixel densities corresponding to each pixel, and render a 3D face image based on the pixel colors and pixel densities.

[0042] In actual implementation, the process of performing 3D face reconstruction on an object based on the face region latent code and the face action latent code through the neural radiation field of the face modeling model to obtain a 3D face image of the object is to input the face region latent code and the face action latent code into the neural radiation field of the face modeling model to perform 3D face reconstruction and obtain a 3D face image.

[0043] In step 204, a 3D reconstruction loss between the 3D face image and the sample face image is obtained, and a face modeling model is trained based on the 3D reconstruction loss.

[0044] Note that training a face modeling model based on the 3D reconstruction loss here means training the encoder and neural radiation field of the face modeling model based on the 3D reconstruction loss between a 3D face image and a sample face image.

[0045] After rendering the 3D face image, a 3D reconstruction loss can be determined based on the difference between the 3D face image and the sample face image, and the 3D reconstruction loss can be used to train a face modeling model.

[0046] In one possible embodiment, the model can be trained using gradient update or backpropagation. When the 3D reconstruction loss function reaches a convergence condition, the training can be terminated to obtain a face modeling model used for 3D face modeling, and the face modeling model can achieve decoupling control for the face. In the subsequent use process, the face action latent code can be changed only when the face action changes, and when a specific local area changes, only the face area latent code corresponding to the local area is changed, thereby achieving 3D face modeling of a long sequence.

[0047] In an embodiment of the present application, feature coding is performed on sample face images from multiple viewpoints to obtain face region latent codes and face action latent codes corresponding to the sample face images, which are used to represent the region features and face action features of different face regions. Then, 3D reconstruction is performed using the face region latent codes and face action latent codes to obtain a 3D face image. Thus, a face modeling model capable of face decoupling control is obtained by training based on the difference between the 3D face image and the sample face image. When the model is used for 3D face modeling, different face regions and facial actions can be decoupled and controlled individually, so that face images with combinations of different face regions and different facial actions can be generated. In the 3D face modeling process, by adjusting only the latent codes that need to be changed, a long sequence of 3D face modeling can be realized.

[0048] In some embodiments, the neural radiation fields include a regional neural radiation field, a motion neural radiation field, and a rendering neural radiation field, and different neural radiation fields are used to realize decoupling control for different faces and reconstruct three-dimensional face images, which will be described below using exemplary embodiments.

[0049] 3, a flowchart of a method for determining 3D reconstruction loss according to another exemplary embodiment of the present application is shown. In this embodiment, the method is performed by an electronic device as an example, and includes the following steps:

[0050] In step 301, sample face images are acquired, and the sample face images include face images of the same object at the same time and from different viewpoints.

[0051] The implementation of this step can refer to step 201 above, and will not be repeated in this embodiment.

[0052] In step 302, the encoder of the face model performs feature encoding on the sample face image to obtain an upper region latent code, a lower region latent code, and a facial action latent code, where the upper region latent code is used to represent facial features of the upper half of the face, and the lower region latent code is used to represent facial features of the lower half of the face.

[0053] In some embodiments, the face is divided into upper and lower halves, i.e., by inputting multi-view sample face images into an encoder and performing feature encoding, an upper region latent code corresponding to the upper half of the face, a lower region latent code corresponding to the lower half of the face, and a facial action latent code corresponding to the facial action can be obtained.

[0054] In one possible embodiment, the encoder first performs feature encoding on multi-view sample face images to obtain face state latent codes to represent the overall face state, and then decouples the face state latent codes to obtain face region latent codes and face action latent codes.

[0055] Illustratively, the image size of the sample face images at three viewpoints is 512 × 374, the latent code size of the face latent code is 256, and the latent code size of the face region latent code and facial action latent code obtained after decoupling is 128.

[0056] In step 303, the facial action latent code is input into the action neural radiation field to obtain the action transformation matrix.

[0057] The latent code-based neural radiation field achieves decoupled control over each part of the face through multiple neural radiation fields, where the neural radiation fields include a motion neural radiation field used to achieve decoupled control over facial motion.

[0058] The motion neural radiation field can perform feature extraction on the facial motion latent code to obtain a motion transformation matrix corresponding to the facial motion, and can use the motion transformation matrix to transform the facial motion.

[0059] Here, the motion neural radiation field is composed of a multi-layer fully connected network. As shown in Figure 4, Figure 4 shows a configuration diagram of the motion neural radiation field according to one exemplary embodiment of the present application, the motion neural radiation field includes a four-layer fully connected network 401, the network width is 128, and the facial motion latent code Z p By inputting into the motion neural radiation field, the motion transformation matrix (R, t) can be obtained.

[0060] In step 304, the three-dimensional face coordinates of the object are obtained based on the sample face image, and a motion transformation is performed on the three-dimensional face coordinates based on the motion transformation matrix to obtain the face motion coordinates.

[0061] In actual implementation, after determining a sample face image, a three-dimensional coordinate system of the sample face image is constructed, and then the three-dimensional facial coordinates of the object corresponding to the sample face image are obtained based on the three-dimensional coordinate system.

[0062] After obtaining the motion transformation matrix, the three-dimensional coordinates of the face can be transformed using the motion transformation matrix to obtain the facial motion coordinates corresponding to the facial motion.

[0063] Here, the conversion method is as follows.

[0064] x′=Rx+t Equation (1) where: TIFF2025535488000002.tif5170 indicates the 3D coordinates before transformation, and (R, t) is the transformation matrix.

[0065] By performing a motion transformation on the three-dimensional face coordinates (x, y, z), the face motion coordinates (x', y', z') can be obtained.

[0066] In another possible embodiment, the coordinates (x,y,z) can be divided by a motion transformation matrix (R,t) to obtain the transformed facial motion coordinates.

[0067] In step 305, the face action coordinates and the face region latent code are input into the regional neural radiation field to obtain a face region slice plane (masked image), and the feature dimension of the face region slice plane is higher than the feature dimension of the face region latent code.

[0068] The neural radiation field further includes a regional neural radiation field, which is used to realize decoupling control for different face regions. Here, the regional neural radiation field is used to perform feature mapping, mapping latent codes to feature vectors in a high-dimensional space, i.e., the slice plane is the mapping of the latent codes in the high-dimensional space. In one possible embodiment, the computer device can input the face action coordinates and the face region latent code into the regional neural radiation field to obtain a face region slice plane in the high-dimensional space, whose feature dimension is higher than that of the face region latent code, and learn more face region features.

[0069] The face regional latent code includes an upper regional latent code and a lower regional latent code, and correspondingly, the regional neural radiation fields can include an upper neural radiation field for controlling the upper half of the face, and also include a lower neural radiation field for controlling the lower half of the face.

[0070] In some embodiments, the facial motion coordinates and the upper region latent code are input into the upper neural radiation field to obtain the upper region slice plane, i.e., the facial motion coordinates (x', y', z') and the upper region latent code z f is input to the upper neural radiation field, and the upper region slice plane w f get.

[0071] In some embodiments, the facial motion coordinates and the lower region latent code are input into the lower neural radiation field to obtain the lower region slice plane, i.e., the facial motion coordinates (x', y', z') and the lower region latent code z m is input to the lower neural radiation field, and the lower region slice plane w m get.

[0072] In some embodiments, the regional neural radiation field also employs a multi-layer fully connected network. For example, as shown in FIG. 5, FIG. 5 shows a block diagram of the regional neural radiation field according to one exemplary embodiment of the present application, the regional neural radiation field includes a six-layer fully connected network 501, the network width is 128, and the facial action coordinates (x', y', z') and the facial region latent code (z f / z m ) is subjected to position coding and input into the regional neural radiation field, and the face region slice plane (w f / w m ) is obtained.

[0073] After the upper region slice surface and the lower region slice surface are obtained, a face region slice surface is obtained based on the upper region slice surface and the lower region slice surface.

[0074] In step 306, the facial motion coordinates and the facial region slice plane are input into the rendering neural radiation field to obtain the facial color and facial density.

[0075] The neural radiation field further includes a rendering neural radiation field, which is used to perform 3D face reconstruction. After obtaining the upper region slice plane and the lower region slice plane, the facial motion coordinates and the facial region slice plane (including the upper region slice plane and the lower region slice plane) are input into the rendering neural radiation field to obtain the facial color and facial density.

[0076] In addition, in the 3D reconstruction process, the captured view direction and appearance code need to be input into the fully connected network of the rendering neural radiation field to perform 3D reconstruction and obtain the pixel color and pixel density of each pixel point corresponding to the 3D face.

[0077] In some embodiments, the rendering neural radiation field also employs a multi-layer fully connected network. For example, as shown in FIG. 6, FIG. 6 shows a configuration diagram of the rendering neural radiation field according to one exemplary embodiment of the present application, the rendering neural radiation field includes a 7-layer fully connected network 601, where the network width is 256, and the facial motion coordinates (x', y', z') and the facial region slice plane (w f / w m ) can be position-coded and input into the rendering neural radiation field, and the view direction (θ, Φ) and appearance code (ψ) can be input into the final layer of a fully connected network to perform 3D reconstruction and obtain the RGB color of the face c and the face density (mask) σ.

[0078] In step 307, a three-dimensional facial image is generated based on the facial color and facial density.

[0079] In step 308, a pixel error loss is determined based on pixel differences between the 3D face image and the sample face image.

[0080] After the 3D reconstruction, a 3D reconstruction loss can be determined based on the difference between the sample face image and the 3D face image. In some embodiments, a pixel error loss can be determined based on pixel difference, where the pixel difference is determined based on the RGB difference between the sample face image and each pixel point of the 3D face image. In one possible embodiment, a pixel-to-pixel mean squared error (MSE) loss can be calculated based on the pixel-to-pixel difference to obtain the pixel error loss.

[0081] In step 309, a mask error loss is determined based on the mask difference between the 3D face image and the sample face image.

[0082] The mask error loss can also be determined based on the mask difference between faces. In one possible embodiment, a mask corresponding to a sample face image can be obtained, and the cross-entropy loss between the mask and the mask of the 3D face image can be calculated to obtain the mask error loss.

[0083] In step 310, a 3D reconstruction loss is determined based on the pixel error loss and the mask error loss.

[0084] Here, the 3D reconstruction loss is the sum of the pixel error loss and the mask error loss, as follows: TIFF2025535488000003.tif6170 where, TIFF2025535488000004.tif7170 is sample face image I and 3D face image TIFF2025535488000005.tif7170 is the pixel error loss between TIFF2025535488000006.tif7170 is the face mask M corresponding to the sample face image. h is the mask error loss between the 3D face image and the mask (opaque image) α, and λ is a balance parameter.

[0085] In actual implementation, after determining the 3D reconstruction loss, a face modeling model is trained based on the 3D reconstruction loss, that is, the encoder and neural radiation field are trained based on the 3D reconstruction loss, which may be to use the 3D reconstruction loss to train the encoder and neural radiation field (including the motion neural radiation field, the area neural radiation field, and the rendering neural radiation field) to obtain a face modeling model.

[0086] In this embodiment, the decoupling control of facial movements and different regions is achieved through the motion neural radiation field and the region neural radiation field, respectively, and then the rendering neural radiation field is used to perform 3D reconstruction to obtain a 3D facial image. Because decoupling control is possible, the efficiency of 3D facial reconstruction can be improved by subsequently modifying only the corresponding face latent code, thereby achieving highly efficient long-sequence 3D facial modeling.

[0087] In order to ensure accurate separation of different attributes of faces, in one possible embodiment, normalization of upper and lower faces is added to decouple the upper and lower faces, that is, the decoupling loss is determined in the manner of a control variable, and the decoupling loss is used to train an encoder, thereby improving the accuracy of the latent codes corresponding to the different attributes obtained by decoupling. This will be described below using an illustrative example.

[0088] 7, there is shown a flowchart of a method for training a facial model according to another exemplary embodiment of the present application. In this embodiment, the method is performed by an electronic device as an example, and includes the following steps:

[0089] In step 701, sample face images are acquired, and the sample face images include face images of the same object at the same time and from different viewpoints.

[0090] The implementation of this step can refer to step 201 above, and will not be repeated in this embodiment.

[0091] In step 702, feature coding is performed on the sample facial image via a coding network to obtain a facial state latent code, which is used to represent the overall facial state of the reader when reading the text at the time corresponding to the sample facial image.

[0092] In practical implementation, the sample face image is subjected to feature coding through the coding network to obtain a face state latent code, that is, the sample face image is input to the coding network to undergo feature coding to obtain a face state latent code.

[0093] The encoder includes a coding network and a decoupling neural radiation field, where the coding network is used to perform feature coding on multi-view sample face images to obtain a face state latent code for representing an overall face state.

[0094] In some embodiments, the encoding network is composed of a convolutional layer and a cascade layer, where the convolutional layer is used to perform two-dimensional convolution on input sample face images, and the cascade layer is used to cascade the intermediate representations corresponding to each image to obtain a face state latent code.

[0095] Illustratively, as shown in FIG. 8, FIG. 8 shows a schematic configuration diagram of an encoder according to one exemplary embodiment of the present application, where sample input images include image 1 to image k, and image 1 to image k with a size of 512×374 are respectively input into the convolutional layer of the encoding network and convolved to obtain an intermediate vector 801 with a size of 4×256, and then the intermediate vector 801 is reshaped and cascaded to obtain a face state latent code Z802 with a size of 256.

[0096] In step 703, the facial state latent code is decoupled via a decoupling neural radiation field to obtain an upper region latent code, a lower region latent code, and a facial action latent code, and a facial region latent code is determined based on the upper region latent code and the lower region latent code, where the upper region latent code is used to represent the facial features of the upper half of the face, and the lower region latent code is used to represent the facial features of the lower half of the face.

[0097] In actual implementation, the facial state latent code is decoupled through the decoupling neural radiation field to obtain the upper region latent code, the lower region latent code, and the facial action latent code; that is, the facial state latent code is input into the decoupling neural radiation field and decoupled to obtain the upper region latent code, the lower region latent code, and the facial action latent code.

[0098] The encoder further includes a decoupling neural radiation field, which is used to decouple the global latent code and obtain face region latent codes (including upper region latent codes and lower region latent codes) and facial action latent codes for different regions.

[0099] In some embodiments, the decoupling neural radiation field is composed of a multi-layer fully connected network, such as a three-layer fully connected network 803 shown in FIG. 8 . After inputting the face state latent code Z into the decoupling neural radiation field, the face state latent code Z is converted into the upper region latent code Z. f , lower region latent code z m , and facial action latent code z p can be decoupled to

[0100] In step 704, a 3D face reconstruction is performed on the object based on the face region latent code and the face action latent code via the neural radiation field of the face modeling model to obtain a 3D face image of the object.

[0101] In step 705, the 3D reconstruction loss between the 3D face image and the sample face image is obtained, and a face modeling model is trained based on the 3D reconstruction loss.

[0102] Here, the implementation of steps 704 to 705 can refer to steps 303 to 309 in the above embodiment, and will not be repeated in this embodiment.

[0103] In step 706, a decoupling loss is determined based on the upper region latent code, the lower region latent code, and the facial action latent code at different times, and the decoupling loss is used to represent the difference in facial decoupling at different times and is used to train a decoupling neural radiation field.

[0104] In one possible embodiment, the decoupling neural radiation field training is realized using a control variable method. When only changing the latent code in a local region, the rendering results in other regions are not affected. Therefore, the rendering results corresponding to the latent code combinations at different times can be used to determine the decoupling loss. Here, this method may include steps 706a to 706e (not shown).

[0105] In step 706a, a first upper region latent code, a first lower region latent code, and a first facial movement latent code corresponding to the sample facial image at a first time are obtained, and a second upper region latent code, a second lower region latent code, and a second facial movement latent code corresponding to the sample facial image at a second time are obtained.

[0106] First, the electronic device uses an encoder to encode sample face images at two different times (a first time t1 and a second time t2) to generate a first upper region latent code, a first lower region latent code, and a first facial action latent code (i.e., TIFF2025535488000007.tif12170), and the second upper region latent code, the second lower region latent code, and the second facial action latent code at the second time (i.e., TIFF2025535488000008.tif12170) can be obtained respectively.

[0107] In step 706b, an upper decoupling loss is determined based on the first upper region latent code, the first facial action latent code, and the second lower region latent code.

[0108] If only the lower region latent code at the first time is modified, the rendering result of the upper half of the face will not be affected. Therefore, in one possible embodiment, the lower region latent code at the first time is used to render a 3D face image based on the modified latent code combination, and the difference between the 3D face image and the upper half of the face in the sample face image corresponding to the first time is used to determine the upper decoupling loss. In one possible embodiment, the method may include the following steps:

[0109] In step 1, a first upper region latent code, a first facial action latent code, and a second lower region latent code are input into a neural radiation field to obtain a first facial image.

[0110] The electronic device changes the first lower region latent code at the first time to the second lower region latent code at the second time, thereby generating a first latent code combination. TIFF2025535488000009.tif12170 is obtained, and then the first latent code combination is input into the neural radiation field to perform 3D face reconstruction, and the first face image is obtained. You can get TIFF2025535488000010.tif10170.

[0111] In step 2, an upper decoupling loss is determined based on the difference between the upper half of the face in the first face image and the upper half of the face in the sample face image at the first time.

[0112] Since we only changed the latent code of the lower half of the face, the upper half of the face image obtained by rendering should be the same as the upper half of the face image in the sample face image, and the upper decoupling loss can be determined based on the difference between the upper half of the face images.

[0113] In one possible embodiment, the top decoupling losses are as follows: TIFF2025535488000011.tif6170 where M f is the upper half of the face mask, TIFF2025535488000012.tif10170 is the first face image rendered based on the first latent code combination, TIFF2025535488000013.tif7170 is a sample face image corresponding to the first time point, TIFF2025535488000014.tif7170 refers to the mean squared error between the two.

[0114] In step 706c, a lower decoupling loss is determined based on the first lower region latent code, the first facial action latent code, and the second upper region latent code.

[0115] Similarly, if only the upper region latent code at the first time is changed, the rendering result of the lower half of the face will not be affected, so in one possible embodiment, the upper region latent code at the first time is used to render a 3D face image based on the changed latent code combination, and the difference between the 3D face image and the lower half of the face in the sample face image corresponding to the first time is used to determine the lower decoupling loss. In one possible embodiment, the method may include the following steps.

[0116] In step 1, the first lower region latent code, the first facial action latent code, and the second upper region latent code are input into a neural radiation field to obtain a second facial image.

[0117] The electronic device changes the first upper region latent code at the first time to a second upper region latent code at the second time to generate a second latent code combination. TIFF2025535488000015.tif12170 is obtained, and then the second latent code combination is input into the neural radiation field to perform 3D face reconstruction, and the second face image is obtained. You can get TIFF2025535488000016.tif10170.

[0118] In step 2, a lower decoupling loss is determined based on the difference between the lower half of the face in the second face image and the lower half of the face in the sample face image at the first time.

[0119] Since only the latent code of the upper half of the face is changed, the lower half of the face image obtained by rendering should be the same as the lower half of the face image in the sample face image corresponding to the first time, and the lower decoupling loss can be determined based on the difference between the lower half of the face image.

[0120] In one possible embodiment, the bottom decoupling losses are as follows: TIFF2025535488000017.tif6170 where M m is the lower half of the face mask, TIFF2025535488000018.tif10170 is the second face image rendered based on the second latent code combination, TIFF2025535488000019.tif7170 is a sample face image corresponding to the first time point, TIFF2025535488000020.tif7170 refers to the mean squared error between the two.

[0121] In step 706d, a motion decoupling loss is determined based on the second facial motion latent code, the first upper region latent code, and the first lower region latent code.

[0122] To improve the accuracy of facial motion decoupling, a motion decoupling loss is also introduced. When a facial motion is changed, the mask of the rendered image is not affected, so a facial motion latent code at a first time point can be used to render a 3D facial image based on the changed latent code combination, and a motion decoupling loss can be determined using the difference between the mask of the 3D facial image and the mask of the sample facial image corresponding to the first time point. The method may include the following steps:

[0123] In step 1, the second facial action latent code, the first upper region latent code, and the first lower region latent code are input into a neural radiation field to obtain a third facial image.

[0124] In one possible embodiment, a first facial movement latent code at a first time point is changed to a second facial movement latent code at a second time point to form a third latent code combination. TIFF2025535488000021.tif12170 can be obtained. Then, the third latent code combination can be input into the neural radiation field for 3D reconstruction to obtain the third face image.

[0125] In step 2, a motion decoupling loss is determined based on the difference between the mask of the third face image and the mask of the sample face image at the first time instant.

[0126] Even if the facial motion changes, the facial projection does not change, i.e., the facial mask corresponding to the third facial image should be the same as the facial mask corresponding to the sample facial image at the first time, and the motion decoupling loss can be determined based on the mask difference between the two.

[0127] Here, the operational decoupling loss is given by the following equation: TIFF2025535488000022.tif7170 where, TIFF2025535488000023.tif10170 is the face mask corresponding to the third face image, and M h(p) is a face mask corresponding to the sample face image corresponding to the first time.

[0128] In addition, this embodiment is not limited to the execution order of determining the above upper decoupling loss, lower decoupling loss, and operational decoupling loss (i.e., steps 706b to 706d); these three steps can be executed synchronously or sequentially. This embodiment only provides an illustrative explanation of the implementation form, and is not limited to the order of steps.

[0129] In step 706e, a decoupling loss is determined based on the top decoupling loss, the bottom decoupling loss, and the operating decoupling loss.

[0130] After determining the upper decoupling loss, the lower decoupling loss, and the operational decoupling loss, the sum of these three is determined as the decoupling loss to be used to train the decoupling neural radiation field.

[0131] In step 707, a face modeling model is trained based on the decoupling loss and the 3D reconstruction loss.

[0132] After determining the decoupling loss, the process of training a face modeling model based on the 3D reconstruction loss may be to train a face modeling model based on the decoupling loss and the 3D reconstruction loss.

[0133] In the above example, we further use the 3D reconstruction loss to train the encoder (including training the decoupled neural radiation fields), so the total loss corresponding to the decoupled neural radiation fields is as follows: TIFF2025535488000024.tif6170 where, TIFF2025535488000025.tif7170 is the 3D reconstruction loss, TIFF2025535488000026.tif7170 is the upper decoupling loss, TIFF2025535488000027.tif7170 is the lower decoupling loss, TIFF2025535488000028.tif7170 is the operational decoupling loss.

[0134] By training the decoupling neural radiation field using the total loss, the accuracy of decoupling of the decoupling neural radiation field can be improved.

[0135] In one possible embodiment, the model structure of the overall face modeling model is as shown in Figure 9, which shows a schematic diagram of the configuration of the face modeling model according to one exemplary embodiment of the present application, and includes an encoder 91 and a neural radiation field 92. The encoder 91 includes an encoding network 910 and a decoupling neural radiation field 911. When inputting sample face images at three viewpoints, the encoding network 910 can encode the face state latent code, and then the decoupling neural radiation field 911 decodes the face state latent code to obtain the upper region latent code z f , lower region latent code z m , and facial action latent code z p can be obtained.

[0136] After completing the encoding, the upper region latent code z f , lower region latent code z m , and facial action latent code z p can be input to the neural radiation field 92, where the neural radiation field 92 includes an action neural radiation field 921, a region neural radiation field 922, and a rendering neural radiation field 923. The action neural radiation field 921 generates a facial action latent code z p Then, the facial 3D coordinate x is transformed to obtain the facial motion coordinate x′, and then the upper region latent code z f , lower region latent code z m, and the facial motion coordinate x′ are input to the regional neural radiation field 922, and the regional neural radiation field 922 is used to calculate the upper regional slice plane w f and the lower region slice plane w m and the upper region slice plane w f , the lower region slice plane w m , and the facial motion coordinate x′ are input into the rendering neural radiation field, which are used to obtain the facial color c and facial density σ. Here, when generating the facial color c, the color is affected by the gaze direction and environmental lighting, etc., so to complete the 3D face reconstruction, the gaze direction and appearance code d(θ, Φ, ψ) also need to be input into the rendering neural radiation field.

[0137] In this embodiment, the latent codes at different times are combined to determine the decoupling loss in the manner of control variables, so that the decoupling loss can be used to train the decoupling neural radiation field to improve the decoupling accuracy, which is helpful in improving the accuracy of the decoupling control.

[0138] In the above embodiments, a facial modeling model can be trained, and the trained facial modeling model can be used in the text-based facial driving process, which will be described below with exemplary embodiments.

[0139] Referring to Figure 10, Figure 10 shows a flowchart of a 3D face modeling method according to one exemplary embodiment of the present application. This embodiment is described as an example in which the method is performed by an electronic device, and the method includes the following steps:

[0140] In step 1001, when an input text is received, the text phonemes corresponding to the input text are determined.

[0141] In one possible embodiment, the electronic device is running a text-based face driving program, and when receiving an input text, the electronic device can determine the text phonemes corresponding to the input text to obtain a text phoneme sequence corresponding to the input text.

[0142] In step 1002, a target region latent code and a target action latent code corresponding to a text phoneme are queried based on a phoneme latent code index, where the phoneme latent code index indicates the correspondence between the phoneme and the latent code sequence, and the latent code sequence is obtained by performing feature encoding on the face image corresponding to the phoneme based on an encoder of a face modeling model.

[0143] After training the facial model, an encoder within the model can encode facial images to obtain different latent codes. In one possible embodiment, the electronic device uses the encoder to encode facial images corresponding to different phonemes to obtain latent code sequences corresponding to different phonemes. Here, the latent code sequences corresponding to each phoneme include a facial region latent code and a facial action latent code. Based on the encoding results of the facial images by the encoder, a phoneme latent code index can be established to indicate the correspondence between the phonemes and the latent code sequences.

[0144] In one possible embodiment, the electronic device can obtain the target region latent code and the target action latent code corresponding to each phoneme by respectively querying the latent code sequence corresponding to each phoneme. For example, if the text contains "pu", the phonemes "p" and "u" are included, and the facial region latent code and the facial action latent code corresponding to "p" and the facial region latent code and the facial action latent code corresponding to "u" are respectively query.

[0145] Alternatively, in another possible embodiment, the phoneme sequence can be determined first, and the corresponding latent code sequence can be searched based on the phoneme sequence, i.e., the latent code sequence corresponding to "pu" can be searched. In some embodiments, a longest identical subsequence search method can be adopted to search for the latent code corresponding to the phoneme.

[0146] Here, the facial images corresponding to phonemes may be images collected in advance. For different pronunciation modes of phonemes, facial images corresponding to the pronunciation modes can be collected respectively. For example, for the phonemes of the same text, facial images corresponding to Mandarin pronunciation, facial images corresponding to dialect pronunciation, and facial images corresponding to singing pronunciation, etc. can be collected respectively, so that the latent code sequence corresponding to the phonemes in the Mandarin pronunciation mode (i.e., obtain the phoneme latent code index corresponding to Mandarin), the latent code sequence corresponding to the phonemes in the dialect pronunciation mode (phoneme latent code index corresponding to the dialect), and the latent code sequence corresponding to the phonemes when singing (phoneme latent code index corresponding to singing) can be encoded to provide the user with different facial activation modes.

[0147] When searching for a latent code sequence corresponding to a phoneme based on a phoneme latent code index, a corresponding phoneme latent code index is determined based on a pronunciation system selected by a user, and a corresponding latent code sequence is searched for accordingly.

[0148] In step 1003, through the neural radiation field of the face modeling model, three-dimensional reconstruction is performed based on the target region latent code and the target action latent code to obtain a face action image corresponding to the text phoneme.

[0149] In actual implementation, the process of performing three-dimensional reconstruction based on the target area latent code and the target action latent code through the neural radiation field of the face modeling model to obtain a facial action image corresponding to the text phoneme is to input the target area latent code and the target action latent code into the neural radiation field of the face modeling model to perform three-dimensional reconstruction to obtain a facial action image corresponding to the text phoneme.

[0150] In the process of determining the latent code sequence corresponding to the text phoneme, the latent code sequence can be input into the neural radiation field to perform 3D reconstruction. In one possible embodiment, the target action latent code is input into the action neural radiation field to obtain an action transformation matrix, and the predetermined facial 3D coordinates are transformed using the action transformation matrix to obtain action-transformed facial action coordinates, and the facial action coordinates and the upper region latent code are input into the upper neural radiation field to obtain an upper region slice plane, and the facial action coordinates and the lower region latent code are input into the lower neural radiation field to obtain a lower region slice plane, and finally, the upper region slice plane, the lower region slice plane and the facial action coordinates are input into the rendering neural radiation field to render a facial action image corresponding to the text phoneme.

[0151] The text phoneme corresponding to the input text includes multiple phonemes. In one possible embodiment, a latent code sequence corresponding to each text phoneme is input into the neural radiation field to obtain a reading facial image corresponding to each text phoneme. When querying the latent code sequence corresponding to the text phoneme, there may be a case where some latent codes corresponding to adjacent phonemes are the same, for example, the upper region latent codes corresponding to two adjacent phonemes are the same. In this case, when reconstructing the facial action image corresponding to the second phoneme, only the modified latent codes (e.g., the lower region latent code and the facial action latent code) can be input into the neural radiation field to perform 3D reconstruction, thereby realizing 3D facial modeling of a long sequence.

[0152] In step 1004, a three-dimensional facial movement animation corresponding to the input text is generated based on the facial movement image of each frame.

[0153] After generating the multi-frame facial movement images, the electronic device combines the multi-frame facial movement images to obtain a 3D facial movement animation corresponding to the input text, thereby providing a realistic 3D facial pronunciation simulation. That is, when the electronic device receives an input text, it can generate a 3D facial reading animation corresponding to the input text, and the animation is changed according to changes in the text, thereby realizing a 3D facial driving method according to the text and improving the realistic simulation.

[0154] In this embodiment, an encoder performs feature encoding on facial images corresponding to different phonemes, thereby obtaining correspondences between different phonemes and latent code sequences, and stores phoneme latent code indexes. When the electronic device receives text entered by a user, it queries the corresponding latent code sequence based on the phoneme latent code index and inputs it into a neural radiation field to obtain a 3D facial motion animation corresponding to the input text, thereby realizing text-based 3D facial driving. The user can change the 3D facial animation by changing the input text, thereby improving the applicability of facial driving. Furthermore, because the neural radiation field enables decoupling control of the face, if only some areas need to be changed, only some latent codes need to be adjusted, making it easier to model 3D facial sequences over long sequences.

[0155] Here, the phoneme latent code index is determined based on the encoding results of different facial images by the encoder in the facial modeling model. Different phoneme latent code indexes can be established for different pronunciation modes, and facial movements corresponding to the different pronunciation modes can be simulated through the different phoneme latent code indexes to generate corresponding 3D facial movement animations. For example, 3D facial movement animations corresponding to modes such as Mandarin, dialects, or singing can be simulated to enrich the diversity of facial movements driven by text. The following describes the process of establishing a phoneme latent code index in the mode of reading text aloud as an example. Referring to FIG. 11, FIG. 11 shows a flowchart of generating a phoneme latent code index according to an exemplary embodiment of the present application, and the process includes the following steps:

[0156] In step 1101, a text reading image is acquired, and the text reading image includes facial images of a reader reading a text, collected at the same time from different viewpoints.

[0157] In one possible embodiment, a multi-view camera collection system can be used to collect facial images of readers when they recite different texts, and in this process, audio and video recording can be performed simultaneously to obtain recitation video and recitation audio, where the recitation video includes the facial images of the readers.

[0158] By aligning the spoken voice with the spoken text, a time axis corresponding to each phoneme can be determined, and a text reading image corresponding to each phoneme can be determined based on the time axis corresponding to each phoneme, where the text reading image corresponding to each phoneme includes face images collected from different viewpoints.

[0159] In step 1102, feature coding is performed on the text reading image based on the encoder to obtain a face latent code sequence corresponding to the text reading image, where the face latent code sequence includes a region latent code and an action latent code.

[0160] In actual implementation, based on the encoder, feature coding is performed on the text reading image to obtain a face latent code sequence corresponding to the text reading image, where the face latent code sequence includes a region latent code and an action latent code, that is, the text reading image is input to the encoder to perform feature coding, and a face latent code sequence corresponding to the text reading image is obtained.

[0161] The electronic device can use an encoder to perform feature encoding on the text reading image corresponding to each phoneme, in which a coding network is first used to encode a facial state latent code to represent the overall state, and then a decoupling neural radiation field is used to decouple the facial region latent code and the facial action latent code to obtain a latent code sequence corresponding to each phoneme.

[0162] In step 1103, the face latent code sequence is stored in association with the reading phonemes corresponding to the text reading image to obtain a phoneme latent code index.

[0163] The electronics can store the face latent code sequences in association with the corresponding spoken phonemes for subsequent use in the text-to-face driving process.

[0164] Referring to Figure 12, Figure 12 shows a block diagram of the structure of a facial modeling model training device according to one exemplary embodiment of the present application. an image acquisition module 1201 configured to acquire sample face images, the sample face images including face images of the same object at the same time but from different viewpoints; a feature encoding module 1202 configured to perform feature encoding on the sample face image by an encoder of the face modeling model to obtain face region latent codes corresponding to different face regions and face action latent codes corresponding to different face actions; a 3D reconstruction module 1203 configured to perform 3D face reconstruction on the object based on the face region latent code and the facial action latent code via a neural radiation field of the face modeling model to obtain a 3D face image of the object; a model training module 1204 configured to obtain a 3D reconstruction loss between the 3D face image and the sample face image, and train the face modeling model based on the 3D reconstruction loss.

[0165] In some embodiments, the neural radiation fields include a regional neural radiation field, a motion neural radiation field, and a rendering neural radiation field; The 3D reconstruction module 1203 further comprises: inputting the facial action latent code into the action neural radiation field to obtain an action transformation matrix; acquire three-dimensional facial coordinates of the object based on the sample facial image, perform motion transformation on the three-dimensional facial coordinates based on the motion transformation matrix, and obtain facial motion coordinates; inputting the facial motion coordinates and the facial region latent code into the regional neural radiation field to obtain a facial region slice plane, wherein the feature dimension of the facial region slice plane is higher than the feature dimension of the facial region latent code; inputting the facial motion coordinates and the facial region slice plane into the rendering neural radiation field to obtain facial color and facial density; The three-dimensional facial image is configured to be generated based on the facial color and the facial density.

[0166] The feature encoding module 1202 further comprises: an encoder of the face modeling model is configured to perform feature encoding on the sample face image to obtain an upper region latent code, a lower region latent code, and the facial action latent code, wherein the upper region latent code is used to represent facial features of an upper half of a face, and the lower region latent code is used to represent facial features of a lower half of a face; The 3D reconstruction module 1203 further comprises: Input the facial motion coordinates and the upper region latent code into an upper neural radiation field to obtain an upper region slice surface; Input the facial motion coordinates and the lower region latent code into the lower neural radiation field to obtain a lower region slice surface; The face region slice plane is configured to be obtained based on the upper region slice plane and the lower region slice plane.

[0167] The model training module 1204 further comprises: determining a pixel error loss based on pixel differences between the 3D face image and the sample face image; determining a mask error loss based on mask differences between the 3D face image and the sample face image; and determining the 3D reconstruction loss based on the pixel error loss and the mask error loss.

[0168] In some embodiments, the encoder comprises a coding network and a decoupling neural radiation field; The feature encoding module 1202 further comprises: performing feature coding on the sample facial image via the coding network to obtain a facial state latent code, which is used to represent an overall facial state of an object corresponding to the sample facial image when reading a text; decoupling the facial state latent code via the decoupling neural radiation field to obtain an upper region latent code, a lower region latent code, and the facial action latent code; and determining the facial region latent code based on the upper region latent code and the lower region latent code; Here, the upper region latent code is used to represent facial features of the upper half of the face, and the lower region latent code is used to represent facial features of the lower half of the face.

[0169] In some embodiments, the device further comprises: a loss determination module configured to determine a decoupling loss based on the upper region latent code, the lower region latent code, and the facial action latent code at different times, the decoupling loss being used to represent differences in facial decoupling at different times and being used to train the decoupling neural radiation field; The model training module 1204 is further configured to train the face modeling model based on the decoupling loss and the 3D reconstruction loss.

[0170] In some embodiments, the loss determination module further comprises: obtaining a first upper region latent code, a first lower region latent code, and a first facial movement latent code corresponding to the sample face image at a first time; obtaining a second upper region latent code, a second lower region latent code, and a second facial movement latent code corresponding to the sample face image at a second time; determining an upper decoupling loss based on the first upper region latent code, the first facial expression latent code, and the second lower region latent code; determining a lower decoupling loss based on the first lower region latent code, the first facial expression latent code, and the second upper region latent code; determining a motion decoupling loss based on the second facial motion latent code, the first upper region latent code, and the first lower region latent code; The decoupling loss is configured to determine the decoupling loss based on the top decoupling loss, the bottom decoupling loss, and the operational decoupling loss.

[0171] In some embodiments, the loss determination module further comprises: inputting the first upper region latent code, the first facial action latent code, and the second lower region latent code into the neural radiation field to obtain a first facial image; determining the upper decoupling loss based on a difference between an upper half of a face in the first face image and an upper half of a face in a sample face image at the first time; In some embodiments, the loss determination module further comprises: inputting the first lower region latent code, the first facial action latent code, and the second upper region latent code into the neural radiation field to obtain a second facial image; The lower decoupling loss is configured to determine the lower decoupling loss based on a difference between a lower half of a face in the second face image and a lower half of a face in the sample face image at the first time.

[0172] In some embodiments, the loss determination module further comprises: inputting the second facial action latent code, the first upper region latent code, and the first lower region latent code into the neural radiation field to obtain a third facial image; The motion decoupling loss is configured to determine the motion decoupling loss based on a difference between a mask of the third facial image and a mask of the sample facial image at the first time.

[0173] In an embodiment of the present application, feature coding is performed on sample face images from multiple viewpoints to obtain face region latent codes and face action latent codes corresponding to the sample face images, which are used to represent the region features and face action features of different face regions. Then, 3D reconstruction is performed using the face region latent codes and face action latent codes to obtain a 3D face image. Thus, a face modeling model capable of face decoupling control is obtained by training based on the difference between the 3D face image and the sample face image. When the model is used for 3D face modeling, different face regions and facial actions can be decoupled and controlled individually, so that face images with combinations of different face regions and different facial actions can be generated. In the 3D face modeling process, by adjusting only the latent codes that need to be changed, a long sequence of 3D face modeling can be realized.

[0174] Referring to Figure 13, Figure 13 shows a block diagram of the structure of a 3D face modeling device according to one exemplary embodiment of the present application, which comprises: a phoneme determination module 1301 configured to, when receiving an input text, determine text phonemes corresponding to said input text; a latent code lookup module 1302 configured to look up a target region latent code and a target action latent code corresponding to the text phoneme based on a phoneme latent code index, where the phoneme latent code index indicates a correspondence between a phoneme and a latent code sequence, and the latent code sequence is obtained by performing feature encoding on a face image corresponding to the phoneme based on an encoder of a face model; a 3D reconstruction module 1303 configured to perform 3D reconstruction based on the target region latent code and the target action latent code through a neural radiation field of the face modeling model to obtain a facial action image corresponding to the text phoneme; and an animation generation module 1304 configured to generate a three-dimensional facial movement animation corresponding to the input text based on the facial movement image of each frame.

[0175] In some embodiments, the device further comprises: an image acquisition module configured to acquire text reading images, the text reading images including facial images of a reader reading the text collected at the same time from different viewpoints; a feature encoding module configured to perform feature encoding on the text reading image based on the encoder to obtain a face latent code sequence corresponding to the text reading image, the face latent code sequence including a region latent code and an action latent code; a storage module configured to store the face latent code sequence in association with a reading phoneme corresponding to the text reading image to obtain the phoneme latent code index.

[0176] In this embodiment, an encoder performs feature encoding on facial images corresponding to different phonemes, thereby obtaining correspondences between different phonemes and latent code sequences, and stores phoneme latent code indexes. When the electronic device receives text entered by a user, it queries the corresponding latent code sequence based on the phoneme latent code index and inputs it into a neural radiation field to obtain a 3D facial motion animation corresponding to the input text, thereby realizing text-based 3D facial driving. The user can change the 3D facial animation by changing the input text, thereby improving the applicability of facial driving. Furthermore, because the neural radiation field enables decoupling control of the face, if only some areas need to be changed, only some latent codes need to be adjusted, making it easier to model 3D facial sequences over long sequences.

[0177] 14, which shows an exemplary structural diagram of an electronic device according to one exemplary embodiment of the present application, which can be realized as a terminal or a server in the above-mentioned embodiments. Here, the electronic device 1400 includes a central processing unit (CPU) 1401, a system memory 1404 including a random access memory 1402 and a read-only memory 1403, and a system bus 1405 connecting the system memory 1404 and the central processing unit 1401. The electronic device 1400 further includes a basic input / output system (I / O) 1406 that supports information transmission between devices in the computer, and a mass storage device 1407 configured to store an operating system 1413, application programs 1414, and other program modules 1415.

[0178] In some embodiments, the basic input / output system 1406 includes a display 1408 configured to display information and an input device 1409, such as a mouse or keyboard, through which a user inputs information. Here, the display 1408 and the input device 1409 are both connected to the central processing unit 1401 by connecting to an input / output controller 1410 on the system bus 1405. The basic input / output system 1406 can further include an input / output controller 1410 for receiving and processing input from a number of other devices, such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1410 can also provide output to a display, printer, or other type of output device.

[0179] The mass storage device 1407 connects to the central processing unit 1401 via a connection to a mass storage controller (not shown) on the system bus 1405. The mass storage device 1407 and its associated computer-readable media provide non-volatile storage for the electronic device 1400. That is, the mass storage device 1407 may include a computer-readable medium (not shown), such as a hard disk or drive.

[0180] Without loss of generality, the computer-readable media may include computer storage media and communication media. Computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes random access memory (RAM), read-only memory (ROM), flash memory or other solid-state drive technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical memory, magnetic tape cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that the computer storage media are not limited to the foregoing. The system memory 1404 and mass storage device 1407 may be collectively referred to as memory.

[0181] The memory stores one or more programs, which are configured to be executed by one or more central processing units 1401, and which contain instructions used to implement the above methods, and the central processing unit 1401 executes the one or more programs to implement the methods according to each of the above method embodiments.

[0182] According to various embodiments of the present application, the electronic device 1400 may also be connected to a remote computer on a network, such as the Internet, and run on the network 1412. That is, the electronic device 1400 may be connected to a network 1412 via a connection to a network interface unit 1411 in the system bus 1405, which in turn may be used to connect to other types of networks or remote computer systems (not shown).

[0183] The memory further includes one or more programs stored in the memory, the one or more programs including steps executed by the electronic device to perform a method according to an embodiment of the present application.

[0184] An embodiment of the present application further provides a computer-readable storage medium, wherein at least one instruction, at least one program, code set or instruction set is stored in the readable storage medium, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to realize the facial modeling model training method described in any of the above embodiments or to realize the 3D face modeling method described in any of the above embodiments.

[0185] An embodiment of the present application provides a computer program product or a computer program including computer instructions stored in a computer-readable storage medium, wherein a processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the electronic device to perform the facial modeling model training method according to the above aspect, or to perform the three-dimensional facial modeling method described in any of the above embodiments.

[0186] Those skilled in the art will understand that all or some of the steps in the various methods of the above embodiments can be performed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, which may be the computer-readable storage medium included in the memory of the above embodiments, or may exist independently without being placed in the computer-readable storage medium of a terminal. The computer-readable storage medium stores at least one instruction, at least one program, code set, or instruction set, which is loaded and executed by a processor to realize the facial modeling model training method described in any of the above method embodiments or the 3D face modeling method described in any of the above embodiments.

[0187] For example, the computer-readable storage medium may include a ROM, a RAM, a solid-state hard disk (SSD), an optical disk, etc. Here, the RAM may include a resistive random access memory (ReRAM) and a dynamic random access memory (DRAM). The numbers of the above embodiments of the present application do not indicate the superiority or inferiority of the embodiments, but are provided for the convenience of explanation.

[0188] Those skilled in the art can understand that all or part of the steps in the above embodiments can be implemented by hardware, and can also be realized by instructing related hardware by a program, and the program can be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disk, etc.

[0189] It should be understood that the term "multiple" referred to herein refers to two or more than two. The term "and / or" indicates a relationship between related objects and indicates that three relationships can exist. For example, A and / or B indicates three cases: when only A exists, when both A and B exist, or when only B exists. The symbol " / " generally indicates that the relationship between related objects before and after it is an "or" relationship. Furthermore, terms such as "first," "second," and the like used herein are used to distinguish between similar objects and are not intended to limit a specific order or sequence. Furthermore, the step numbers used herein exemplify possible execution priorities between steps. In some other embodiments, the steps described above may not be performed in the order shown. For example, two steps with different numbers may be performed simultaneously, or two steps with different numbers may be performed in the opposite order to that shown in the drawings. The embodiments of the present application are not limited thereto.

[0190] The above are merely examples of some embodiments of the present application, and do not limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall all be included in the protection scope of the present application.

Claims

1. 1. A method for training a facial modeling model, performed by an electronic device, comprising: acquiring sample face images, the sample face images including face images of the same object taken at the same time and from different viewpoints; performing feature encoding on the sample face images by an encoder of the face modeling model to obtain face region latent codes corresponding to different face regions and face action latent codes corresponding to different face actions; performing 3D face reconstruction on the object based on the face region latent code and the facial action latent code via a neural radiation field of the face modeling model to obtain a 3D face image of the object; obtaining a 3D reconstruction loss between the 3D face image and the sample face image; and training the face model based on the 3D reconstruction loss.

2. the neural radiation fields include a regional neural radiation field, an operational neural radiation field, and a rendering neural radiation field; performing 3D face reconstruction on the object based on the face region latent code and the facial action latent code via a neural radiation field of the face modeling model to obtain a 3D face image of the object, inputting the facial action latent code into the action neural radiation field to obtain an action transformation matrix; acquiring three-dimensional face coordinates of the object based on the sample face image, and performing motion transformation on the three-dimensional face coordinates based on the motion transformation matrix to obtain face motion coordinates; inputting the face motion coordinates and the face region latent code into the regional neural radiation field to obtain a face region slice plane, wherein a feature dimension of the face region slice plane is higher than a feature dimension of the face region latent code; inputting the facial motion coordinates and the facial region slice plane into the rendering neural radiation field to obtain facial color and facial density; and generating the three-dimensional facial image based on the facial color and the facial density.

3. performing feature encoding on the sample face image by an encoder of the face modeling model to obtain face region latent codes corresponding to different face regions and face action latent codes corresponding to different face actions, performing feature encoding on the sample face image by an encoder of the face model to obtain an upper region latent code, a lower region latent code, and the facial action latent code, wherein the upper region latent code is used to represent facial features of an upper half of a face, and the lower region latent code is used to represent facial features of a lower half of a face; inputting the face motion coordinates and the face region latent code into the regional neural radiation field to obtain a face region slice plane, inputting the facial motion coordinates and the upper region latent code into an upper neural radiation field to obtain an upper region slice plane; inputting the facial motion coordinates and the lower region latent code into a lower neural radiation field to obtain a lower region slice surface; and obtaining the face region slice plane based on the upper region slice plane and the lower region slice plane.

4. The step of obtaining a 3D reconstruction loss between the 3D face image and the sample face image includes: determining a pixel error loss based on pixel differences between the 3D face image and the sample face image; determining a mask error loss based on mask differences between the 3D face image and the sample face image; The method of training a facial modeling model according to claim 1 , further comprising: determining the 3D reconstruction loss based on the pixel error loss and the mask error loss.

5. the encoder includes a coding network and a decoupling neural radiation field; performing feature encoding on the sample face image by an encoder of the face modeling model to obtain face region latent codes corresponding to different face regions and face action latent codes corresponding to different face actions, performing feature coding on the sample facial image via the coding network to obtain a facial state latent code, the facial state latent code being used to represent an overall facial state of an object corresponding to the sample facial image when reading a text; decoupling the facial state latent code via the decoupling neural radiation field to obtain an upper region latent code, a lower region latent code, and the facial action latent code; and determining the facial region latent code based on the upper region latent code and the lower region latent code; 4. The method for training a facial modeling model according to claim 1, wherein the upper region latent code is used to represent facial features of an upper half of a face, and the lower region latent code is used to represent facial features of a lower half of a face.

6. The method for training a facial modeling model includes: determining a decoupling loss based on the upper region latent code, the lower region latent code, and the facial action latent code at different times, wherein the decoupling loss is used to represent differences in facial decoupling at different times, and is used to train the decoupling neural radiation field; training the face modeling model based on the 3D reconstruction loss, The method for training a facial model of claim 5 , further comprising: training the facial model based on the decoupling loss and the 3D reconstruction loss.

7. determining a decoupling loss based on the upper region latent code, the lower region latent code, and the facial action latent code at the different times, obtaining a first upper region latent code, a first lower region latent code, and a first facial movement latent code corresponding to the sample facial image at a first time, and obtaining a second upper region latent code, a second lower region latent code, and a second facial movement latent code corresponding to the sample facial image at a second time; determining an upper decoupling loss based on the first upper region latent code, the first facial expression latent code, and the second lower region latent code; determining a lower decoupling loss based on the first lower region latent code, the first facial expression latent code, and the second upper region latent code; determining a motion decoupling loss based on the second facial motion latent code, the first upper region latent code, and the first lower region latent code; and determining the decoupling loss based on the upper decoupling loss, the lower decoupling loss, and the motion decoupling loss.

8. determining an upper decoupling loss based on the first upper region latent code, the first facial expression latent code, and the second lower region latent code, inputting the first upper region latent code, the first facial action latent code, and the second lower region latent code into the neural radiation field to obtain a first facial image; determining the upper decoupling loss based on a difference between an upper half of a face in the first face image and an upper half of a face in a sample face image at the first time; determining a lower decoupling loss based on the first lower region latent code, the first facial expression latent code, and the second upper region latent code, inputting the first lower region latent code, the first facial action latent code, and the second upper region latent code into the neural radiation field to obtain a second facial image; and determining the lower decoupling loss based on a difference between a lower half of a face in the second face image and a lower half of a face in a sample face image at the first time.

9. determining a motion decoupling loss based on the second facial motion latent code, the first upper region latent code, and the first lower region latent code, inputting the second facial action latent code, the first upper region latent code, and the first lower region latent code into the neural radiation field to obtain a third facial image; and determining the motion decoupling loss based on a difference between a mask of the third facial image and a mask of the sample facial image at the first time.

10. 1. A method for modeling a three-dimensional face performed by an electronic device, comprising: when receiving input text, determining text phonemes corresponding to said input text; querying a target region latent code and a target action latent code corresponding to the text phoneme based on a phoneme latent code index, where the phoneme latent code index indicates a correspondence between a phoneme and a latent code sequence, and the latent code sequence is obtained by performing feature encoding on a face image corresponding to the phoneme based on an encoder of a face model; performing three-dimensional reconstruction based on the target region latent code and the target action latent code through the neural radiation field of the face modeling model to obtain a facial action image corresponding to the text phoneme; generating a three-dimensional facial movement animation corresponding to the input text based on the facial movement image of each frame.

11. The three-dimensional face modeling method includes: acquiring text reading images, the text reading images including facial images of a reader reading the text collected at the same time from different viewpoints; performing feature coding on the text reading image based on the encoder to obtain a face latent code sequence corresponding to the text reading image, the face latent code sequence including a region latent code and an action latent code; 11. The method of claim 10, further comprising the step of: storing the facial latent code sequence in association with a reading phoneme corresponding to the text reading image to obtain the phoneme latent code index.

12. A facial modeling model training device, an image acquisition module configured to acquire sample facial images, the sample facial images including facial images of the same object at the same time but taken from different viewpoints; a feature encoding module configured to perform feature encoding on the sample face image by an encoder of the face modeling model to obtain face region latent codes corresponding to different face regions and face action latent codes corresponding to different facial actions; a 3D reconstruction module configured to perform 3D face reconstruction on the object based on the face region latent code and the facial action latent code via a neural radiation field of the face modeling model to obtain a 3D face image of the object; a model training module configured to obtain a 3D reconstruction loss between the 3D facial image and the sample facial image, and to train the facial modeling model based on the 3D reconstruction loss.

13. A three-dimensional face modeling device, comprising: a phoneme determination module configured to, when receiving an input text, determine text phonemes corresponding to said input text; a latent code lookup module configured to look up a target region latent code and a target action latent code corresponding to the text phoneme based on a phoneme latent code index, wherein the phoneme latent code index indicates a correspondence between a phoneme and a latent code sequence, and the latent code sequence is obtained by performing feature encoding on a face image corresponding to the phoneme based on an encoder of a face model; a 3D reconstruction module configured to perform 3D reconstruction based on the target region latent code and the target action latent code via a neural radiation field of the face modeling model to obtain a facial action image corresponding to the text phoneme; and an animation generation module configured to generate a three-dimensional facial movement animation corresponding to the input text based on the facial movement image of each frame.

14. An electronic device, a memory storing computer-executable instructions; and a processor that executes computer-executable instructions stored in the memory to perform the method for training a facial modeling model of any one of claims 1 to 9 or the method for modeling a three-dimensional face of any one of claims 10 to 11.

15. A computer-readable storage medium having stored thereon computer-executable instructions that, when executed by a processor, cause the processor to perform the facial modeling model training method of any one of claims 1 to 9 or the three-dimensional face modeling method of any one of claims 10 to 11.

16. A computer program product comprising a computer program or computer executable instructions which, when executed by a processor, causes the processor to perform the method for training a facial modeling model according to any one of claims 1 to 9 or the method for modeling a three-dimensional face according to any one of claims 10 to 11.

Citation Information

Patent Citations

  • Generative adversarial neural network assisted video reconstruction

    CN113542759A

  • Communication networks, and devices for converting text to voice and text to video.

    JP2010519791A

  • Image processing device and method, image processing system and program

    JP2022119067A