Face driving model training method and device, electronic equipment, medium and product

By obtaining video training data from Chinese Internet platforms, extracting identity and lip shape feature information, and using generative adversarial networks to train a three-dimensional model, the problem of poor lip shape effect of English audio driving Chinese audio is solved, and high-quality face driving effect is achieved.

CN120707703APending Publication Date: 2025-09-26CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410296695.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the existing technology, the neural network based on English audio cannot effectively drive the corresponding lip shape of Chinese audio, resulting in poor face driving effect.

Method used

By obtaining video training data from Chinese Internet platforms, extracting identity feature information and lip shape feature information, generating a three-dimensional model sequence frame, and using a generative adversarial network to train the model, a three-dimensional model corresponding to the target audio is generated.

Benefits of technology

The face-driven quality of Chinese audio has been improved, the difference between the 3D model simulated lip shape and the target audio lip shape has been reduced, and the face-driven effect has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707703A_ABST
    Figure CN120707703A_ABST
Patent Text Reader

Abstract

The invention discloses a face-driven model training method and device, electronic equipment, a medium and a product. The method comprises the steps that video training data is acquired from a target internet platform, the video training data at least comprises a face image and a target audio, the target audio is the audio corresponding to pronunciation of pictographs, the target internet platform is used for storing video data, and the audio in the video data is the target audio; extracting identity feature information and mouth shape feature information corresponding to the face image from the video training data; generating a three-dimensional model sequence frame corresponding to the target audio based on the identity feature information and the mouth feature information; and training the initial face driving model based on the three-dimensional model sequence frame to obtain a target face driving model. According to the method provided by the invention, face driving based on pictographic audio can be realized, and the quality of face driving is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, and in particular relates to a training method, device, electronic device, medium and product for a face-driven model. Background Art

[0002] Audio-based face driving refers to inferring the possible lip shape based on the audio input by the user, and driving the virtual 3D model based on the lip shape so that the lip shape of the virtual 3D model matches the voice input by the user.

[0003] However, in the related art, in the process of driving faces through audio, English audio is usually used as the audio input object, and in the process of using neural networks to predict lip shapes, English data sets are also used as training data. However, for audio corresponding to the pictographic system (for example, Chinese voice), it is impossible to effectively drive faces, the face driving effect is poor, and the quality of face driving is reduced. Summary of the Invention

[0004] The embodiments of the present application provide a training method, device, electronic device, medium and product for a face driving model, which can realize face driving based on pictographic audio and improve the quality of face driving.

[0005] In a first aspect, an embodiment of the present application provides a training method for a face-driven model, the method comprising: obtaining video training data from a target Internet platform, wherein the video training data includes at least a face image and a target audio, the target audio is the audio corresponding to the pronunciation of a pictogram, the target Internet platform is used to store video data, and the audio in the video data is the target audio; extracting identity feature information and lip shape feature information corresponding to the face image from the video training data; generating a three-dimensional model sequence frame corresponding to the target audio based on the identity feature information and the lip shape feature information; and training an initial face-driven model based on the three-dimensional model sequence frame to obtain a target face-driven model.

[0006] In the second aspect, an embodiment of the present application provides a training device for a face-driven model, which includes: a data acquisition module for acquiring video training data from a target Internet platform, wherein the video training data includes at least a face image and a target audio, the target audio is the audio corresponding to the pronunciation of the pictogram, the target Internet platform is used to store video data, and the audio in the video data is the target audio; an information extraction module for extracting identity feature information and lip feature information corresponding to the face image from the video training data; a face model generation module for generating a three-dimensional model sequence frame corresponding to the target audio based on the identity feature information and the lip feature information; a model training module for training the initial face-driven model based on the three-dimensional model sequence frame to obtain a target face-driven model.

[0007] In a third aspect, an embodiment of the present application provides an electronic device comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the training method of the face driving model as described in the first aspect is implemented.

[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the training method of the face-driven model as described in the first aspect is implemented.

[0009] In a fifth aspect, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the training method of the face driving model as described in the first aspect.

[0010] In the embodiment of the present application, the target audio is the audio corresponding to the pictogram. That is, the target face driving model trained using the method provided in the embodiment of the present application can achieve face driving for the audio corresponding to the pronunciation of the pictogram, thereby improving the face driving effect based on the audio corresponding to the pictogram and improving the quality of face driving. In addition, for the same audio data, objects of different identities have different lip shapes. In the embodiment of the present application, identity feature information and lip shape feature information are extracted from the video training data, the identity feature information and lip shape feature information are separated from each other, and a three-dimensional model sequence frame is generated based on the identity feature information and lip shape feature information to reduce the difference between the simulated lip shape of the three-dimensional model and the lip shape corresponding to the target audio, thereby improving the quality of face driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 This is a flowchart of a method for training a face-driven model provided by one embodiment of the present application;

[0013] Figure 2 This is a principle block diagram of a method for training a face-driven model provided by one embodiment of the present application;

[0014] Figure 3 1 is a structural diagram of a training device for a face-driven model provided in another embodiment of the present application;

[0015] Figure 4 This is a structural diagram of an electronic device provided in yet another embodiment of the present application. DETAILED DESCRIPTION

[0016] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.

[0017] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.

[0018] For ease of understanding, before explaining the solution provided in this application, the background of the solution provided in this application is first explained.

[0019] Audio (e.g., voice) driven three-dimensional face refers to inferring the corresponding lip shape of a virtual three-dimensional model based on any audio segment freely input by the user, and outputting the lip shape animation of the virtual three-dimensional model. In the related art, in the research on lip shape driven audio, the syllables in the audio are usually divided into different phonemes, and the appropriate lip shape is matched to each phoneme. This method is relatively mechanical and cannot adapt to scenarios where different phonemes cause different lip shape changes. With the development of neural networks, in the related art, deep learning methods can be combined to calculate the correspondence between phonemes and lip shapes, or the facial animation of the three-dimensional model can be predicted directly based on the audio band without using phoneme mapping at all.

[0020] However, in related technologies, most audio-driven lip-shape methods use English audio as input to predict the lip shape of a three-dimensional model, thereby achieving face-driven recognition. However, the scheme for mapping phonemes to lip shapes, as well as the lip shape prediction algorithm based on English phonemes, are completely unsuitable for Chinese networks. Subsequent prediction algorithms using neural networks also use English datasets as training data. Even if some algorithms can be used across languages, their effectiveness is more reflected in alphabetical languages ​​with similar origins to English, such as French and Russian, and they cannot accurately predict the lip shapes corresponding to pictographic systems, such as Chinese audio.

[0021] Based on the above content, we can see that in related technologies, there are mainly three voice-based face driving solutions:

[0022] Solution 1: Based on a preset mapping relationship between speech elements and facial drive parameters, the current facial drive parameters corresponding to the speech elements contained in the current speech information are obtained. This method relies on speech element expansion, and different languages ​​contain different speech elements. The speech element classification rules for English are not applicable in the Chinese environment. Furthermore, speech elements are not sensitive to audio speed and volume, making the drive system based on speech elements less effective.

[0023] Solution 2: Coordinating signals between computing devices in a voice-driven computing environment. In this solution, the first and second digital assistants can detect input audio signals, perform signal quality checks, and provide instructions that the first and second digital assistants can operate to process the input audio signals. In this solution, considering only English as input content is not effective in a Chinese-based environment.

[0024] Solution 3: Use the speech-driven dataset to perform several rounds of speech-driven training on the deep learning network model. After training, the speech-driven model is generated and used to create facial and expression animation data. The mouth animation data is then integrated with the facial and expression animation data to render a natural-looking 3D digital human speech-driven animation. This solution does not consider the impact of identity information on lip shape data. In speech-driven tasks, the same speech-driven lip shape will produce different results on different faces.

[0025] In order to solve the problems of the prior art, the embodiments of the present application provide a method, device, electronic device, medium and product for training a face-driven model.

[0026] In the solution provided in the embodiment of the present application, the audio corresponding to the pictogram is used as input data, and the facial movements of the three-dimensional model are inferred based on the two-dimensional video containing audio and images and the above-mentioned audio. It can be seen that the solution provided in the embodiment of the present application does not need to divide the speech elements, and directly extracts features from the audio segments, which can adapt to sound waves of different speaking speeds and different pitch frequencies, and has a higher degree of adaptation to the mouth shape; in addition, in the solution provided in the embodiment of the present application, the audio corresponding to the pictogram is used as input data, which can realize the face driving of Chinese speech in the pictographic system and improve the quality of face driving; finally, in the embodiment of the present application, in the process of training the face driving model, the audio data and identity data in the two-dimensional video data are stripped, which can optimize the driving effect of the face driving model and improve the quality of face driving.

[0027] The following is an introduction to the training method of the face driving model provided in the embodiment of the present application. It should be noted that the method provided in the embodiment of the present application can be applied to but not limited to driving the lip shape of three-dimensional game characters, driving the digital population shape, and driving the virtual population shape. In the following, the face driving platform is used as an example to explain the execution of the method provided in the embodiment of the present application, wherein the face driving platform can be composed of a server and a client, the client is used to receive the user's control instructions, and the server user obtains video data from the Internet platform and uses the video data to train the face driving model.

[0028] Figure 1 FIG. 1 is a flow chart showing a method for training a face-driven model according to an embodiment of the present application. Figure 1 As shown, the method includes the following steps:

[0029] Step S101: Obtain video training data from a target Internet platform.

[0030] In step S101, the video training data is two-dimensional video data, i.e., video data including a facial image and target audio. The target audio is the audio corresponding to the pronunciation of a pictographic character, for example, Chinese speech. The target internet platform is used to store the video data, and the audio in the video data is the target audio. In this embodiment of the present application, the target internet platform can be a Chinese internet platform.

[0031] It should be noted that, in the embodiments of the present application, the target audio is explained as Chinese speech.

[0032] In one example, the face-driven platform can obtain video data that meets a first preset condition from multiple video data stored in the target Internet platform to obtain first video data; then remove the head area whose lip shape information does not match the single audio data from the head area contained in the first video data to obtain second video data; finally, the second video data is identity-labeled based on the facial image of the head area contained in the second video data to obtain video training data.

[0033] In the above embodiment, the first preset condition includes at least: including a head area, the mouth area in the head area is in an unobstructed state, and including a single audio data, wherein the video data including a single audio data means that the video data only contains the audio of a single person speaking, and does not contain any background audio, such as soundtrack.

[0034] In the above embodiment, the target Internet platform is a Chinese Internet platform, which stores various types of video data, such as variety show video data and news video data. The face-driven platform can directly extract video data with clear heads, only one person talking, no background music, and no mouth obstruction from the Chinese Internet platform. Among them, since the host's face is clear and the pronunciation is standard in the news video data, it can be used as video training data. Therefore, in the embodiment of the present application, news video data is obtained from the Chinese Internet platform first. After obtaining the news video data, the face-driven platform uses a preset detection code to extract the head position of the object in the video, removes the area where the audio does not correspond to the object, and marks the identity of the object, and uses the marked video data as video training data.

[0035] Step S102: extracting identity feature information and lip shape feature information corresponding to the face image from the video training data.

[0036] In step S102, the lip shape is not only related to the audio content, but also to the face shape of the subject. For different audio, different subjects will produce different lip movements when speaking. For example, the lip shapes of men and women are different, and the lip shapes of overweight people and thin people are different. To prevent the subject's identity information from affecting the effect of lip shape generation, in this embodiment of the application, after obtaining video training data from the target Internet platform, the subject's identity information in the video training data is also annotated. The subject's identity information will be used in subsequent model training to improve the accuracy of the face-driven data generated by the model, thereby improving the quality of face-driven data.

[0037] It should be noted that in the embodiment of the present application, identity feature information and lip shape feature information are extracted from the video training data, and the identity feature information and lip shape feature information are separated from each other to reduce the difference between the simulated lip shape of the three-dimensional model and the lip shape corresponding to the target audio, thereby improving the quality of face driving.

[0038] Step S103: Generate a three-dimensional model sequence frame corresponding to the target audio based on the identity feature information and the lip shape feature information.

[0039] In step S103, the face driving platform constructs a three-dimensional model for each frame image in the video training data, and the three-dimensional model can be a three-dimensional face model; then, the time corresponding to each frame image is determined as the time corresponding to the three-dimensional model; then, the multiple three-dimensional models corresponding to the video training data are sorted according to the playback order of the images in the video training data to obtain a three-dimensional model sequence frame.

[0040] It should be noted that in the embodiment of the present application, since the video training data consists of facial images and target audio, after constructing the three-dimensional model sequence frames based on the facial images in a certain segment of video data, the three-dimensional model sequence frames can be associated with the audio based on the association relationship between the facial images and the audio, thereby obtaining the three-dimensional model sequence frames corresponding to the audio.

[0041] Step S104: training the initial face driving model based on the 3D model sequence frames to obtain a target face driving model.

[0042] In step S104, after obtaining the three-dimensional model sequence frames corresponding to the video training data, the face driving platform can calculate the difference between the three-dimensional model sequence frames constructed using the method provided in the embodiment of the present application and the real three-dimensional model sequence frames corresponding to the video training data, and adjust the model parameters of the initial face driving model based on the difference, so as to obtain a target face driving model that can accurately predict the mouth shape.

[0043] Based on the scheme defined in steps S101 to S104 above, it can be learned that in the embodiment of the present application, the target audio is the audio corresponding to the pictogram, that is, the target face drive model trained using the method provided in the embodiment of the present application can achieve face drive for the audio corresponding to the pictogram, thereby improving the face drive effect based on the audio corresponding to the pictogram and improving the quality of face drive. In addition, for the same audio data, objects of different identities have different lip shapes. In the embodiment of the present application, identity feature information and lip shape feature information are extracted from the video training data, the identity feature information and lip shape feature information are separated from each other, and a three-dimensional model sequence frame is generated based on the identity feature information and lip shape feature information to reduce the difference between the simulated lip shape of the three-dimensional model and the lip shape corresponding to the target audio, thereby improving the quality of face drive.

[0044] As an example, Figure 2 The principle block diagram of the method provided in the embodiment of the present application is shown. Figure 2As can be seen, the training process of the target face-driven model mainly includes four stages, namely, the data accuracy stage M1, the feature processing stage M2, the feature generation stage M3, and the discriminant generation stage M4. The following is a detailed explanation of each step of the method provided in the embodiment of this application in combination with the above four stages.

[0045] After obtaining the video training data, the face driving platform extracts identity feature information and lip shape feature information corresponding to the face image from the video training data.

[0046] Specifically, the face-driven platform performs data separation on the video training data to obtain two-dimensional video data containing facial images and target audio. It then performs model conversion on the two-dimensional video data through detailed expression capture and the animated DECA network to obtain the initial three-dimensional model sequence frame corresponding to the two-dimensional video data. Next, under the constraints of the identity space loss function, the face-driven platform encodes the identity information contained in the initial three-dimensional model sequence frame to obtain identity feature information. Under the constraints of the audio space loss function, the platform encodes the lip shape information contained in the initial three-dimensional model sequence frame based on the target audio to obtain lip shape feature information.

[0047] In the above embodiment, the initial three-dimensional model sequence frames include at least lip shape information and identity information for identifying the identity of the face in the face image, wherein the identity information is obtained by identity labeling the video data obtained from the target Internet platform before obtaining the video training data.

[0048] As an example, Figure 2 As shown, during the data accuracy phase, the face-driven platform separates the video training data into target audio and facial images without the audio. The facial images are then converted into 3D models using the DECA (Detailed Expression Capture and Animation) network. The 3D models are then sorted based on the playback order of the facial images corresponding to the 2D video data, resulting in an initial 3D model sequence frame.

[0049] It should be noted that the DECA network is a network specifically used for three-dimensional face reconstruction. It can perform three-dimensional face reconstruction on a single input color face image, thereby obtaining a three-dimensional model of the face. At the same time, the DECA network has low computing power requirements and can reach 120fps on an Nvidia Quadro RTX 5000. It can quickly convert the input two-dimensional video into a three-dimensional model, and then obtain a three-dimensional model sequence frame. The converted three-dimensional model sequence frame corresponds to the original video training data.

[0050] Further, such as Figure 2As shown in , after obtaining the initial 3D model sequence frames, the face driving platform processes the initial 3D model sequence frames in the feature processing stage to obtain identity feature information and lip feature information. Figure 2 As shown, in the feature processing stage, the face driving platform extracts identity information from the identity feature space corresponding to the initial three-dimensional model sequence frame, and extracts mouth shape information from the audio feature space corresponding to the initial three-dimensional model sequence frame; then, two encoders are used in the two feature spaces to encode the identity information and mouth shape information corresponding to the initial three-dimensional model sequence frame to obtain identity feature information and mouth shape feature information.

[0051] It should be noted that in different feature spaces, the face-driven platform uses different loss functions for constraints. For example, Figure 2 In the identity feature space, the face-driven platform has a loss function L in the identity space. id The identity feature information is determined under the constraints of the face driving platform in the identity space loss function L w The lip shape feature information is determined under the constraints of

[0052] Specifically, in the identity feature space, the face driving platform encodes the identity information contained in the initial three-dimensional model sequence frames to obtain initial identity feature information; then, based on the initial identity feature information and the preset target identity feature information, the function value of the identity space loss function is calculated; when the function value of the identity space loss function is within a first numerical range, the identity feature information is determined to be the initial identity feature information.

[0053] It should be noted that, in the above embodiment, the function value of the identity space loss function is used to characterize the feature similarity corresponding to the initial identity feature information and the target identity feature information, wherein the function value of the identity space loss function is negatively correlated with the feature similarity corresponding to the identity feature information, that is, the lower the function value of the identity space loss function, the higher the feature similarity corresponding to the identity feature information.

[0054] As an example, Figure 2 As shown, in the identity feature space, the identity feature encoder E id Encode the identity information to obtain identity feature information f id For any identity information (i.e., the i-th identity feature information), the positive pairs with the same identity label are recorded as The negative pairs with different identities are recorded as Then the identity space loss function can be expressed as follows:

[0055]

[0056] From the above expression of identity space loss function, we can see that in the identity space loss function L id Under the constraint of , in the identity feature space, the distance between identity feature information with the same identity tag can be reduced, and the distance between identity feature information with different identity tags can be increased. In the embodiment of the present application, the distance between identity feature information is measured by the cosine function.

[0057] It should be noted that, in practical applications, the distance between identity feature information can be determined by the vector distance between the feature vectors corresponding to the identity feature information, wherein the vector distance between the feature vectors corresponding to the identity feature information can reflect the similarity between the two identity feature information.

[0058] In addition, it should be noted that identity tags can be used to identify objects in the video. The identity tag features can be encoded by the face driving platform according to the object attributes of the object in the video training data according to preset encoding rules. The object attributes may include but are not limited to gender, age group, and physical characteristics (for example, height, shortness, fatness, thinness). After encoding the object in the video to obtain the identity tag, the face driving platform compares the identity tag with the identity tag of the three-dimensional model stored in the face driving platform. If the face driving platform does not store the identity tag, it stores the three-dimensional model corresponding to the identity tag.

[0059] In the audio feature space, the face-driven platform performs audio encoding on the target audio to obtain audio feature information, and performs lip feature encoding on the lip shape information contained in the initial three-dimensional model sequence frames to obtain initial lip shape feature information; then, based on the audio feature information and the initial lip shape feature information, the function value of the audio space loss function is calculated; then, when the function value of the audio space loss function is within a second numerical range, the lip shape feature information is determined to be the initial lip shape feature information.

[0060] It should be noted that, in the above embodiment, the function value of the audio space loss function is used to characterize the feature similarity corresponding to the audio feature information and the initial lip shape feature information. Similar to the identity feature information, the function value of the audio space loss function is negatively correlated with the feature similarity corresponding to the audio feature information and the initial lip shape feature information, that is, the lower the function value of the audio space loss function, the higher the feature similarity corresponding to the audio feature information and the initial lip shape feature information.

[0061] As an example, Figure 2 As shown, in the audio feature space, the lip feature encoder E wv Encode the lip shape information in the initial 3D model sequence frame to obtain the initial lip shape feature information f v ; Audio feature encoder Ewa Encode the target audio to obtain audio feature information f a Similar to identity feature information, for any lip feature information (i.e., the ith lip shape feature information), the positive pair of audio feature information matching it is recorded as Unmatched negative pairs are recorded as Then the audio space loss function can be expressed as follows:

[0062]

[0063] From the above expression of audio space loss function, we can see that in the audio space loss function L w Under the constraints of , in the audio feature space, the characteristic distance between the audio feature information and the lip-shaped feature information that matches the audio feature information can be reduced, and the characteristic distance between the audio feature information and the lip-shaped feature information that does not match the audio feature information can be increased. In this embodiment of the application, the distance between the audio feature information and the lip-shaped feature information is measured using the cosine function.

[0064] Furthermore, after obtaining the identity feature information and lip shape feature information, the face driving platform can generate a three-dimensional model sequence frame corresponding to the target audio.

[0065] In an embodiment of the present application, the face driving platform can process identity feature information and lip feature information through a generator in a generative adversarial network to generate a three-dimensional model sequence frame corresponding to the target audio.

[0066] It should be noted that GAN (Generative Adversarial Nets) includes a generator and a discriminator, wherein the generator is used to generate new data instances, and the discriminator is used to determine whether the data instances are real or generated by the model. The generator and the discriminator compete with each other and finally reach a Nash equilibrium, that is, the data instances generated by the generator are indistinguishable from the real instances, and the discriminator cannot distinguish whether the data instances are real or generated by the generator. In the embodiment of the present application, Figure 2 As shown, in the feature generation stage, the generator G id It is used to generate a three-dimensional model sequence frame based on identity feature information and lip feature information; the discriminator D is used to compare the three-dimensional model sequence frame with the initial three-dimensional model sequence frame obtained by DECA network conversion to obtain a comparison result (i.e., the discrimination result of the discriminator D).

[0067] In the process of training the initial face-driven model based on the three-dimensional model sequence frames to obtain the target face-driven model, the three-dimensional model sequence frames are compared with the initial three-dimensional model sequence frames converted by the DECA network through the discriminator to obtain the comparison results; at the same time, the target loss function corresponding to the initial face-driven model is constructed based on the identity space loss function, the audio space loss function and the loss function corresponding to the generative adversarial network; under the constraint of the target loss function, the model parameters of the initial face-driven model are iterated multiple times based on the comparison results until the target loss function converges to obtain the target face-driven model.

[0068] In the above embodiment, the comparison result is used to represent the similarity between the 3D model sequence frames and the initial 3D model sequence frames.

[0069] As an example, in the embodiment of the present application, the target loss function of the face driving model consists of three parts, namely, the identity space constraint function L for extracting identity feature information id , audio space constraint function L for extracting lip shape feature information and audio feature information w And the adversarial loss function for network generation adversarial training (i.e., the loss function corresponding to the generation adversarial network) L GAN . That is, the target loss function L can be expressed as follows:

[0070] L=L GAN +L id +L w

[0071] In the above formula, L GAN represents the generation adversarial loss function from the generator and the discriminator, which can be expressed as:

[0072]

[0073] In the above formula, z represents random noise. In the embodiment of the present application, z is a 3D model sequence frame; x is real data (i.e., a real instance). In the embodiment of the present application, x is an initial 3D model sequence frame converted by the DECA network; D represents the discriminant function corresponding to the discriminator; G represents the generation function corresponding to the generator; is the mathematical expectation. The larger the value of , the greater the probability that the discriminator believes that x is real data, and the stronger the discriminant ability of the discriminator; the smaller the value of D(G(z)), the stronger the discriminant ability of the discriminator; The larger the value of , the stronger the discriminant ability of the discriminator.

[0074] It should be noted that, from the above content, the target loss function is generally composed of the above three parts. Figure 2 As shown, in Figure 2 In the figure, the thin arrow part indicates that the network participates in training optimization, and continuously updates the parameters through forward propagation and back propagation under the supervision of the loss function; the thick arrow part represents that the network does not participate in training optimization, and this part only performs a forward calculation based on the input data and then remains fixed.

[0075] This concludes the explanation of the face driving method proposed in the embodiments of the present application.

[0076] Furthermore, after obtaining the target face driving model, the face driving platform can use the target face driving model to generate face driving data corresponding to the audio to be played to drive the three-dimensional face model.

[0077] Specifically, the face driving platform first obtains the audio to be played and determines the target three-dimensional model from multiple stored three-dimensional models; then, the audio to be played is input into the target face driving model, and the face driving data generated by the target face driving model is obtained; finally, the target three-dimensional model is face driven based on the face driving data to obtain a face animation.

[0078] As an example, when a user hopes that the mouth shape of the selected target 3D model can match the input voice to be played, the user can input the audio to be played into the face driving platform. The audio to be played can be the audio corresponding to the hieroglyph. In addition, the user can also select the target 3D model to be driven through the face driving platform, wherein the determination of the target 3D model is essentially the determination of identity information. After determining the audio to be played and the target 3D model, the face driving platform inputs the audio to be played into the target face driving model, and then obtains the face driving data output by the target face driving model. The face driving platform can then use the face driving data to drive the selected target 3D model, and the lip shape of the target 3D model matches the audio to be played.

[0079] As can be seen from the above content, the method provided in the embodiment of the present application filters out two-dimensional video data that meets certain conditions from the Chinese Internet and marks the identity information of the objects in the two-dimensional video data. In the data accuracy stage, the face driving platform separates images and audio from the video training data, and then uses the DECA network to convert the image into a three-dimensional model. Then, in the feature generation stage, feature information about identity information and lip shape information is generated in two feature spaces respectively. The generator combines the two feature information to form a new three-dimensional face sequence frame. The discriminator judges the quality of the three-dimensional face sequence frame to determine whether to generate a face driving model.

[0080] The method provided in the embodiments of the present application has at least the following advantages:

[0081] (1) In response to the problem of poor Chinese three-dimensional voice driving effect in related technologies, this application uses video data containing Chinese audio (i.e., audio corresponding to hieroglyphs) as video training data to realize the training of the face driving model, so that the trained face driving model can drive the three-dimensional face model based on Chinese audio, which improves the quality of face driving compared with related technologies.

[0082] (2) In the embodiment of the present application, video data with clear pronunciation, natural expression, and no music or noise is collected from the Chinese Internet, and then the audio and video data are separated, and audio data and image data are extracted from them. The audio data is used as input audio, and the image data is directly converted into a three-dimensional model using a DECA network as input for the lip shape data of the three-dimensional model, which solves the problem in related technologies that there is less video training data corresponding to Chinese audio.

[0083] (3) Since the lip shape is not only related to the audio information, but also to the identity of the speaker, different speakers will produce different effects for the same audio. Therefore, in order to avoid the influence of identity information on the lip shape generation effect, in the embodiment of the present application, different encoders are used to extract identity information and lip shape information respectively, and finally the identity features and lip shape features are used as input to add to the generator to generate a new three-dimensional model, thereby improving the quality of face driving.

[0084] (4) In the embodiment of the present application, identity features and audio features are calculated using two different feature spaces. In the audio feature space, audio features from the audio feature encoder and video features from the video feature encoder and lip-sync feature encoder that contain the same audio content are constrained by the loss function to shorten the feature distance, thereby reducing the error caused by lip-sync prediction; the generator network that receives the features then further reduces the lip-sync error during generative adversarial training.

[0085] (5) In the embodiment of the present application, a generative adversarial network is used to generate a sequence set of lip shape prediction results (i.e., a three-dimensional model sequence frame). The generator of the generative adversarial network receives the identity features and speech features calculated by feature metrics, generates a three-dimensional model sequence set, and the discriminator distinguishes the generated sequence frames from the real sequence frames. Through the generative adversarial process, the output quality of the face-driven model is improved.

[0086] The present application also provides a training device for a face driving model, such as Figure 3 As shown, the device 300 includes: a data acquisition module 301, an information extraction module 302, a face model generation module 303 and a model training module 304.

[0087] A data acquisition module 301 is configured to acquire video training data from a target Internet platform, wherein the video training data includes at least a facial image and target audio, wherein the target audio is audio corresponding to the pronunciation of the pictogram. The target Internet platform is configured to store video data, and the audio in the video data is the target audio.

[0088] Information extraction module 302, for extracting identity feature information and lip shape feature information corresponding to the face image from the video training data;

[0089] A face model generation module 303 is configured to generate a three-dimensional model sequence frame corresponding to the target audio based on the identity feature information and the lip shape feature information;

[0090] The model training module 304 is used to train the initial face driving model based on the 3D model sequence frames to obtain the target face driving model.

[0091] In one example, the information extraction module includes: a data separation module, a model conversion module, a first encoding module, and a second encoding module. The data separation module is used to perform data separation on the video training data to obtain two-dimensional video data containing a facial image and target audio; the model conversion module is used to perform model conversion on the two-dimensional video data through detailed expression capture and an animation DECA network to obtain an initial three-dimensional model sequence frame corresponding to the two-dimensional video data, wherein the initial three-dimensional model sequence frame includes at least lip shape information and identity information for identifying the identity of the face in the facial image; the first encoding module is used to perform identity feature encoding on the identity information contained in the initial three-dimensional model sequence frame under the constraints of an identity space loss function to obtain identity feature information; and the second encoding module is used to perform lip shape feature encoding on the lip shape information contained in the initial three-dimensional model sequence frame based on the target audio under the constraints of an audio space loss function to obtain lip shape feature information.

[0092] In one example, the first encoding module is specifically used to perform identity feature encoding on the identity information contained in the initial three-dimensional model sequence frames to obtain initial identity feature information; based on the initial identity feature information and the preset target identity feature information, the function value of the identity space loss function is calculated, wherein the function value of the identity space loss function is used to characterize the feature similarity corresponding to the initial identity feature information and the target identity feature information; when the function value of the identity space loss function is within a first numerical range, the identity feature information is determined to be the initial identity feature information.

[0093] In one example, the second encoding module is specifically used to perform audio encoding on the target audio to obtain audio feature information; perform lip shape feature encoding on the lip shape information contained in the initial three-dimensional model sequence frames to obtain initial lip shape feature information; based on the audio feature information and the initial lip shape feature information, calculate the function value of the audio space loss function, wherein the function value of the audio space loss function is used to characterize the feature similarity corresponding to the audio feature information and the initial lip shape feature information; when the function value of the audio space loss function is within a second numerical range, determine that the lip shape feature information is the initial lip shape feature information.

[0094] In one example, a generative adversarial network includes a generator and a discriminator, wherein the face model generation module is specifically used to process identity feature information and lip shape feature information through the generator in the generative adversarial network to generate a three-dimensional model sequence frame corresponding to the target audio.

[0095] In one example, the model training module is specifically used to compare the three-dimensional model sequence frames and the initial three-dimensional model sequence frames converted by the DECA network through a discriminator to obtain a comparison result, wherein the comparison result is used to characterize the degree of similarity between the three-dimensional model sequence frames and the initial three-dimensional model sequence frames; based on the identity space loss function, the audio space loss function and the loss function corresponding to the generative adversarial network, a target loss function corresponding to the initial face-driven model is constructed; under the constraint of the target loss function, the model parameters of the initial face-driven model are iterated multiple times based on the comparison results until the target loss function converges to obtain the target face-driven model.

[0096] In one example, the training device for the face driving model also includes: a face driving module, which is used to obtain audio to be played and determine a target three-dimensional model from multiple stored three-dimensional models; input the audio to be played into the target face driving model to obtain face driving data generated by the target face driving model; and perform face driving on the target three-dimensional model based on the face driving data to obtain a face animation.

[0097] The training device for the face driving model provided in the embodiment of the present application can implement each process implemented in the aforementioned method embodiment. To avoid repetition, it will not be described here.

[0098] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0099] Figure 4 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.

[0100] The electronic device may include a processor 401 and a memory 402 storing computer program instructions.

[0101] Specifically, the processor 401 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0102] Memory 402 may include a large capacity memory for data or instructions. By way of example and not limitation, memory 402 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, memory 402 is a non-volatile solid-state memory.

[0103] The memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.

[0104] The processor 401 implements any one of the face-driven model training methods in the above embodiments by reading and executing computer program instructions stored in the memory 402 .

[0105] In one example, the electronic device may further include a communication interface 403 and a bus 410. Figure 4 As shown, the processor 401 , the memory 402 , and the communication interface 403 are connected via a bus 410 and communicate with each other.

[0106] The communication interface 403 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0107] Bus 410 comprises hardware, software or both, couples the parts of electronic equipment to each other.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations.In suitable cases, bus 410 can comprise one or more buses.Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.

[0108] In addition, in conjunction with the face-driven model training method in the above embodiments, embodiments of the present application may provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when executed by a processor, the computer program instructions implement any of the face-driven model training methods in the above embodiments.

[0109] In addition, in conjunction with the face drive model training method in the above embodiments, embodiments of the present application may provide a computer program product for implementation. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes and implements any of the face drive model training methods in the above embodiments.

[0110] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.

[0111] The functional modules shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in unit, a function card or the like. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0112] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0113] The above describes various aspects of the present disclosure with reference to the flowcharts and / or block diagrams of the training methods, devices, electronic devices, media and products of the face-driven model according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more boxes in the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by special-purpose hardware that performs the specified function or action, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0114] The above description is only a specific embodiment of the present application. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the scope of protection of the present application.

Claims

1. A method for training a face-driven model, characterized in that: include: Obtaining video training data from a target Internet platform, wherein the video training data includes at least a facial image and target audio, the target audio being audio corresponding to the pronunciation of the pictogram; the target Internet platform is used to store the video data, and the audio in the video data is the target audio; Extracting identity feature information and lip shape feature information corresponding to the face image from the video training data; generating a three-dimensional model sequence frame corresponding to the target audio based on the identity feature information and the lip shape feature information; An initial face driving model is trained based on the three-dimensional model sequence frames to obtain a target face driving model.

2. The method according to claim 1, characterized in that Extracting identity feature information and lip shape feature information corresponding to the face image from the video training data includes: Performing data separation on the video training data to obtain two-dimensional video data containing the face image and the target audio; Performing model conversion on the two-dimensional video data through detailed expression capture and animation DECA network to obtain an initial three-dimensional model sequence frame corresponding to the two-dimensional video data, wherein the initial three-dimensional model sequence frame includes at least lip shape information and identity information for identifying the identity of the face in the face image; Under the constraint of the identity space loss function, performing identity feature encoding on the identity information contained in the initial three-dimensional model sequence frames to obtain the identity feature information; Under the constraint of the audio space loss function, lip shape feature encoding is performed on the lip shape information contained in the initial 3D model sequence frames based on the target audio to obtain the lip shape feature information.

3. The method according to claim 2, characterized in that Under the constraint of the identity space loss function, identity feature encoding is performed on the identity information contained in the initial three-dimensional model sequence frames to obtain the identity feature information, including: Performing identity feature encoding on the identity information contained in the initial three-dimensional model sequence frames to obtain initial identity feature information; Calculating a function value of the identity space loss function based on the initial identity feature information and the preset target identity feature information, wherein the function value of the identity space loss function is used to represent the feature similarity corresponding to the initial identity feature information and the target identity feature information; When the function value of the identity space loss function is within a first numerical range, the identity feature information is determined to be the initial identity feature information.

4. The method according to claim 2, characterized in that Under the constraint of the audio space loss function, performing lip shape feature encoding on the lip shape information contained in the initial 3D model sequence frames based on the target audio to obtain the lip shape feature information, including: Performing audio encoding on the target audio to obtain audio feature information; Performing lip shape feature encoding on the lip shape information contained in the initial three-dimensional model sequence frames to obtain initial lip shape feature information; Calculating a function value of the audio space loss function based on the audio feature information and the initial lip shape feature information, wherein the function value of the audio space loss function is used to represent the feature similarity corresponding to the audio feature information and the initial lip shape feature information; When the function value of the audio space loss function is within a second numerical range, the lip-shape feature information is determined to be the initial lip-shape feature information.

5. The method according to claim 2, characterized in that The generative adversarial network includes a generator and a discriminator, wherein generating a three-dimensional model sequence frame corresponding to the target audio based on the identity feature information and the lip feature information includes: The identity feature information and the lip shape feature information are processed by a generator in the generative adversarial network to generate a three-dimensional model sequence frame corresponding to the target audio.

6. The method according to claim 5, characterized in that Training an initial face driving model based on the three-dimensional model sequence frames to obtain a target face driving model includes: Comparing the three-dimensional model sequence frames with the initial three-dimensional model sequence frames converted by the DECA network through the discriminator to obtain a comparison result, wherein the comparison result is used to represent the similarity between the three-dimensional model sequence frames and the initial three-dimensional model sequence frames; Constructing a target loss function corresponding to the initial face-driven model based on the identity space loss function, the audio space loss function, and the loss function corresponding to the generative adversarial network; Under the constraint of the target loss function, the model parameters of the initial face-driven model are iterated multiple times based on the comparison result until the target loss function converges, thereby obtaining the target face-driven model.

7. The method according to claim 1, characterized in that After training the initial face driving model based on the three-dimensional model sequence frames to obtain the target face driving model, the method further includes: Obtaining audio to be played, and determining a target three-dimensional model from a plurality of stored three-dimensional models; Inputting the audio to be played into the target face driving model to obtain face driving data generated by the target face driving model; The target three-dimensional model is face driven based on the face driving data to obtain a face animation.

8. A training device for a face-driven model, characterized in that: include: a data acquisition module, configured to acquire video training data from a target Internet platform, wherein the video training data includes at least a facial image and target audio, wherein the target audio is audio corresponding to the pronunciation of the pictogram; the target Internet platform is configured to store video data, and the audio in the video data is the target audio; An information extraction module, configured to extract identity feature information and lip shape feature information corresponding to the face image from the video training data; A face model generation module, configured to generate a three-dimensional model sequence frame corresponding to the target audio based on the identity feature information and the lip shape feature information; The model training module is used to train the initial face driving model based on the three-dimensional model sequence frame to obtain the target face driving model.

9. An electronic device, characterized in that: The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the training method of the face driving model as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that Computer program instructions are stored on a computer-readable storage medium, and when the computer program instructions are executed by a processor, the training method of the face-driven model according to any one of claims 1 to 7 is implemented.

11. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the training method of the face-driven model as described in any one of claims 1 to 7.