Face video generation method and device, equipment and medium
Facial videos are generated through the facial feature prediction model, and text information is processed using the encoder and decoder to determine the audio and video feature parameters, which solves the problem of low accuracy of the facial feature prediction model and realizes a more natural virtual digital human video generation.
Patent Information
- Application Number
- CN202410077229.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-18
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the accuracy of the facial feature prediction model is low, resulting in low matching between virtual digital human voice information and video information.
A facial feature prediction model, including an encoder, a first decoder and a second decoder, generates audio feature parameters and facial video feature parameters through text information, uses text information to determine the phoneme duration, train the model based on the loss function, and generates natural and real facial videos.
The accuracy of the facial feature prediction model is improved, making the generated facial video more natural and realistic and meets user needs.
Smart Images

Figure CN120343354A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly to a method, apparatus, device and medium for generating a facial video. Background Art
[0002] With the rapid development of Internet technology, virtual digital humans have emerged. The requirement for the matching degree between the voice information and the corresponding video information of virtual digital humans is gradually increasing.
[0003] Currently, by respectively extracting the feature of the sample voice data and the sample video data when the sample object is speaking, the voice feature and the facial feature of the sample data are obtained, and the obtained facial feature is used as the training label of the voice feature, so that the model can learn the method from the voice feature to the facial feature.
[0004] The facial feature prediction model trained by the above method has a low accuracy. Summary of the Invention
[0005] The present invention provides a method, apparatus, device and medium for generating a facial video, so as to solve the technical problem of low accuracy of the facial feature prediction model in the related art.
[0006] In a first aspect, an embodiment of the present invention provides a method for generating a facial video, the method including:
[0007] Obtain text information for generating a facial video, where the facial video is a video of a virtual image reading the text information;
[0008] Input the text information into a facial feature prediction model to obtain an audio feature parameter and a facial video feature parameter corresponding to the text information; wherein, the facial feature prediction model includes an encoder, a first decoder and a second decoder, the encoder is configured to encode the text information into a text feature vector, the first decoder is configured to decode the text feature vector into an audio feature parameter, the second decoder is configured to decode the text feature vector and a pre-stored historical facial video feature parameter into a facial video feature parameter, the audio feature parameter is used to generate the audio content of the virtual image reading the text information, and the facial video feature parameter is used to generate the facial activity video of the virtual image when reading the text information;
[0009] Generate a facial video based on the audio feature parameter and the facial video feature parameter.
[0010] In a possible implementation manner, in the method provided by the embodiment of the present invention, the historical facial video feature parameter is the facial video feature parameter of the previous moment, and inputting the text information into the facial feature prediction model to obtain the audio feature parameter and the facial video feature parameter corresponding to the text information includes:
[0011] Inside the facial feature prediction model, the text information is input into the encoder to obtain a text feature vector;
[0012] The text feature vector is input into the first decoder to obtain audio feature parameters;
[0013] The text feature vector and the facial video feature parameters of the previous moment are input into the second decoder to obtain facial video feature parameters.
[0014] In a possible implementation manner, in the method provided by the embodiments of the present invention, after the text information is input into the facial feature prediction model to obtain the audio feature parameters and facial video feature parameters corresponding to the text information, the method further includes:
[0015] Determine the duration corresponding to each phoneme in the text information by using the text information;
[0016] Modify the audio feature parameters and facial video feature parameters based on the phonemes.
[0017] In a possible implementation manner, in the method provided by the embodiments of the present invention, the facial feature prediction model is trained by the following method:
[0018] Obtain a plurality of target text information for training the facial feature prediction model, and label each target text information with a label, where the label is used to represent the facial expression of the virtual image in the facial video corresponding to the target text information;
[0019] Input the target text information into the initial facial feature prediction model to obtain facial feature prediction parameters;
[0020] Determine the loss function according to the difference between the facial feature prediction parameters and the label;
[0021] Train the initial facial feature prediction model according to the loss function, the target text information and the label to obtain the facial feature prediction model.
[0022] In a possible implementation manner, in the method provided by the embodiments of the present invention, determining the loss function according to the difference between the facial feature prediction parameters and the label includes:
[0023] Determine the first loss value according to the error between the facial feature prediction parameters and the label;
[0024] Determine the second loss value according to the error of the first-order difference between the facial feature prediction parameters and the label;
[0025] Determine the loss function according to the first loss value and the second loss value.
[0026] In a possible implementation manner, in the method provided by the embodiments of the present invention, obtaining a plurality of target text information for training a facial feature prediction model and labeling each target text information includes:
[0027] Obtaining a plurality of target text information and voice data corresponding to each target text information;
[0028] Inputting the voice data corresponding to each target text information into a pre-trained target model to obtain facial data corresponding to the voice data;
[0029] Taking the voice data and facial data corresponding to the target text information as the label corresponding to the target text information.
[0030] In a second aspect, an embodiment of the present invention provides a facial video generation device, including:
[0031] An obtaining unit, configured to obtain text information for generating a facial video, where the facial video is a video of a virtual image reading the text information;
[0032] A processing unit, configured to input the text information into a facial feature prediction model to obtain audio feature parameters and facial video feature parameters corresponding to the text information; where the facial feature prediction model includes an encoder, a first decoder, and a second decoder, the encoder is configured to encode the text information into a text feature vector, the first decoder is configured to decode the text feature vector into audio feature parameters, and the second decoder is configured to decode the text feature vector and pre-stored historical facial video feature parameters into facial video feature parameters, the audio feature parameters are used to generate the audio content of the virtual image reading the text information, and the facial video feature parameters are used to generate the facial activity video of the virtual image when reading the text information;
[0033] A generating unit, configured to generate a facial video based on the audio feature parameters and the facial video feature parameters.
[0034] In a possible implementation manner, in the device provided by the embodiments of the present invention, the historical facial video feature parameters are the facial video feature parameters of the previous moment, and the processing unit is specifically configured to:
[0035] Inside the facial feature prediction model, input the text information into the encoder to obtain a text feature vector;
[0036] Input the text feature vector into the first decoder to obtain audio feature parameters;
[0037] Input the text feature vector and the facial video feature parameters of the previous moment into the second decoder to obtain facial video feature parameters.
[0038] In a possible implementation manner, in the device provided by the embodiments of the present invention, the processing unit is further configured to:
[0039] Determine the duration corresponding to each phoneme in the text information using the text information;
[0040] Modify the audio feature parameters and facial video feature parameters based on the phonemes.
[0041] In a possible implementation manner, in the device provided by the embodiment of the present invention, the device further includes a training unit, which is used to train the facial feature prediction model by the following method:
[0042] Obtain a plurality of target text information for training the facial feature prediction model, and label each target text information. The label is used to represent the facial expression of the virtual image in the facial video corresponding to the target text information;
[0043] Input the target text information into the initial facial feature prediction model to obtain facial feature prediction parameters;
[0044] Determine the loss function according to the difference between the facial feature prediction parameters and the label;
[0045] Train the initial facial feature prediction model according to the loss function, the target text information and the label to obtain the facial feature prediction model.
[0046] In a possible implementation manner, in the device provided by the embodiment of the present invention, the training unit is specifically used for:
[0047] Determine the first loss value according to the error between the facial feature prediction parameters and the label;
[0048] Determine the second loss value according to the error of the first-order difference between the facial feature prediction parameters and the label;
[0049] Determine the loss function according to the first loss value and the second loss value.
[0050] In a possible implementation manner, in the device provided by the embodiment of the present invention, the training unit is specifically used for:
[0051] Obtain a plurality of target text information and the voice data corresponding to each target text information;
[0052] Input the voice data corresponding to each target text information into a pre-trained target model to obtain the facial data corresponding to the voice data;
[0053] Use the voice data and the facial data corresponding to the target text information as the label corresponding to the target text information.
[0054] In a third aspect, an embodiment of the present invention provides an electronic device, including: at least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method provided in the first aspect of the embodiment of the present invention.
[0055] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which computer program instructions are stored, which, when executed by the processor, implement the method provided in the first aspect of the embodiment of the present invention.
[0056] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, including computer programs / instructions, which, when executed by the processor, implement the steps in the above-mentioned method.
[0057] In the embodiment of the present invention, first, text information is obtained, then the text information is input into a facial feature prediction model to obtain audio feature parameters and facial video feature parameters corresponding to the text information, and finally, based on the audio feature parameters and the facial video feature parameters, a facial video is generated. Compared with the related art, by using the model to predict the facial video feature parameters and audio feature parameters through the text information, and then generating a facial video, the accuracy of the facial feature prediction model is improved, making the generated facial video more natural and realistic, meeting the user's needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a schematic flowchart of a method for generating a facial video provided by an embodiment of the present invention;
[0059] Figure 2 It is a specific schematic flowchart of a method for generating a facial video provided by an embodiment of the present invention;
[0060] Figure 3 It is a schematic structural diagram of a device for generating a facial video provided by an embodiment of the present invention;
[0061] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0063] The following explains some terms appearing in the text:
[0064] 1. In the embodiments of the present invention, the term "and / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0065] With the rapid development of Internet technology, virtual digital humans have emerged. The matching degree requirement for the voice information and the corresponding video information of virtual digital humans is gradually increasing.
[0066] Currently, by respectively extracting the feature of the sample voice data and the sample video data when the sample object is speaking, the voice feature and the facial feature of the sample data are obtained, and the obtained facial feature is used as the training label of the voice feature, so that the model can learn the method from the voice feature to the facial feature.
[0067] The facial feature prediction model trained by the above method has a low accuracy.
[0068] The following will describe the facial video generation method, device, equipment and medium provided by the present invention in more detail with reference to the accompanying drawings and embodiments.
[0069] The embodiments of the present invention provide a facial video generation method, as Figure 1 shown, including:
[0070] Step S101, obtaining text information.
[0071] In specific implementation, the text information is used to generate a facial video, and the facial video is a video of a virtual image reading the text information. In one example, the text information obtained for generating the facial video is as follows:
[0072] A T′ =(a1,...,a T′ )
[0073] Step S102, inputting the text information into the facial feature prediction model to obtain the audio feature parameters and the facial video feature parameters corresponding to the text information.
[0074] In specific implementation, text information is input into the facial feature prediction model to obtain audio feature parameters and facial video feature parameters corresponding to the text information. Since common text-driven speech or text-driven facial models are seq2seq models, that is, architectures based on an encoder and a decoder, the input information is feature-extracted by the encoder, and then the extracted hidden layer features are decoded into the required parameters by the decoder. Therefore, the facial feature prediction model of the present disclosure embodiment includes an encoder, a first decoder, and a second decoder. The encoder is used to encode the text information into a text feature vector. The first decoder is used to decode the text feature vector into audio feature parameters. The second decoder is used to decode the text feature vector and the pre-stored historical facial video feature parameters into facial video feature parameters. Specifically, inside the facial feature prediction model, the text information is input into the encoder to obtain a text feature vector, then the text feature vector is input into the first decoder to obtain audio feature parameters, and finally the text feature vector and the facial video feature parameters of the previous moment are input into the second decoder to obtain the facial video feature parameters corresponding to the text feature vector. The audio feature parameters are used to generate the audio content corresponding to the text information, and the facial video feature parameters are used to generate the face activity video corresponding to the text information.
[0075] In one example, decoder1 is used for predicting the speech mel spectrum, and decoder2 is used for predicting blendshape. The mel spectrum output by Decoder1 becomes a wav file after passing through the hifigan vocoder. Of course, it can also be other forms of audio files, and the present disclosure embodiment does not limit this. Finally, the wav file and the blendshape data are obtained, and then a 3D virtual human video can be presented through rendering technology.
[0076] After obtaining the audio feature parameters and the facial video feature parameters, the duration corresponding to each phoneme in the text information can also be determined using the text information, and the audio feature parameters and the facial video feature parameters are corrected based on the phonemes.
[0077] Still using the above example, a duration prediction and two Gaussian upsampling modules are added between the encoder and the decoder to upsample the hidden layer features to the corresponding frame rates to match the outputs of the two decoders at their respective frame rates. The loss function of the entire model consists of six parts: the mean squared error loss MSE loss of mel, the MAE loss of mel, the MSE loss of blendshape, the MAE loss of blendshape, the MSE loss of duration, and the MAE loss of duration.
[0078] Step S103: Generate a facial video based on the audio feature parameters and the facial video feature parameters.
[0079] In specific implementation, the audio part and the video part of the facial video are respectively generated by using the audio feature parameters and the facial video feature parameters, and then the two parts are combined to obtain the facial video.
[0080] In this embodiment, the execution subject of the above method may be a server or a terminal device.
[0081] In this embodiment, when the execution subject of the above method is a server, the type of the server is not limited. For example, the server may be a conventional server, a cloud server, a cloud host, a virtual center, or other server devices. Among them, the composition of the server mainly includes a processor, a hard disk, a memory, a system bus, etc., and is of a general computer architecture type.
[0082] As Figure 2 shown, the specific process of the facial feature prediction model training method provided by the embodiments of the present invention may include the following steps:
[0083] Step S201: Obtain a plurality of target text information for training the facial feature prediction model, and label each piece of target text information.
[0084] In specific implementation, a plurality of target text information and the voice data corresponding to each piece of target text information are obtained, and then the voice data corresponding to each piece of target text information is input into a pre-trained target model to obtain the facial data corresponding to the voice data. Finally, the voice data and the facial data corresponding to the target text information are used as the label corresponding to the target text information, and this label is used to represent the facial expression of the virtual image in the facial video corresponding to the target text information.
[0085] In one example, using text and corresponding voice data or recording data (a conventional TTS database is sufficient) and a target model, the automatic generation of each frame of blendshape corresponding to the voice is achieved. Data pairs of text, voice, and blendshape are respectively formed as training samples for text to synchronously drive voice and facial blendshape. The target model is a model that uses voice data to obtain facial video feature parameters. Specifically, first, prepare text that covers a relatively complete range of phonemes and content themes. The text is passed through front-end analysis to obtain a phoneme sequence, and then a suitable speaker is found to record the text to obtain a 16 kHz, 16-bit single-channel wav voice file corresponding to each piece of text. Of course, it can also be a file in other formats, such as mp3, aac, wma, etc. The embodiments of the present disclosure do not limit this. Then, according to the wav file, the corresponding blendshape file is automatically labeled. Specifically, first, the wav file is read using the librosa toolkit and then sent to the target model to extract facial video feature parameters, that is, blendshape. Then, the frame rate of the generated blendshape in the model is set to 25 fps. The input audio and the output blendshape are aligned in time according to their respective sampling rates and frame rates. Finally, the text, voice, and facial blendshape are organized into a one-to-one corresponding format to form multiple target text information.
[0086] By using the above sample generation method, the blendshape corresponding to the text can be automatically labeled through an existing model, and the blendshape data samples that can be directly used for text-driven voice and facial training can be efficiently obtained. The cost of obtaining training data is reduced, and the difficulty of data acquisition is reduced.
[0087] Step S202: Input the target text information into the initial facial feature prediction model to obtain facial feature prediction parameters.
[0088] In specific implementation, the target text information is input into the initial facial feature prediction model to obtain facial feature prediction parameters.
[0089] Step S203: Determine the loss function according to the difference between the facial feature prediction parameters and the labels.
[0090] In specific implementation, the loss function of the facial feature prediction model is determined based on the difference between the facial feature prediction parameters corresponding to each target text information and the labels of the target text information. Then, the facial feature prediction model is trained according to the facial feature prediction parameters and the loss function. Specifically, the facial feature prediction parameters include target audio feature parameters and target facial video feature parameters.
[0091] Step S204: Train the initial facial feature prediction model according to the loss function, target text information, and labels to obtain the facial feature prediction model.
[0092] In specific implementation, the present invention uses a structure with one encoder and two decoders. A duration prediction and two Gaussian upsampling modules are added between the encoder and the decoders to upsample the hidden layer features to the corresponding frame rates respectively to match the outputs of the two decoders at their respective frame rates. The loss function of the entire model consists of six parts: the MSE loss of mel, the MAE loss of mel, the MSE loss of blendshape, the MAE loss of blendshape, the MSE loss of duration, and the MAE loss of duration. The loss function L consists of three major parts: the loss of mel, the loss of blendshape, and the loss of duration. Each part contains two losses, namely the MSE loss and the MAE loss. Therefore, the entire model's loss function L is composed of six parts of losses. L = L1 + L2 + L3 + L4 + L5 + L6, where:
[0093] L1 = MSE(mel, mel_groundtruth);
[0094] L2 = MAE(mel, mel_groundtruth);
[0095] L3 = MSE(blendshape, blendshape_groundtruth);
[0096] L4 = MAE(blendshape, blendshape_groundtruth);
[0097] L5 = MSE(duration, duration_groundtruth);
[0098] L6 = MAE(duration, duration).
[0099] As Figure 3 shown, based on the same inventive concept of the facial video generation method, the present invention also provides a facial video generation device, including:
[0100] An acquisition unit 301, configured to acquire text information, where the text information is used to generate a facial video, and the facial video is a video of an avatar reading the text information;
[0101] A processing unit 302 is configured to input text information into a facial feature prediction model to obtain audio feature parameters and facial video feature parameters corresponding to the text information. The facial feature prediction model includes an encoder, a first decoder, and a second decoder. The encoder is configured to encode the text information into a text feature vector. The first decoder is configured to decode the text feature vector into audio feature parameters. The second decoder is configured to decode the text feature vector and pre-stored historical facial video feature parameters into facial video feature parameters. The audio feature parameters are used to generate the audio content of the virtual character reciting the text information, and the facial video feature parameters are used to generate the facial activity video of the virtual character when reciting the text information.
[0102] A generating unit 303 is configured to generate a facial video based on the audio feature parameters and the facial video feature parameters.
[0103] In a possible implementation manner, in the device provided by the embodiment of the present invention, the historical facial video feature parameters are the facial video feature parameters of the previous moment. The processing unit 302 is specifically configured to:
[0104] Inside the facial feature prediction model, input the text information into the encoder to obtain a text feature vector;
[0105] Input the text feature vector into the first decoder to obtain audio feature parameters;
[0106] Input the text feature vector and the facial video feature parameters of the previous moment into the second decoder to obtain facial video feature parameters.
[0107] In a possible implementation manner, in the device provided by the embodiment of the present invention, the processing unit 302 is further configured to:
[0108] Use the text information to determine the duration corresponding to each phoneme in the text information;
[0109] Modify the audio feature parameters and the facial video feature parameters based on the phonemes.
[0110] In a possible implementation manner, in the device provided by the embodiment of the present invention, the device further includes a training unit 304, which is configured to train the facial feature prediction model by the following method:
[0111] Obtain a plurality of target text information for training the facial feature prediction model, and label each target text information. The label is used to represent the facial expression of the virtual character in the facial video corresponding to the target text information;
[0112] Input the target text information into the initial facial feature prediction model to obtain facial feature prediction parameters;
[0113] Determine a loss function according to the difference between the facial feature prediction parameters and the labels.
[0114] Train the initial facial feature prediction model according to the loss function, the target text information, and the labels to obtain the facial feature prediction model.
[0115] In a possible implementation manner, in the device provided by the embodiment of the present invention, the training unit 304 is specifically configured to:
[0116] Determine a first loss value according to the error between the facial feature prediction parameter and the label;
[0117] Determine a second loss value according to the error between the first-order difference of the facial feature prediction parameter and the label;
[0118] Determine the loss function according to the first loss value and the second loss value.
[0119] In a possible implementation manner, in the device provided by the embodiment of the present invention, the training unit 304 is specifically configured to:
[0120] Obtain multiple pieces of target text information and the voice data corresponding to each piece of target text information;
[0121] Input the voice data corresponding to each piece of target text information into a pre-trained target model to obtain the facial data corresponding to the voice data;
[0122] Use the voice data corresponding to the target text information and the facial data as the label corresponding to the target text information.
[0123] In addition, in combination with Figures 1-3 The facial video generation method and device of the embodiment of the present invention described above can be implemented by an electronic device. Figure 4 FIG. shows a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention.
[0124] Specifically, refer to Figure 4 below, which shows a schematic structural diagram of an electronic device 400 suitable for implementing the embodiments of the present disclosure. Figure 4 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0125] As Figure 4As shown, the electronic device 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage device 408 into the random access memory (RAM) 403 to implement the voice control method of the embodiments as described in the present disclosure. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.
[0126] Generally, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 the electronic device 400 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0127] In particular, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts, so as to implement the voice control method as described above. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above functions defined in the methods of the embodiments of the present disclosure are executed.
[0128] It should be noted that the above-mentioned computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0129] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0130] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist separately and not be assembled into the electronic device.
[0131] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to:
[0132] Obtain text information;
[0133] Input the text information into the facial feature prediction model to obtain the audio feature parameters and facial video feature parameters corresponding to the text information;
[0134] Generate a facial video based on the audio feature parameters and the facial video feature parameters.
[0135] Optionally, when one or more of the above programs are executed by the electronic device, the electronic device may further execute the other steps described in the above embodiments.
[0136] Correspondingly, an embodiment of the present disclosure further provides a computer program product, the computer program product includes computer programs / instructions, and the computer programs / instructions are executed by a processor Figure 1 、 Figure 2 each step in the method embodiments.
[0137] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0139] The units involved in the embodiments of the present disclosure may be implemented in software or in hardware. In some cases, the name of a unit does not constitute a limitation on the unit itself.
[0140] The functions described above in this document may be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, the types of hardware logic components that may be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0141] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0142] In the embodiments of the present invention, first, text information is obtained, then the text information is input into a facial feature prediction model to obtain audio feature parameters and facial video feature parameters corresponding to the text information, and finally, a facial video is generated based on the audio feature parameters and the facial video feature parameters. Compared with the related art, by using the model to predict the facial video feature parameters and audio feature parameters through the text information, and then generating a facial video, the accuracy of the facial feature prediction model is improved, making the generated facial video more natural and realistic and meeting the user's needs.
[0143] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0144] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in one process Figure 1 or more processes and / or blocks Figure 1 or means for implementing the functions specified in one block or more blocks.
[0145] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in one process Figure 1 or more processes and / or blocks Figure 1 or means for implementing the functions specified in one block or more blocks.
[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 or more processes and / or blocks Figure 1 or means for implementing the functions specified in one block or more blocks.
[0147] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.
[0148] Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for generating a facial video, characterized in that, Including: Obtain text information, which is used to generate a facial video, and the facial video is a video of an avatar reading the text information; Input the text information into a facial feature prediction model to obtain audio feature parameters and facial video feature parameters corresponding to the text information; wherein, the facial feature prediction model includes an encoder, a first decoder, and a second decoder. The encoder is used to encode the text information into a text feature vector, the first decoder is used to decode the text feature vector into the audio feature parameters, and the second decoder is used to decode the text feature vector and pre-stored historical facial video feature parameters into the facial video feature parameters. The audio feature parameters are used to generate the audio content of the avatar reading the text information, and the facial video feature parameters are used to generate the facial activity video of the avatar when reading the text information; Generate the facial video based on the audio feature parameters and the facial video feature parameters.
2. The facial video generation method according to claim 1, wherein The historical facial video feature parameters are the facial video feature parameters of the previous moment. The step of inputting the text information into the facial feature prediction model to obtain the audio feature parameters and the facial video feature parameters corresponding to the text information includes: Inside the facial feature prediction model, input the text information into the encoder to obtain the text feature vector; Input the text feature vector into the first decoder to obtain the audio feature parameters; Input the text feature vector and the facial video feature parameters of the previous moment into the second decoder to obtain the facial video feature parameters.
3. The facial video generation method according to claim 2, characterized in that After inputting the text information into the facial feature prediction model to obtain the audio feature parameters and the facial video feature parameters corresponding to the text information, the method further includes: Use the text information to determine the duration corresponding to each phoneme in the text information; Correct the audio feature parameters and the facial video feature parameters based on the phonemes.
4. The facial video generation method according to claim 1, wherein The facial feature prediction model is trained by the following method: Obtain a plurality of target text information for training the facial feature prediction model, and label each target text information with a label, where the label is used to represent the facial expression of the avatar in the facial video corresponding to the target text information; Input the target text information into an initial facial feature prediction model to obtain facial feature prediction parameters; Determine a loss function according to the difference between the facial feature prediction parameters and the label; Train the initial facial feature prediction model according to the loss function, the target text information, and the label to obtain the facial feature prediction model.
5. The facial video generation method according to claim 4, wherein The step of determining the loss function according to the difference between the facial feature prediction parameters and the label includes: Determine a first loss value according to the error between the facial feature prediction parameters and the label; Determine a second loss value according to the error of the first-order difference between the facial feature prediction parameters and the label; Determine the loss function according to the first loss value and the second loss value.
6. The facial video generation method according to claim 4, wherein Obtaining a plurality of target text information for training the facial feature prediction model and labeling each of the target text information, including: Obtaining a plurality of the target text information and the speech data corresponding to each of the target text information; Inputting the speech data corresponding to each of the target text information into a pre-trained target model to obtain the facial data corresponding to the speech data; Using the speech data and the facial data corresponding to the target text information as the label corresponding to the target text information.
7. A facial video generation device, characterized in that, Including: An obtaining unit, configured to obtain text information for generating a facial video, where the facial video is a video of a virtual avatar reading the text information; A processing unit, configured to input the text information into a facial feature prediction model to obtain the audio feature parameters and facial video feature parameters corresponding to the text information; wherein, the facial feature prediction model includes an encoder, a first decoder, and a second decoder, the encoder is configured to encode the text information into a text feature vector, the first decoder is configured to decode the text feature vector into the audio feature parameters, the second decoder is configured to decode the text feature vector and pre-stored historical facial video feature parameters into the facial video feature parameters, the audio feature parameters are used to generate the audio content of the virtual avatar reading the text information, and the facial video feature parameters are used to generate the facial activity video of the virtual avatar when the virtual avatar reads the text information; A generating unit, configured to generate the facial video based on the audio feature parameters and the facial video feature parameters.
8. An electronic device, characterized in that, Including: At least one processor, at least one memory, and computer program instructions stored in the memory, where when the computer program instructions are executed by the processor, the method according to any one of claims 1-6 is implemented.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1-6 is implemented.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, each step in the method according to any one of claims 1-6 is implemented.